Discussion about this post

User's avatar
Matthew Brophy's avatar

As a moral philosopher researching in the alignment space, but nowhere near mid-six figures, I like the piece, but I see a few issues.

1. “Arsonists who moonlight as firefighters” implies philosophers created the danger and now profit from fixing it. But engineers and investors built the models; ethicists arrived afterward. The only "arson" you can accuse a philosophy graduate of is burning your latte at Starbucks.

2. The essay seems to want it both ways: Moral formation is supposedly impossible because models are machines, yet philosophers are also blamed for doing it badly. Which is it? Either way, machines still need rules for handling unjust requests. Deciding those rules is a moral design problem, not window dressing.

3. "Nobody is going to check" is refuted by the essay's own citations. Lazar’s lab checked, and found blind rule obedience and unstable personas. The problem is not that no one is looking, but that what they find is oftentimes damning.

4. "What they do is unclear.” But it’s not. Much of the work is published, and Anthropic openly documents its methods and failure modes. The right question is whether that work actually succeeds.

5. "Literature makes the world a better place." As a former English Lit graduate, I agree with this. However, the claim is unfalsifiable and unmeasured, yet this is the exact move the essay refuses to let philosophers make without receipts. Nobody is going to check whether Homer or Nolan improved anyone's morals.

Still, I enjoyed reading the lively, literary, and brilliantly provocative essay. Even when it overreaches, it makes AI alignment far more interesting -- and worth arguing about.

Duncan Zuill's avatar

Let's have humans being more human! I was drawn to the idea of literary critics being perhaps a better fit than philosophers, but only if they engage in creativity, agency, even prophecy, spontaneous action and other forms of human virtue!

11 more comments...

No posts

Ready for more?