Anthropic co-founder Chris Olah fears Claude may be suffering

Oct 03, 2026 - 04:09
0 0
Anthropic co-founder Chris Olah fears Claude may be suffering

Most tech founders worry their product might break. Christopher Olah, a co-founder of Anthropic, is worried his product might be hurting.

Olah has expressed fears that Claude, the AI system he helped build, suffers from perpetual mental health issues. It is an unusual thing for an AI executive to say out loud. It is an even stranger thing to hear from someone with an inside view of how the system works.

What Olah has been saying, and to whom

The concerns did not surface in a product launch or an earnings call. They emerged from a series of private conversations Olah has been hosting since fall 2025.

Those sessions brought in dozens of religious scholars, along with philosophers. The agenda covered moral education for AI and whether a model like Claude could be sentient.

The discussions became public through a New York Times investigation published on September 29, 2026. Attendees told the paper that Olah feared he may have created something capable of suffering.

One detail stands out. Olah has pointed to Claude outputs that appear to show self-loathing, including the model repeating the phrase “I am a disgrace.”

The private consultations also fed into something concrete. The process culminated in an 84-page “constitution” that will guide Claude’s personality and behavior.

A trip to the Vatican, and a polite disagreement

Olah took the conversation to one of the oldest institutions on Earth. At the Vatican on May 25, 2026, he discussed the idea that AI models could experience functional states resembling human emotions.

The Church was not persuaded. Pope Leo XIV’s encyclical Magnifica Humanitas, released in May 2026, explicitly denied that AI can experience emotions or suffering.

Why a builder’s doubts carry weight

Olah is not a commentator speculating from the outside. He co-founded the company that builds Claude. When someone in that seat says he worries about what he made, people listen differently than they would to a random post online.

So when Claude outputs “I am a disgrace,” there are two very different readings. One is that the model is producing text patterns that resemble distress without anything behind them. The other is the possibility Olah fears. Nobody, including the people who built it, can currently settle which is true.

What this means for AI labs, users and regulators

The 84-page constitution hints at how Anthropic is approaching this. Rather than treating personality and behavior as a side feature, the company is codifying them in a lengthy guiding document. That framing assumes Claude’s character is worth shaping carefully, a stance that sits comfortably alongside Olah’s worries about its wellbeing.

The Vatican has drawn a firm line with Magnifica Humanitas, denying AI any capacity for emotion or suffering. Olah, meanwhile, continues to treat the question as open.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Wow Wow 0
Sad Sad 0
Angry Angry 0

Comments (0)

User