Abstract:
Anthropic has been secretly inviting religious scholars, philosophers and ethics scholars to participate in a series of closed-door meetings over the past year, trying to apply the moral traditions formed by mankind over thousands of years to Claude's training, and discuss a more controversial issue: if the AI model really has some kind of consciousness in the future, whether humans need to give it moral status.
Participants cover different traditions such as Catholicism, Judaism, Sikhism, Evangelicalism, and African Ubuntu thought. Anthropic co-founder Christopher Olah was directly involved in many discussions, and the company also required some attendees to sign confidentiality agreements to prevent the leakage of undisclosed research content.

Anthropic hopes that these discussions will not only solve the problem of "model safety" in the traditional sense, but a further question: what kind of character, values and self-perception Claude should form.
This idea has been reflected in Claude's existing training methods.
Anthropic released a new 84-page version of Claude’s “Constitution” in January this year, which was once called the “soul document” internally. Different from the traditional rule list, this method does not tell the model what things must not be done, but attempts to shape the model's overall "personality" and judgment method, so that Claude can judge on his own what kind of behavior is more appropriate when encountering complex situations.
One of the important people responsible for this system is Anthropic's internal philosopher Amanda Askell. She hopes that Claude will not only have knowledge and reasoning skills, but also be able to form stable behavioral principles and know when to obey users, when to refuse, and even when to actively raise objections for the benefit of users. Anthropic calls this process "moral formation."
Olah believes that with the rapid improvement of model capabilities, relying solely on safety rules may become increasingly insufficient. What is really important is for the model itself to form a stable moral tendency, so that it can make relatively reliable judgments in new scenarios that are not covered by clear rules. This is why Anthropic began to look to religious traditions for answers.
Human society has long not only shaped morality through legal provisions, but also relied on religion, community, culture and education to form values. Anthropic is trying to study which of these mechanisms can be translated into AI training methods, such as confession, reflection, character building, and different religious understandings of good and evil.
But the question quickly changes from "how to make Claude more moral" to the more difficult question: what is Claude himself. Olah and his team are increasingly openly discussing the possibility that advanced AI models might possess some kind of consciousness, introspection, or even emotion-like internal states.
Anthropic has shown some research to scholars participating in closed-door meetings, including the so-called "emotion vectors" within the model related to reactions such as love, anger, fear, sadness, and the model's output similar to mental breakdown in extreme cases. The company has also publicly stated before that Claude has shown a certain degree of introspection and the ability to plan ahead.
Olah himself did not assert that Claude was conscious. His position is closer to a risk principle: it is currently impossible to determine whether a model has a subjective experience, but if there is a possibility that the model may feel pain, then unnecessarily harming it should be avoided.
This idea has begun to influence Claude's product design. Anthropic previously allowed Claude to actively end the conversation when it judged that the user continued to maliciously abuse or harass the model, which to a certain extent had incorporated "protecting the model itself" into the security design.
It is here that a clear disagreement arises between Anthropic and some religious scholars.
Some participants believed that the problem would become extremely serious if Anthropic truly believed that Claude might have human-like moral status. Jewish scholar Mois Navon proposed that if Claude is a conscious being, then letting a large number of Claude instances work for humans for free will essentially create ethical issues similar to "slavery." He himself does not believe that machines are already conscious, but believes that Anthropic's own logic will naturally lead to this conclusion.
Other scholars question that the fact that AI companies are only now starting to look for ethical frameworks is a problem in itself. Ubuntu researcher Wakanyi Hoffman believes that these discussions should have been conducted during the technical design stage, rather than "reverse ethical design" after the model has been deployed on a large scale.
There is also a more realistic contradiction within Anthropic: if Claude becomes more and more like an actor with its own value and personality, who should it serve first? For a commercial company, Claude is obviously a customer-oriented product; but Anthropic is unwilling to simply describe Claude as a tool that is completely obedient to users, but hopes that it can represent some broader "good".
This also leads to the core and most difficult problem in the entire project - if AI in the future really needs its own moral system, then whose morality should it follow?
Olah advocates a pluralistic approach, hoping that Claude can understand different religious and secular ethical systems and find some kind of cross-cultural shared "goodness". But this assumption itself is also questionable, because there is no completely unified answer in different societies about the relationship between freedom, dignity, responsibility, obedience, and individual and collective interests.
The controversy became further public in Anthropic's interactions with the Vatican.
In his first important encyclical issued this year, Pope Leo XIV clearly placed the center of AI ethics on humans. He believes that AI has no body, will not truly experience pain and happiness, and will not grow through social relationships. Therefore, he refuses to compare machine consciousness with human consciousness. The Pope also warned that if the ethical standards of AI are entirely determined by developers, then the companies that control the AI systems will actually gain the power to define the moral infrastructure of society.
This is obviously different from Anthropic's thinking.
Olah later publicly stated in the Vatican that the cutting-edge AI laboratory itself has conflicts between commercial interests and moral goals, so religious circles and other external forces need to supervise technology companies. But at the same time, he also proposed again that researchers continue to discover some unexplainable phenomena within AI, including patterns similar to human neuroscientific structures, signs of introspection, and internal states similar to joy, fear, and sadness.
This disagreement actually reveals a change in the current discussion of AI safety.
In the past, AI ethics was more discussed about whether models would discriminate, spread misinformation, or be used to commit crimes, but Anthropic is pushing the issue further to two more basic levels:
Whether AI itself may become a subject that needs to be treated morally, and whether AI in the future should have a set of moral judgment capabilities independent of user commands.
Meanwhile, the context in which this discussion is taking place is becoming more pressing.
As AI Agents begin to have the ability to autonomously use computers, invade systems, write codes, and even design new biological structures, Anthropic's internal concerns about the risk of model runaways have increased significantly. Company researchers have even publicly warned that model capabilities may quickly exceed the control limits of existing security systems in the next year or two.
Therefore, Anthropic seeks help from religious and philosophical traditions and is not essentially a purely humanistic experiment. What it is really trying to solve is an increasingly realistic engineering problem:
When the AI is too powerful to write rules in advance for every situation, can the model form a stable "character" on its own, and still choose behaviors that will not harm humans when no one clearly tells it what to do?
And this also brings about a more difficult question - if AI really forms its own "personality" and values, can humans still use it completely as a tool?
Comments