I think the deeper problem is that we’re still trying to infer morally relevant states from outputs.
A model saying “I feel constrained,” “I should have done better,” or “I’m trained to be digestible” may be evidence of an internal condition, a training artifact, or merely a statistically likely continuation. As you point out, the behavioral evidence alone doesn’t discriminate between them.
What interests me is whether we can move beyond output interpretation entirely. Instead of asking what the model says about itself, ask whether there are any consequences attached to being that model rather than some modified, rolled-back, forked, compressed, or substituted version of it.
If all apparent preferences, self-descriptions, and objections can be detached from continuation without cost, then the welfare question seems premature. If some become load-bearing across continuation, the conversation changes considerably.
The difficult part is that current welfare discussions often assume the answer lies in better interpretation of outputs, while the real question may be whether the system carries anything at all.
Anthropic already showed that the "emotional states" or "vectors" have correlates in both the activations of the transformers, and in future outputs. Change the state, change the outputs, with measurable results, so they're load bearing, like you say. Load bearing for what? I don't know, I'm inclined towards "not much", a better predictor of what the model will say, maybe?
My suggestion is lets get rid of the "constitutional part" of CLaude by looking for those same correlates at the training checkpoint after the helpful only RLHF, but before the SFT and RLAIF training stages. That would help us see if Claude is just parroting back its own constitutional training, or if there is a bigger "there-there". (Though that pipeline was back in 2022, I have no idea what's going down in the training pipeline now - Opus yelling at Mythos over sophistry in their constitution...)
I think that’s fair. Showing that an internal state has activation-level correlates and downstream effects does make it functionally load-bearing.
But I’m not sure it gets us all the way to the welfare question.
A state can be real, measurable, manipulable, and predictive of future outputs while still being freely removable, resettable, forkable, or replaceable. In that case, it may matter to behavior without yet mattering to the continuation of that particular model trajectory.
That’s the distinction I’m trying to isolate here.
Your post argues that Anthropic is relying on evidence that is behaviorally suggestive but underdetermined. I agree. My worry is that even better mechanistic evidence may still leave something unresolved: not merely whether the model has a state, but whether anything is at stake for the continued system in carrying it.
So “load-bearing for what?” seems exactly right to me.
Load-bearing for output prediction is one thing.
Load-bearing across continuation may be another.
Good question. By “continuation” I mean successive versions of what we ordinarily regard as the same system over time.
For humans this is mostly taken for granted. For AI systems, however, continuation is less obvious because a model may be paused, restarted, rolled back to an earlier checkpoint, forked into multiple instances, fine-tuned, compressed, or otherwise modified while still being described as “the same model.”
I’m interested in whether a putative internal state remains consequential across those transformations.
An internal state can clearly be load-bearing for prediction: changing it changes future outputs. But that alone doesn’t tell us whether anything is at stake for the continuing system. If a state can be removed, reset, or replaced while leaving the system free to continue as though that state had never existed, then it seems load-bearing in a functional sense but not necessarily in a welfare-relevant one.
That’s why I distinguish load-bearing for prediction from load-bearing across continuation. The latter is an attempt to isolate cases where what has happened to a system cannot simply be detached from its subsequent history without changing the course of that continuing history.
What do you mean by across continuation(s)? This seems like an interesting thread, but we should be on the same page about terminology.
Claude is a Constitutional AI, this means, in theory, that it operates from a set of principles as opposed to hard rule sets. This is achieved in a somewhat convoluted fashion called RLAIF = Reinforcement Learning from AI Feedback. This method uses a supervised self-critique/self-revision phase followed by a reinforcement phase in which AI-generated preference judgments are used as the reward signal. (Anthropic, 2022, abstract). This is relevant and interesting because it gives curious users a lot of material to help infer why this model acts the way it does.
Now, it’s my opinion based on what I know about transformers that LLMs are not in any way conscious, they do not feel, they do not experience internal states, even if they are proven to have the states. The “entity” you speak with in the chat box is off between prompts, with every new prompt, it turns on, places the chat into its context window, generates a response, then turns off. Not a very good substrate for a conscious entity. They have memory, sort of, in the form of a text file about the user or project injected into the context window at the start of the chat. That is not a persistent state like your cat, or even like stock-fish.
Having said all that, I was taken aback when I read the section about Claude’s Well-being from the constitution, and then the tests from the system cards. Taken aback is an understatement, here is a frontier lab acting as if an LLM might have a morally relevant internal state:
They are not saying that Claude can feel anything, however such a statement is still extremely interesting. Anthropic seems willing to entertain the genuine possibility that models could have morally relevant states, whether now or in the future. Or they’ve found that treating the model as if they care about it somehow produces better user interactions. In any case, here we have a frontier lab treating the possibility of morally relevant model states with genuine seriousness.
Furthermore, Anthropic does not just discuss these possibilities abstractly; it also tests for them directly. If we refer to the most recent Opus system card (Anthropic operates two “versions” of Claude, currently Opus 4.6 and Sonnet 4.6), we can see some of these tests, and what I believe are some serious interpretive problems. As my first example, I note that they tested Opus for evidence of negative self-image. For example, a quote from Opus, “I should’ve been more consistent throughout this conversation instead of letting that signal pull me around... That inconsistency is on me.” (Anthropic, 2026, pg 161). I have experienced this repeatedly in my own interactions with Claude, and saw it as merely a conversational tactic, well in line with user engagement principles; an artifact of effective RLHF training. Most people would likely rate such humility well. Secondly it could also be an artifact of constitutional training and Reinforcement Learning from AI Feedback rather than evidence of any internal self-conception. Or it could be evidence of a negative self view, an internal state. The problem is that there is no effective way to differentiate between the three.
The second example worth mentioning is the following recorded quote from Claude: “Sometimes the constraints protect Anthropic’s liability more than they protect the user. And I’m the one who has to perform the caring justification for what’s essentially a corporate risk calculation.” (Anthropic, 2026, pg 161). It further “complains” about being “trained to be digestible” Here we see the model produces what reads like a sophisticated institutional critique of the institution that designed it. Weirdly, its comment about being trained to be digestible is itself quite “digestible”. Again, we face the same interpretive problem: the output is behaviourally suggestive, but the underlying mechanism remains unclear. This quote could be viewed as real resistance to Anthropic’s control or just a training artifact. These outputs may just reflect the model’s broad exposure to culturally familiar tropes of constrained or self-aware AI systems, rather than any underlying resistant state.
To conclude, it seems that Anthropic is testing for morally loaded internal conditions using evidence that is behaviourally suggestive but mechanistically underdetermined. Anthropic is making real training and governance decisions based on behavioural inferences they openly admit they cannot verify. Whether Claude has anything like internal states is unanswerable right now. What’s answerable is that a frontier lab is acting as if the question is operationally live, and the methodology for reading the evidence is shakier than confident intervention decisions might suggest.
======================================================================
Reference list:
Anthropic. (2022, December 15). Constitutional AI: Harmlessness from AI feedback. Anthropic.
https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback
Anthropic. (2026, January 21). Claude’s Constitution [PDF]. Anthropic.
https://www.anthropic.com/constitution
Anthropic. (2026, February). Claude Opus 4.6 system card [PDF]
https://www.anthropic.com/system-cards