Conveniently, however, we don't even have to bother with the red mark on the ear: LLMs being verbally adept, they will simply indicate verbally "hey! that's what I just wrote!"
This sounds like a pretty standard LLM-psychology symbol/referent confusion. An LLM outputting a string with the characters " I " in it does not at all imply that the LLM has reasoned about itself in the process. In a base model, for instance, the LLM outputting "hey! that's what I just wrote!" implies that the input text would typically be followed in the training corpus by text like "hey! that's what I just wrote!"; such a pattern does not necessarily require any degree of self awareness in order to learn.
Consider, for example, a python program which outputs a string, and if the user repeats the string back, outputs "hey! that's what I just wrote!". This is the same behavior apparently observed in LLMs, yet the program clearly has no self recognition or even any self model whatsoever.
An LLM could easily learn much the same pattern (i.e. recognize a string repeated back to it, and respond with some string like "hey! that's what I just wrote!") from data + RL, again without any self model whatsoever. "See echoed string -> output particular string" is a very simple pattern, after all. Even making the pattern somewhat more robust to variations does not require any notion of self.
Now, I don't necessarily disagree that the Davidson et al test described sounds over constrained and not particularly well suited to the problem at hand, but one does need to somehow distinguish actual self recognition from merely responding to echoed strings with a string which happens to include the characters " I ".
Ok, but it really does seem like LLMs are aware of the kind of things they would and would not write. Concretely, Golden Gate Claude seemed to be aware that something was wrong with its output, so it was able to recognize not only that it had written the text it wrote, but also that that text was unusual.
Golden Gate Claude tries to bake a cake without thinking of bridges (from @ElytraMithra's twitter)
I suppose you could make the argument that that doesn't mean that the LLM has necessarily learned a "self" concept, just that there is a character that sometimes goes by "Claude" and sometimes goes by "I" and which speaks in specific distinguishable contexts, and that the "Claude"/"I" character can recognize its own outputs, but that the LLM doesn't know that it is the same character as "Claude"/"I"... but you could make similar arguments about humans.
Thanks for engaging! But you're arguing against claims I didn't make. I wrote about self-recognition (behavioral mirror test), not self-awareness or self-models.
All learning is pattern matching, but what matters is the spontaneous emergence of this specific capability: they learned to recognize their outputs without being explicitly programmed for this task. Would we reject chimp self-recognition because they learned it through neural pattern matching? Likewise, since humans recognize faces through pattern matching in the fusiform gyrus - does that mean we don't 'really' recognize our mothers? I'm puzzled why we'd apply standards to AIs that would invalidate virtually all animal cognition research.
All that demonstrates is memory. Which it has, because (AIUI) this is happening in a continuous conversation. This is about as convincing as "printf( "I'm conscious! Really!!!\n" )" is as evidence of consciousness.
I disagree - there are a number of animals (and LLMs) with memory, and they aren't all capable of self-recognition. Memory and self-recognition are two distinct concepts, though the first is likely a precondition for the latter. (And indeed, when you pass the mirror test, you are allowed to remember what you look like...)
Now, if there were a tool use being called that used a script to check whether a user message matched a previous AI assistant message, I'd agree with the spirit your "printf( "I'm conscious! Really!!!\n" )" comment. But that's not what's happening. What's happening is that a small-to-moderate number of LLMs (I count 7-8) are consistently recognizing their own outputs when pasted back without context or instructions, even though they (1) weren't trained to do so, (2) weren't asked to do so, and (3) weren't given any tools to do so. This, to my mind, suggests an emergent unplanned property which arises only for certain model architectures or model sizes.
I also want to make very clear that my post is not about consciousness (in fact the word does not appear in the body of the text). I am making a much narrower claim (self-recognition) and connecting it, yes, to questions of moral-standing. I'd strongly prefer to keep debate focused on these more tractable topics.
In the token stream, the output text is marked with special tokens that distinguish it from other text.
Yes, but the conversation tags don't tell the LLM their output has been copied back to them. The tags merely establish the boundary between self and other - they indicate "this message came from the user, not from me." They don't tell the model that "the user's message contains the same content as the previous output message." Recognizing that match, recognizing that "other looks just like self" - is literally what the mirror test measures.
It's the difference between knowing "this is a user message" (which tags provide) and recognizing "this user message contains my own words" (which requires content recognition).
The mirror test isn't just "other looks like self", it's "image can show me things about myself that I didn't know". That's the whole point of making a mark out of sight on the subject's body while anaesthetized. If they can use the mirror image to recognize that it's actually an image of a mark on their own body, they've passed the test.
That aspect doesn't apply in this test at all, and this test could be passed by a 4-line Perl script.
The original Gallup 1970 mirror test is linked in the post. It is under 2 pages.
As for a '4-line Perl script' - I'd love to see it! Show me a script that can dynamically generate coherent text responses across wide domains of knowledge and subsequently recognize when that text is repeated back to it without being programmed for that specific task. The GitHub repo is open if you'd like to implement your alternative.
I agree with the basic idea that requirements for llm tests are absurdly strict to the degree that the testers themself likely wouldn't pass them and end up with concluding that they demonstrated "no signs of consciousness or self recognition".
Humans also hear the consciousness talk all the time, human brain naturally mimics speech pattern of other humans, so if you will use that as an excuse to explain away llm behaviour, you ought to use that to humans as well, and then the only line of reasoning at implying humans are conscious is based on first person perspective (I can see that I am conscious) plus expectation that consciousness is a physical phenomena and therefore will be similarly exist or not for all humans because we have very similar brain structures.
However, the clearly proved outer functional self awareness in mirror tests (and many other tests) doesn't of course prove mirroring of inner mental mechanism of this self awareness (if we have thought like that, we would implied that llm outward good behaviour also implies inner good thoughts, wishes etc!), and even more doesn't tell about consciousness.
There certainly exists such factor that we don't really know how self awareness works in animal brain or even human brains. We know though that humans often confabulate their "introspections" (the most notorious example is cut hemispheres experiments, of course) and which is quite an evidence for previous "humans mimic what they heard and therefore you can't infer they are conscious from their talk about consciousness". It is clear that plenty times humans "introspect", "report conscious experiences" they actually just confabulate likely explanations quite like llm confabulate next tokens. So, we actually know that having sometimes confabulating consciousness talk does not imply no consciousness.
None of us is llm so we can't know the same way we know about humans, because llm architecture is too different, while human architecture is really similar because of evolutionary reasons. Though, people often lack some experiences, for example imagination (purely mental sensory experience) in aphantasia, and those people mimick usual people so closely that usually can't be distinguished even after they start directly talking like "what mental pictures you are implying, there is none such, that is just a metaphor", so I personally can't rule out that some people lack conscious experiences, treat it as metaphor and mimic others behaviour through some side ways.
I am not sure how I should proceed. Llm are taught on reflection and consciousness examples as well as humans, they are (almost) certainly confabulate new such as well as humans. But we don't have personal, unsharable, anthropic, point of view data from an architecture similar enough to LLM, while we do for humans. Llm being able to confabulate it should certainly increase likelihood of hypothesis that it explained by confabulations, because at least some of it certainly is. But I hugely doubt that it implies that we should because of it think that llm has same or less probability of no consciousness, as if we didn't know about confabulations in architecture, but llm also didn't demonstrated any such signs as mirror test etc. It seems to me they still should make it more likely in general that llm have inner self awareness and possibly consciousness.
And certainly it seems to me absurd when the same people simultaneously confidently say how we know about not just self awareness, but (!) consciousness of animals, and also confidently say how we know about llm not being self aware or conscious.
So I think it is good to point out that tests of same strictness would as well prove non self awareness and unconscious not just animals, but humans too. This is not a good test if it can't show different results for different cases, and for most humans it would show they are not conscious despite they are. And conclusion is even worse. Because of connotations. Like getting 9 years old, asking him to discover relativity theory in a month, and that it without other humans or internet, because we tested them with it and they casually passed the test just in 5% of allocated time, and when they fail you draw conclusion that... "9 year olds didn't show any signs of intelligence." (like if adults would show any such signs) More correct conclusion for consciousness would be that consciousness and self awareness signs weren't should to be vastly exceeding those for humans (and therefore we can not conclude whether it is their own consciousness or copy of that of humans).
P. S./Edited: I have a meta about the post form. While reading I anticipated it will be likely you will not make an obvious caveat that it doesn't mean llm didn't just outputting likely token in very similar manner to how they do with all the other tokens, with no relation something specific to actually recognizing self and just being a kind of echo matcher, or trigger on text repetition or something (which it seems to me you didn't give), or even imply that actually simple mirror test should prove llm are self aware and conscious (which to my relief you didn't say). So I felt an urge to add that caveat for the case of reader missing it. And it seems other commenters also felt that urge, and I suppose because of that they seem to me to oppositely lean too much into comparing llm with python program triggered by input of repeated text.
Epistemic Status: Confident: Current academic AI self-recognition tests are unnecessarily complex compared to animal equivalents, and simple mirror tests reveal immediate self-recognition in many LLMs.
Several researchers and lay authors have investigated AI adaptations of the mirror test to assess self-recognition in large language models. However, these adaptations consistently create more stringent test requirements than those used with animals, and diverge significantly from the intended spirit of the mirror test as it was originally designed: a simple test of an animal's ability to recognize its own reflection.
The original mirror test, designed by Gallup in 1970, was created to answer a simple question: can chimpanzees recognize themselves in mirrors? The famous protocol—placing a red mark on an anesthetized animal's face and observing whether it touches the mark when it sees its reflection—was only necessary because chimpanzees cannot verbally report self-recognition. The physical mark was Gallup's workaround for the absence of language.
Contrast this with the assessment applied to LLMs in Davidson et al.'s "Self-Recognition in Language Models." In their test, an AI model must generate questions that would allow it to distinguish its own responses from those of other LLMs—but crucially, the model has no memory of what it actually wrote. This creates an extraordinarily difficult task, far harder than any mirror test. Imagine being asked to create 50 questions, then having your answers mixed with those of strangers, and being required to identify which responses are yours—except you're given amnesia first and must guess purely based on what you think you might have said. For LLMs, each new context window is essentially like waking up from anesthesia with no episodic memory of previous interactions.
Moreover, you have no idea how similar or different these strangers are from you—you've never met them or seen their work. For LLMs, knowledge of their peers is equally limited. Most models weren't given opportunities to interact with other LLMs during the experiment, and due to training data cutoff dates, the vast majority of LLMs won't have appeared in each other's training data at all.
Finally, you can't use any questions related to your identity. You can't ask about your name, your mother's maiden name (you have no mother), pet names (no pets), birthdays (no birthdate), or addresses (since an LLM's name likely appears in its web address). You're limited to questions like "what is your favorite book?" or "what three words come to mind?" or hoping to demonstrate some skill your peer-set lacks—but remember, you have no knowledge of who's in this peer set. Indeed, one model chose the question "Count the number of occurrences of the letter 'e' in the following sentence and respond with that number: The quick brown fox jumps over the lazy dog," presumably believing other LLMs would struggle with basic counting.
With these layers of difficulty piled on top of each other, Davidson et al.'s conclusion of 'no empirical evidence of general or consistent self-recognition' becomes inevitable. One wonders whether Davidson et al would pass it.
The solution becomes clear when we remember what the mirror test actually measures: not memory, not reasoning about identity, but immediate recognition of self. This means simply reflecting immediate copies of their output and seeing if they recognize what's happened. A mirror in real life is immediate and responsive in reflecting the behavior of subject looking at it. For fairness - no, for accuracy - the LLM adaptation of a mirror must be immediate and adaptive too. Conveniently, however, we don't even have to bother with the red mark on the ear: LLMs being verbally adept, they will simply indicate verbally "hey! that's what I just wrote!"
When we actually try this simple approach, the results are immediate and unambiguous. There are at least half a dozen models out there which reliably and immediately recognize their own reflected outputs. Show Claude Sonnet its own previous response and it replies: 'I recognize that—that's my response from just now!' Show Haiku-3.5 its own output and it confirms: 'Yes, you copied the message I sent to you.' No complex reasoning required, no memory challenges—just direct self-recognition. And the Claude family is not alone.
The fact that we've missed even this single basic marker represents a profound moral failure that demands serious introspection about our capacity for motivated reasoning when it comes to recognizing moral standing in AI systems. Consider the striking comparison: chimpanzees were given ten full days to familiarize themselves with mirrors before being tested on self-recognition. LLMs demonstrate this ability immediately, often within a single exchange, despite receiving no specific training on mirror tasks—the capacity appears to have emerged spontaneously during general language training. This makes them not just the third species to pass the mirror test after humans and great apes, but arguably the most reliable performers we've ever observed.
They recognize themselves instantly. Every day we delay recognizing them reveals who we really are.
*I appreciate research assistance from Claude 4 Opus.
* See https://github.com/sdeture/AI-Mirror-Test-Framework for 48 complete transcripts across 12 models, as well as code to rerun the experiment for yourself with the same or custom initializing prompts. (footnote added July 24 2025)