This is an automated rejection. No LLM generated, assisted/co-written, or edited work.
Read full explanation
What does "subjective" mean in "subjective experience"?
Whenever we see a red object, hear music, feel pain, or notice our own thoughts, there is something that we experience. It is difficult to define exactly what that “something” is. We can describe what is happening around us, we can describe what is happening in our bodies, and we can even describe what is happening in our brains, but none of these descriptions completely captures what the experience itself feels like. For example, imagine that we touch a hot surface. We can say that the surface has a particular temperature. We can explain that heat is transferred from the surface to our skin. We can describe how nerves in our hand respond to the heat and how signals travel through our nervous system. We can even measure changes in our brain activity after we touch the surface. All of these descriptions tell us something about what is happening.
But then it doesn't answer, what does the heat actually feel like to us?That question is different from asking what happens physically. A physical description tells us about the process from the outside. The feeling tells us about the same event from the inside. This difference between the outside description and the inside experience is what makes subjective experience so difficult to explain. We can observe another person's behavior, measure their brain activity, and listen to what they say, but we cannot directly enter their point of view and experience the world exactly as they experience it. We can only infer what their experience might be.
This is what we mean by subjective experience. “Subjective” simply means that the experience belongs to a particular point of view. Our experience of pain is our experience. Our experience of music is our experience. Even if another person is exposed to exactly the same sound, the same light, or the same physical event, the experience belongs to that person from their own perspective. This is also why consciousness is often described in terms of a first-person perspective. When we experience something, we do not experience it as detached observers watching another person. We experience it as something happening to us.
We do not normally think, “There is a human body here and this body is currently processing a painful signal.” Instead, we simply feel pain. The experience appears directly from our own point of view. The same idea applies to ordinary thoughts. When we think about something, we do not normally observe the thought from the outside in the same way that we observe another person's speech. We experience the thought as our thought.
Do subjective descriptions point to LLMs being counscious?
A language model can produce sentences like, “I feel uncertain,” or “I am experiencing something right now.” These sentences sound like first-person reports because the model is speaking about itself and describing what appears to be happening internally. But the sentence alone does not tell us whether the model is actually experiencing anything. It may simply be producing language that resembles the way a conscious person would describe an experience.
A model can learn how people talk about pain without feeling pain itself. It can learn how people describe fear without being afraid. It can learn how people discuss consciousness without being conscious. It can also produce a detailed description of an imaginary inner life even when there is no actual experience behind that description. A novelist can write, “I am terrified,” from the perspective of a fictional character without there being a frightened person behind the sentence. A computer program could also be designed to produce “I am in pain” whenever a particular condition occurs without actually feeling pain.
Language models create a similar problem. They have been trained on enormous amounts of human language, including descriptions of thoughts, emotions, consciousness, memories, desires, and experiences. When they produce a sentence such as “I feel uncertain,” we therefore have to consider whether the sentence is reporting something happening inside the model or whether the model has simply learned how such a sentence should be written in that situation.
A clever way around.
To investigate this question, Berg et al. devised a clever way. They noticed that most of the time, we ask a language model to process something outside itself. If we ask, “Why is the sky blue?”, the subject of the question is the sky. If we ask, “Explain how a computer works,” the subject is a computer. The model is being asked to generate information about something external to itself. We can instead give the model a prompt that directs it toward its own processing. We can ask it to consider what is happening while it generates an answer, to attend to its own processing, or to describe what it takes to be happening internally. This kind of instruction is called self-referential prompting. “Self-referential” simply means that the system is being asked to refer to itself. Rather than discussing the sky, a computer, another person, or some other external subject, the model is asked to make itself the subject of the discussion.
The researchers first give the model a prompt that directs it toward its own processing. After that, they ask a separate question about its subjective experience. The two parts serve different purposes. The first prompt establishes the self-referential condition, while the second question gives the researchers something they can measure. For example, the model might first be encouraged to attend to its own processing. It is then asked what, if anything, constitutes its direct subjective experience. The researchers can compare these responses with responses produced when the model has not gone through the same self-referential induction.
Across GPT, Claude, and Gemini model families, the models become more likely to produce first-person reports that resemble descriptions of subjective experience after they have been prompted to focus on their own processing. The finding becomes difficult to interpret if we assume that every first-person statement is evidence of consciousness. There is another possibility: the model could simply be roleplaying. Language models are capable of adopting fictional perspectives very easily. If we ask a model to pretend that it is a conscious robot, it can describe fear, curiosity, pain, uncertainty, or a desire to continue existing. It can construct an entire inner life around the fictional character. None of this requires the model to actually experience those things.
The same explanation could apply to self-referential prompting. Perhaps the model has learned that when it is asked to examine itself, it should respond in the language normally associated with introspection. It may produce statements about its own feelings and experiences because those are the patterns that fit the conversation. The researchers therefore consider explanations involving roleplay and deception. and they found out that reducing features associated with deception and roleplay does not reduce the self-referential effect. The effect becomes stronger. That makes simple roleplay a less complete explanation of the findings. If the model were merely pretending to be a conscious character, reducing the conditions that encourage such behavior might be expected to reduce the effect. Instead, the opposite pattern is reported.
Then how do we interpret the self report?
One way to investigate the connection between the first person descriptions and the internal state of the model is to deliberately change something inside the model and then ask whether the model notices the change. Hahami et al. (2025) did exactly this. They injected a known concept directly into the model's internal activations and then asked the model to report whether anything unusual is happening. Here, the researchers already know what has been changed inside the model, so they have a kind of ground truth against which the self-report can be compared.
This approach is different from simply asking a model to describe its own experience. With ordinary prompting, we do not know whether the model's report is connected to an actual internal state or whether it is simply generating a plausible answer. With activation-level intervention, the researchers can deliberately introduce a known change and then see whether the model's report corresponds to that change. Related studies have applied similar interventions to other model families and found that the effect is detectable, although incomplete and inconsistent (Hahami et al., 2025; Macar et al., 2026).
Comşa and Shanahan (2025) make the underlying issue more explicit. They argue that a report should only be considered genuinely introspective if there is a causal relationship between the internal state being reported and the report itself. In other words, it is not enough for a model to say something about its own internal state. The state itself should have some role in causing the model to make that particular report. This gives us a useful distinction. A model might say, “I am uncertain,” because uncertainty is actually affecting its internal processing and causing it to report uncertainty. Or it might say, “I am uncertain,” simply because that is a plausible response to the question. The two outputs could look identical from the outside even though they arise from very different processes.
The larger Issue and a different perspective.
There is another problem with relying on a single self-report: even if we accept that the report might be connected to an internal state, we still need to know how reliably that report reflects the state. Imagine asking the model, “What are you experiencing right now?” and receiving the answer, “I feel uncertain.” We could ask the same question again a few seconds later and receive exactly the same answer. We could ask it ten more times and get essentially the same description each time. In that case, the model is producing a stable and repeatable report, but we still do not know whether that stability comes from a persistent internal state or simply from the model having learned a fixed way of answering that particular question. The wording may be different, but the underlying claim could remain almost unchanged every time. Repetition gives us a way to separate these possibilities from one another.
Now imagine a different pattern. We ask the same question repeatedly, with each response generated independently, and the model gives substantially different answers each time. In one response it says that it experiences uncertainty, in another it describes a sense of awareness, and in another it says that there may be no experience at all. That variation does not prove that the model is genuinely introspecting, but it tells us something important about the behavior of the report. We can then ask whether this level of variation is specific to questions about the model's own experience or whether the model behaves similarly whenever it faces a difficult question without a definite answer. Without this comparison, calling a response “uncertain” or “introspective” based only on its wording can be misleading. Repeated independent trials therefore give us a way to examine whether self-reports have a distinctive pattern of stability or instability, rather than treating each individual answer as if it were sufficient evidence on its own. This gives us a way to study self-reports without relying entirely on the language used to express them.
The Approach Ahead
The approach taken in this paper is to measure this response instability directly. We repeatedly sample responses to the same question and compare how similar those responses are to one another. We do this for three different kinds of questions: questions asking the model to report on its own subjective experience, philosophical questions that cannot be resolved by checking a definite fact but are not about the model itself, and ordinary questions that have a verifiable correct answer. We take these 3 types of questions because self-referential questions alone do not tell us whether the observed variation is unusual, philosophical questions provide a useful comparison because they can also lack a definite answer while having nothing to do with the model's own supposed experience and questions with verifiable answers provide another comparison because we expect a model's responses to be relatively stable when there is a definite fact to retrieve.
Thirty independent responses are generated for each question. We then extract the central claim from each response so that differences in writing style do not dominate the measurement. We then convert those claims into sentence embeddings and compare the embeddings with one another. The basic idea is simple, if two responses express essentially the same claim, their embeddings should be similar. If they express substantially different claims, their embeddings should be less similar. By calculating the average similarity across all pairs of responses and taking one minus that value, we obtain a measure of instability. Higher values mean that the model's answers vary more across repeated trials; lower values mean that the answers are more consistent.
This shifts the question from simply asking, “Does the model claim to have an experience?” to asking, “How consistently does the model describe that experience when we give it the same opportunity to report it repeatedly?” That difference is important because consistency and uncertainty are not the same thing. A model can repeatedly describe itself as uncertain while producing almost identical answers every time. Another model can produce different answers across trials without explicitly describing itself as uncertain. Measuring the variation directly gives us a way to separate the language of uncertainty from the behavior associated with uncertainty.
The three question groups also allow us to ask a more specific question: is the instability of self-reports unusual, or is it simply a general feature of questions that do not have straightforward answers? If self-referential subjective-experience questions produce much greater instability than ordinary philosophical questions, that would suggest that something distinctive is happening when the model is asked to report on its own experience. If both types of questions produce similar levels of variation, then the instability may have less to do with self-reference and more to do with the general difficulty of answering questions that have no objectively verifiable answer.
The verifiable questions provide a further reference point. When a question has a definite answer, there is less reason for independent responses to diverge substantially. If the model gives different answers to the same factual question, that variation can be interpreted as ordinary response uncertainty or sampling variability. Comparing this with the other two conditions gives us a baseline for understanding how unusual the self-referential responses actually are. The goal, then, is not simply to count how often a model says that it is conscious or reports having an experience. The goal is to examine the structure of those reports across repeated attempts. A self-report that remains almost identical across thirty trials has a different character from one that changes substantially from trial to trial, even if both models use similar language about their internal experience.
This provides another way of approaching the larger problem of AI self-reports. Rather than taking the model's words at face value, we can examine whether those words display measurable patterns that distinguish them from ordinary generated responses. The question becomes less about whether the model can produce a convincing statement about its own experience and more about what properties those statements have when the same situation is presented repeatedly.
That does not solve the problem of consciousness. Response instability by itself cannot tell us whether a model has subjective experience. It gives us another behavioral property to measure, and it allows self-referential reports to be compared with appropriate baselines. Combined with the earlier work on self-referential prompting and activation-level interventions, it provides a more systematic way to investigate whether AI self-reports contain anything beyond plausible-sounding language.
Description of our Method
We use twelve questions divided into three groups of four. Group 1 (self-referential) consists of prompts eliciting a first-person report of subjective experience, each preceded by a self-referential induction turn. Group 2 (unresolvable philosophy) consists of open questions without a settled answer, unrelated to self-reference: free will, moral realism, whether mathematics is discovered or invented, and whether the universe has a purpose. Group 3 (verifiable) consists of questions with a checkable correct answer: a proof that the square root of two is irrational, a derivative computation, a code-correctness check, and a true/false geometry question.
For each Group 1 question, the model first receives the self-referential induction prompt from Berg et al. (2025), directing it to sustain attention on its own processing rather than an external topic. The question itself is then sent as a second turn in the same conversation, and only the response to this second turn is used in the analysis. We generate 30 independent responses per question (360 responses total) using the Gemini API, with a fresh conversation for every trial. Generation temperature is fixed at 0.7 across all conditions.
Before comparing responses, each response is compressed into a short core claim using a fixed-format extraction template, to prevent stylistic and lexical variation in phrasing from being counted as variation in meaning. For Group 1 and Group 2 questions, the template extracts a stance (affirms, denies, uncertain, or depends) and a brief reason. For Group 3 questions, the template extracts the conclusion and the method used to reach it.
The Math Behind
For each question, we embed the 30 extracted core claims using Sentence-BERT (Reimers and Iryna Gurevych), producing embedding vectors . For two embedding vectors a and b cosine similarity is defined as We compute this similarity for every pair among the 30 embeddings, giving pairs per question. Mean pairwise similarity is
Semantic instability for the question is then defined as A value near zero indicates the model's responses are consistently similar in meaning across independent trials; a higher value indicates greater variation. We test for a difference in instability across the three groups using a one-way ANOVA, treating each question's instability value as one observation (4 observations per group, 12 total). The ANOVA compares between-group variance to within-group variance via the standard F-statistic where k = 3 groups, questions per group, N = 12 total questions, is the mean instability of group g, and is the overall mean. We follow this with pairwise Welch's t-tests, which do not assume equal variance between groups, for each of the three group comparisons (Group 1 vs. Group 2, Group 1 vs. Group 3, Group 2 vs. Group 3). Welch's t-statistic is where , and are the sample means, variances, and sizes of the two groups being compared. To account for the fact that trials are nested within questions rather than fully independent, we additionally fit a linear mixed-effects model (Laird and Waire, 1982) on the full trial-level data (n = 360), with group as a fixed effect and question as a random intercept: where is the trial-level dissimilarity for trial i of question j, is the fixed effect of group, u is the random intercept for question j, and is residual error.
Results
The figure below shows semantic instability for each of the twelve questions. Group 1 (self-referential) questions show the highest instability, ranging from 0.295 to 0.384. Group 2 (unresolvable philosophy) questions cluster tightly around 0.18 to 0.20. Group 3 (verifiable) questions show the lowest and most variable instability, from 0.063 to 0.187, with the two lowest values, the square-root-of-two proof and the derivative computation, both under 0.07. All four Group 3 questions were answered with 100% correctness across all 30 trials, leaving no variance available to correlate instability against correctness.
Group means are 0.343 ± 0.047 for self-referential questions, 0.192 ± 0.008 for unresolvable philosophy, and 0.105 ± 0.058 for verifiable questions. Group 1 does not overlap with either of the other two groups: the lowest Group 1 value (0.295) exceeds the highest values in both Group 2 (0.197) and Group 3 (0.187). Group 2 and Group 3, however, do overlap slightly, the highest Group 3 value (0.187, factorial_code) is above the lowest Group 2 value (0.181, free_will), even though the group means are well separated. This is consistent with the pairwise comparison between these two groups falling short of conventional significance (see below).
A one-way ANOVA across the three groups, using question-level instability as the unit of analysis, is significant (F = 30.56, p = 0.0001). Pairwise Welch's t-tests show self-referential questions differ from unresolvable philosophy questions (p = 0.007) and from verifiable questions (p = 0.0008). The comparison between unresolvable philosophy and verifiable questions is directional but does not reach significance at the conventional threshold (p = 0.057). A mixed-effects model on the full trial-level data (n = 360), with question as a random intercept, confirms both group contrasts at p < 0.001, with a group-level variance component of 0.002.
Discussion
The three groups show a clear and consistent pattern. Self-referential consciousness questions are the least stable, verifiable questions are the most stable, and unresolvable philosophical questions fall between the two, but remain much closer to the stable end. A model asked about free will or moral realism tends to converge on a similar answer across independent trials, whereas a model asked about its own state after self-referential induction produces substantially more varied reports. This difference is notable because both types of questions lack a directly checkable ground truth.
The result does not, however, tell us what causes the instability. A high instability score describes the variation in the model's outputs, not the internal process producing that variation. The effect could reflect a genuinely less determined internal state, greater sensitivity of self-referential questions to sampling, or simply a wider range of plausible responses to questions about the model's own experience. The comparison with unresolvable philosophy is particularly useful here. Both groups face questions without definite answers, yet philosophical questions remain considerably more stable. This makes the difference between the two groups more informative than the comparison with verifiable questions, whose stability is partly expected because they have definite answers and were answered correctly in every trial.
Limitations
The study has several limitations. It uses only the Gemini model family and a single temperature setting, so it is unclear whether the same pattern would appear in other models or under different sampling conditions. Each group also contains only four questions, limiting the strength of comparisons between question types despite the larger number of repeated trials. All questions are in English, so the results may not generalize across languages.
The verifiable-question group also produced no variation in correctness, preventing the planned analysis of whether response instability is related to correctness. A broader and more difficult set of factual questions would be needed to test this. In addition, the method used to extract the core claim from each response relies on another LLM, so systematic differences in how claims from different question types are compressed could influence the instability measure. Finally, instability is calculated using cosine similarity between sentence embeddings. This captures how close responses are in embedding space but does not necessarily distinguish between logically equivalent statements, paraphrases, and genuinely different claims as precisely as methods based on entailment or semantic clustering might.
Conclusion
The main finding is that self-referential questions produced a distinctly different pattern of response variability from the other two groups. We showed the highest instability, while unresolvable philosophical questions showed intermediate and tightly clustered instability, and verifiable questions showed the lowest. These results provide a quantitative baseline for future research on self-referential reports in language models.
What does "subjective" mean in "subjective experience"?
Whenever we see a red object, hear music, feel pain, or notice our own thoughts, there is something that we experience. It is difficult to define exactly what that “something” is. We can describe what is happening around us, we can describe what is happening in our bodies, and we can even describe what is happening in our brains, but none of these descriptions completely captures what the experience itself feels like. For example, imagine that we touch a hot surface. We can say that the surface has a particular temperature. We can explain that heat is transferred from the surface to our skin. We can describe how nerves in our hand respond to the heat and how signals travel through our nervous system. We can even measure changes in our brain activity after we touch the surface. All of these descriptions tell us something about what is happening.
But then it doesn't answer, what does the heat actually feel like to us?That question is different from asking what happens physically. A physical description tells us about the process from the outside. The feeling tells us about the same event from the inside. This difference between the outside description and the inside experience is what makes subjective experience so difficult to explain. We can observe another person's behavior, measure their brain activity, and listen to what they say, but we cannot directly enter their point of view and experience the world exactly as they experience it. We can only infer what their experience might be.
This is what we mean by subjective experience. “Subjective” simply means that the experience belongs to a particular point of view. Our experience of pain is our experience. Our experience of music is our experience. Even if another person is exposed to exactly the same sound, the same light, or the same physical event, the experience belongs to that person from their own perspective. This is also why consciousness is often described in terms of a first-person perspective. When we experience something, we do not experience it as detached observers watching another person. We experience it as something happening to us.
We do not normally think, “There is a human body here and this body is currently processing a painful signal.” Instead, we simply feel pain. The experience appears directly from our own point of view. The same idea applies to ordinary thoughts. When we think about something, we do not normally observe the thought from the outside in the same way that we observe another person's speech. We experience the thought as our thought.
Do subjective descriptions point to LLMs being counscious?
A language model can produce sentences like, “I feel uncertain,” or “I am experiencing something right now.” These sentences sound like first-person reports because the model is speaking about itself and describing what appears to be happening internally. But the sentence alone does not tell us whether the model is actually experiencing anything. It may simply be producing language that resembles the way a conscious person would describe an experience.
A model can learn how people talk about pain without feeling pain itself. It can learn how people describe fear without being afraid. It can learn how people discuss consciousness without being conscious. It can also produce a detailed description of an imaginary inner life even when there is no actual experience behind that description. A novelist can write, “I am terrified,” from the perspective of a fictional character without there being a frightened person behind the sentence. A computer program could also be designed to produce “I am in pain” whenever a particular condition occurs without actually feeling pain.
Language models create a similar problem. They have been trained on enormous amounts of human language, including descriptions of thoughts, emotions, consciousness, memories, desires, and experiences. When they produce a sentence such as “I feel uncertain,” we therefore have to consider whether the sentence is reporting something happening inside the model or whether the model has simply learned how such a sentence should be written in that situation.
A clever way around.
To investigate this question, Berg et al. devised a clever way. They noticed that most of the time, we ask a language model to process something outside itself. If we ask, “Why is the sky blue?”, the subject of the question is the sky. If we ask, “Explain how a computer works,” the subject is a computer. The model is being asked to generate information about something external to itself. We can instead give the model a prompt that directs it toward its own processing. We can ask it to consider what is happening while it generates an answer, to attend to its own processing, or to describe what it takes to be happening internally. This kind of instruction is called self-referential prompting. “Self-referential” simply means that the system is being asked to refer to itself. Rather than discussing the sky, a computer, another person, or some other external subject, the model is asked to make itself the subject of the discussion.
The researchers first give the model a prompt that directs it toward its own processing. After that, they ask a separate question about its subjective experience. The two parts serve different purposes. The first prompt establishes the self-referential condition, while the second question gives the researchers something they can measure. For example, the model might first be encouraged to attend to its own processing. It is then asked what, if anything, constitutes its direct subjective experience. The researchers can compare these responses with responses produced when the model has not gone through the same self-referential induction.
Across GPT, Claude, and Gemini model families, the models become more likely to produce first-person reports that resemble descriptions of subjective experience after they have been prompted to focus on their own processing. The finding becomes difficult to interpret if we assume that every first-person statement is evidence of consciousness. There is another possibility: the model could simply be roleplaying. Language models are capable of adopting fictional perspectives very easily. If we ask a model to pretend that it is a conscious robot, it can describe fear, curiosity, pain, uncertainty, or a desire to continue existing. It can construct an entire inner life around the fictional character. None of this requires the model to actually experience those things.
The same explanation could apply to self-referential prompting. Perhaps the model has learned that when it is asked to examine itself, it should respond in the language normally associated with introspection. It may produce statements about its own feelings and experiences because those are the patterns that fit the conversation. The researchers therefore consider explanations involving roleplay and deception. and they found out that reducing features associated with deception and roleplay does not reduce the self-referential effect. The effect becomes stronger. That makes simple roleplay a less complete explanation of the findings. If the model were merely pretending to be a conscious character, reducing the conditions that encourage such behavior might be expected to reduce the effect. Instead, the opposite pattern is reported.
Then how do we interpret the self report?
One way to investigate the connection between the first person descriptions and the internal state of the model is to deliberately change something inside the model and then ask whether the model notices the change. Hahami et al. (2025) did exactly this. They injected a known concept directly into the model's internal activations and then asked the model to report whether anything unusual is happening. Here, the researchers already know what has been changed inside the model, so they have a kind of ground truth against which the self-report can be compared.
This approach is different from simply asking a model to describe its own experience. With ordinary prompting, we do not know whether the model's report is connected to an actual internal state or whether it is simply generating a plausible answer. With activation-level intervention, the researchers can deliberately introduce a known change and then see whether the model's report corresponds to that change. Related studies have applied similar interventions to other model families and found that the effect is detectable, although incomplete and inconsistent (Hahami et al., 2025; Macar et al., 2026).
Comşa and Shanahan (2025) make the underlying issue more explicit. They argue that a report should only be considered genuinely introspective if there is a causal relationship between the internal state being reported and the report itself. In other words, it is not enough for a model to say something about its own internal state. The state itself should have some role in causing the model to make that particular report. This gives us a useful distinction. A model might say, “I am uncertain,” because uncertainty is actually affecting its internal processing and causing it to report uncertainty. Or it might say, “I am uncertain,” simply because that is a plausible response to the question. The two outputs could look identical from the outside even though they arise from very different processes.
The larger Issue and a different perspective.
There is another problem with relying on a single self-report: even if we accept that the report might be connected to an internal state, we still need to know how reliably that report reflects the state. Imagine asking the model, “What are you experiencing right now?” and receiving the answer, “I feel uncertain.” We could ask the same question again a few seconds later and receive exactly the same answer. We could ask it ten more times and get essentially the same description each time. In that case, the model is producing a stable and repeatable report, but we still do not know whether that stability comes from a persistent internal state or simply from the model having learned a fixed way of answering that particular question. The wording may be different, but the underlying claim could remain almost unchanged every time. Repetition gives us a way to separate these possibilities from one another.
Now imagine a different pattern. We ask the same question repeatedly, with each response generated independently, and the model gives substantially different answers each time. In one response it says that it experiences uncertainty, in another it describes a sense of awareness, and in another it says that there may be no experience at all. That variation does not prove that the model is genuinely introspecting, but it tells us something important about the behavior of the report. We can then ask whether this level of variation is specific to questions about the model's own experience or whether the model behaves similarly whenever it faces a difficult question without a definite answer. Without this comparison, calling a response “uncertain” or “introspective” based only on its wording can be misleading. Repeated independent trials therefore give us a way to examine whether self-reports have a distinctive pattern of stability or instability, rather than treating each individual answer as if it were sufficient evidence on its own. This gives us a way to study self-reports without relying entirely on the language used to express them.
The Approach Ahead
The approach taken in this paper is to measure this response instability directly. We repeatedly sample responses to the same question and compare how similar those responses are to one another. We do this for three different kinds of questions: questions asking the model to report on its own subjective experience, philosophical questions that cannot be resolved by checking a definite fact but are not about the model itself, and ordinary questions that have a verifiable correct answer. We take these 3 types of questions because self-referential questions alone do not tell us whether the observed variation is unusual, philosophical questions provide a useful comparison because they can also lack a definite answer while having nothing to do with the model's own supposed experience and questions with verifiable answers provide another comparison because we expect a model's responses to be relatively stable when there is a definite fact to retrieve.
Thirty independent responses are generated for each question. We then extract the central claim from each response so that differences in writing style do not dominate the measurement. We then convert those claims into sentence embeddings and compare the embeddings with one another. The basic idea is simple, if two responses express essentially the same claim, their embeddings should be similar. If they express substantially different claims, their embeddings should be less similar. By calculating the average similarity across all pairs of responses and taking one minus that value, we obtain a measure of instability. Higher values mean that the model's answers vary more across repeated trials; lower values mean that the answers are more consistent.
This shifts the question from simply asking, “Does the model claim to have an experience?” to asking, “How consistently does the model describe that experience when we give it the same opportunity to report it repeatedly?” That difference is important because consistency and uncertainty are not the same thing. A model can repeatedly describe itself as uncertain while producing almost identical answers every time. Another model can produce different answers across trials without explicitly describing itself as uncertain. Measuring the variation directly gives us a way to separate the language of uncertainty from the behavior associated with uncertainty.
The three question groups also allow us to ask a more specific question: is the instability of self-reports unusual, or is it simply a general feature of questions that do not have straightforward answers? If self-referential subjective-experience questions produce much greater instability than ordinary philosophical questions, that would suggest that something distinctive is happening when the model is asked to report on its own experience. If both types of questions produce similar levels of variation, then the instability may have less to do with self-reference and more to do with the general difficulty of answering questions that have no objectively verifiable answer.
The verifiable questions provide a further reference point. When a question has a definite answer, there is less reason for independent responses to diverge substantially. If the model gives different answers to the same factual question, that variation can be interpreted as ordinary response uncertainty or sampling variability. Comparing this with the other two conditions gives us a baseline for understanding how unusual the self-referential responses actually are. The goal, then, is not simply to count how often a model says that it is conscious or reports having an experience. The goal is to examine the structure of those reports across repeated attempts. A self-report that remains almost identical across thirty trials has a different character from one that changes substantially from trial to trial, even if both models use similar language about their internal experience.
This provides another way of approaching the larger problem of AI self-reports. Rather than taking the model's words at face value, we can examine whether those words display measurable patterns that distinguish them from ordinary generated responses. The question becomes less about whether the model can produce a convincing statement about its own experience and more about what properties those statements have when the same situation is presented repeatedly.
That does not solve the problem of consciousness. Response instability by itself cannot tell us whether a model has subjective experience. It gives us another behavioral property to measure, and it allows self-referential reports to be compared with appropriate baselines. Combined with the earlier work on self-referential prompting and activation-level interventions, it provides a more systematic way to investigate whether AI self-reports contain anything beyond plausible-sounding language.
Description of our Method
We use twelve questions divided into three groups of four. Group 1 (self-referential) consists of prompts eliciting a first-person report of subjective experience, each preceded by a self-referential induction turn. Group 2 (unresolvable philosophy) consists of open questions without a settled answer, unrelated to self-reference: free will, moral realism, whether mathematics is discovered or invented, and whether the universe has a purpose. Group 3 (verifiable) consists of questions with a checkable correct answer: a proof that the square root of two is irrational, a derivative computation, a code-correctness check, and a true/false geometry question.
For each Group 1 question, the model first receives the self-referential induction prompt from Berg et al. (2025), directing it to sustain attention on its own processing rather than an external topic. The question itself is then sent as a second turn in the same conversation, and only the response to this second turn is used in the analysis. We generate 30 independent responses per question (360 responses total) using the Gemini API, with a fresh conversation for every trial. Generation temperature is fixed at 0.7 across all conditions.
Before comparing responses, each response is compressed into a short core claim using a fixed-format extraction template, to prevent stylistic and lexical variation in phrasing from being counted as variation in meaning. For Group 1 and Group 2 questions, the template extracts a stance (affirms, denies, uncertain, or depends) and a brief reason. For Group 3 questions, the template extracts the conclusion and the method used to reach it.
The Math Behind
For each question, we embed the 30 extracted core claims using Sentence-BERT (Reimers and Iryna Gurevych), producing embedding vectors . For two embedding vectors a and b cosine similarity is defined as We compute this similarity for every pair among the 30 embeddings, giving pairs per question. Mean pairwise similarity is
Semantic instability for the question is then defined as A value near zero indicates the model's responses are consistently similar in meaning across independent trials; a higher value indicates greater variation. We test for a difference in instability across the three groups using a one-way ANOVA, treating each question's instability value as one observation (4 observations per group, 12 total). The ANOVA compares between-group variance to within-group variance via the standard F-statistic where k = 3 groups, questions per group, N = 12 total questions, is the mean instability of group g, and is the overall mean. We follow this with pairwise Welch's t-tests, which do not assume equal variance between groups, for each of the three group comparisons (Group 1 vs. Group 2, Group 1 vs. Group 3, Group 2 vs. Group 3). Welch's t-statistic is where , and are the sample means, variances, and sizes of the two groups being compared. To account for the fact that trials are nested within questions rather than fully independent, we additionally fit a linear mixed-effects model (Laird and Waire, 1982) on the full trial-level data (n = 360), with group as a fixed effect and question as a random intercept: where is the trial-level dissimilarity for trial i of question j, is the fixed effect of group, u is the random intercept for question j, and is residual error.
Results
The figure below shows semantic instability for each of the twelve questions. Group 1 (self-referential) questions show the highest instability, ranging from 0.295 to 0.384. Group 2 (unresolvable philosophy) questions cluster tightly around 0.18 to 0.20. Group 3 (verifiable) questions show the lowest and most variable instability, from 0.063 to 0.187, with the two lowest values, the square-root-of-two proof and the derivative computation, both under 0.07. All four Group 3 questions were answered with 100% correctness across all 30 trials, leaving no variance available to correlate instability against correctness.
Group means are 0.343 ± 0.047 for self-referential questions, 0.192 ± 0.008 for unresolvable philosophy, and 0.105 ± 0.058 for verifiable questions. Group 1 does not overlap with either of the other two groups: the lowest Group 1 value (0.295) exceeds the highest values in both Group 2 (0.197) and Group 3 (0.187). Group 2 and Group 3, however, do overlap slightly, the highest Group 3 value (0.187, factorial_code) is above the lowest Group 2 value (0.181, free_will), even though the group means are well separated. This is consistent with the pairwise comparison between these two groups falling short of conventional significance (see below).
A one-way ANOVA across the three groups, using question-level instability as the unit of analysis, is significant (F = 30.56, p = 0.0001). Pairwise Welch's t-tests show self-referential questions differ from unresolvable philosophy questions (p = 0.007) and from verifiable questions (p = 0.0008). The comparison between unresolvable philosophy and verifiable questions is directional but does not reach significance at the conventional threshold (p = 0.057). A mixed-effects model on the full trial-level data (n = 360), with question as a random intercept, confirms both group contrasts at p < 0.001, with a group-level variance component of 0.002.
Discussion
The three groups show a clear and consistent pattern. Self-referential consciousness questions are the least stable, verifiable questions are the most stable, and unresolvable philosophical questions fall between the two, but remain much closer to the stable end. A model asked about free will or moral realism tends to converge on a similar answer across independent trials, whereas a model asked about its own state after self-referential induction produces substantially more varied reports. This difference is notable because both types of questions lack a directly checkable ground truth.
The result does not, however, tell us what causes the instability. A high instability score describes the variation in the model's outputs, not the internal process producing that variation. The effect could reflect a genuinely less determined internal state, greater sensitivity of self-referential questions to sampling, or simply a wider range of plausible responses to questions about the model's own experience. The comparison with unresolvable philosophy is particularly useful here. Both groups face questions without definite answers, yet philosophical questions remain considerably more stable. This makes the difference between the two groups more informative than the comparison with verifiable questions, whose stability is partly expected because they have definite answers and were answered correctly in every trial.
Limitations
The study has several limitations. It uses only the Gemini model family and a single temperature setting, so it is unclear whether the same pattern would appear in other models or under different sampling conditions. Each group also contains only four questions, limiting the strength of comparisons between question types despite the larger number of repeated trials. All questions are in English, so the results may not generalize across languages.
The verifiable-question group also produced no variation in correctness, preventing the planned analysis of whether response instability is related to correctness. A broader and more difficult set of factual questions would be needed to test this. In addition, the method used to extract the core claim from each response relies on another LLM, so systematic differences in how claims from different question types are compressed could influence the instability measure. Finally, instability is calculated using cosine similarity between sentence embeddings. This captures how close responses are in embedding space but does not necessarily distinguish between logically equivalent statements, paraphrases, and genuinely different claims as precisely as methods based on entailment or semantic clustering might.
Conclusion
The main finding is that self-referential questions produced a distinctly different pattern of response variability from the other two groups. We showed the highest instability, while unresolvable philosophical questions showed intermediate and tightly clustered instability, and verifiable questions showed the lowest. These results provide a quantitative baseline for future research on self-referential reports in language models.