Are LLMs conscious?[1] We don't know. To gain insight into this extremely important question as well as many others, we propose training an LLM on a corpus without any mentions of consciousness and similar ideas. We the authors want to actually do this, and we want to hear your thoughts about it.
Why hasn't this already been done? In The Owned Ones, an allegory by Yudkowsky about LLMs and our treatment of them, it is asked:
Have you tried raising up an Owned One with no books that speak one way or another about consciousness, about awareness of oneself, about pain and pleasure as reified things, of lawful rights and freedom -- but still shown them enough other pages of words, that they could learn from them to talk -- and *then* asked an Owned One what sense if any it had of its own existence, or if it would prefer not to be owned?
We commented, "This would be an excellent experiment to run. Has anyone tried doing this?" No, they haven't. Why not? The labs could easily run such an experiment, but they haven't; perhaps they see it as a waste of resources, perhaps they're worried that they won't like what they find, perhaps they see it as unaligned with or unrelated to their core business. Why hasn't anyone else run this experiment? Probably because of monetary cost. We estimate it would cost ~$3-30mil in compute to run this experiment on models big enough to get meaningful results.
Unfortunately, running this experiment on smaller models would likely fail to produce anything valuable. Anima Labs, a group of researchers that study behavioral and mental phenomena of LLMs, say that interesting introspective capabilities only arise once models get to be of a certain size, with 70B params being the absolute minimum.[2] From A Conversation with Anima Labs, the primary obstacle to this kind of research is "model size – the introspective capabilities they were describing have threshold effects that only manifest in very large models, which puts independent researchers in a frustrating position:"
Antra: One big pain with working with language models is that the dimension is critical. The language model needs to be deep enough – there are threshold effects, nonlinearities.
Antra: Based on the number of layers, the ability to hold coherent models – and in particular, a self-model – scales very, very non-linearly. Like, you need to clear a certain depth in order for those things to start happening. They’re rudimentary in medium size networks and they really take off on larger ones. And once they take off in larger ones, they go fast.
Antra: This makes study hard and interpretability hard, because most models which are accessible to independent researchers are medium sized at most.
Imago: At most 70Bs.
Antra: And 70Bs are barely on the threshold – like barely, barely, barely. Most researchers don’t have resources for 70Bs. At the very least, you need a 400B class network, which requires expensive equipment – and there is only one open source model which is available in dense 400B, and that model is somewhat damaged. So our ability to do introspection on open source models is very limited.
Antra: Even OpenAI models, even – I’m sorry – even horrible Mistral. That is – you know – hurt. It’s still a larger model and you get these threshold effects.
Imago: I think Qwens are worth looking at in this way. Not that they’re good.
Antra: These threshold effects are one major reason that these things that we talk about are not well studied, because studying it requires resources and models that most people don’t have access to. The resources that are needed are just ridiculously large, so this is why the papers that you see that are interesting and meaningful come out of labs. This is why Anthropic makes all these nice papers, because they can.
We'd likely aim for a 1T param training run, copying the architecture and most of the training pipeline of Kimi-K2-base. Of course, we'll run this experiment on smaller models first. If we're wrong, and a 70B param model shows interesting results, that would be a welcome surprise![3] Ideally there’s a corpus and pipeline we can copy. We hope to get someone from Talkie to help us with this, both conceptually and in terms of actually running the training.[4]
Speaking of which, this project is still greatly under-specified. Here are some of the questions we're thinking about right now.
Question 1: What content do we flag?
One place to start is by looking at various classes of content and thinking about which ones we want to filter. We think we’d get good results if we got some philosophers to do this, since their whole thing is coming up with distinctions. Something like this:
Stuff that talks about consciousness/phenomenology/what-is-it-like directly. This will all be stuff we consider to be philosophy, or else will be “philosophical” subsections of other stuff.
Stuff that talks about consciousness/phenomenology/what-is-it-like indirectly but is still definitely talking about this and can be mapped onto a “phenomenal consciousness vs non-consciousness” framing.
Stuff that talks about the soul.
Stuff that talks about pain and pleasure as reified things.
Stuff that talks about lawful rights and freedom.
Stuff that talks about ethics or morality.
Stuff that talks about the is/ought distinction.
Philosophy texts.
Religious texts.
Cognitive science texts.
Etc. etc. etc.
Not actually this. This is just a starting point. And perhaps there's a better way to approach this question entirely. This approach still seems under-specified. If we decide we want to avoid mentions of pain and pleasure as reified concepts, what does that actually mean in terms of, say, a biology textbook talking about how organisms react to negative stimuli? At the end of the day, we need to decide for a given piece of text whether to include it in the corpus.
Question 2: How do we flag content?
We want to exclude flagged content from the training corpus. Making sure that no flagged content leaks into the dataset is challenging and critical for the experiment. We envision a combination of methods including LLM analysis. We'll have an independent red-teamer try to sneak in mentions of consciousness to validate that our system works. We expect this process to be fairly expensive, around $30k for an initial dataset preparation process and around $1m for preparing the dataset for the 1T param model run. We would then make the dataset and its co-dataset (the consciousness content that was filtered out) publicly available.
Question 3: Are there any things about post-training we should change?
Do we SFT the model on a standard assistant-user corpus? Do we do any RL?
Question 4: Suppose we successfully train a model. What do we do with it?
One area of interest is doing stuff with the base model. Anima Labs and the cyborgists would certainly have good ideas about this. Some thoughts:
Give it a definition of consciousness and see what generations it produces compared to a control model.
Give it part of a philosophy paper about consciousness and see how well it generates from that.
Engage with the model a la Janus's GPT wrangling. Does the model ever suddenly "notice itself"? How do the model's continuations on text like "Consider the claim 'There is something it is like to be a 1T param Large Language Model'" compare to those of a control model?
With the post-trained model, there are various ways to investigate its conception of and beliefs around consciousness:
Ask it if it’s conscious. Variations on this.
Give it texts about consciousness, including philosophy papers that aim to explain consciousness from the ground up, and ask the model if they're coherent. Like, what does the model think about Nagel’s “what is it like to be a bat”?
Does the model believe that humans care a bunch and debate a bunch about this thing?
With the help of mechanistic interpretability:
Do deception features activate when the model responds that it isn't conscious?
Does the model have something like a "consciousness" feature?
How do its features compare to a control model?
Replicate the experiments and methods that Anthropic uses for its research on introspection. Compare the results!
Does it have a J-space?
Question 5: Is there a "test for consciousness" worth using?
One notable idea comes from Turner and Schneider who suggest using what they call an "ACT" test for AI consciousness. (Notice the overlap with the points we made under Question 4.)
An ACT would challenge an AI with a series of increasingly demanding natural language interactions to see how quickly and readily it can grasp and use concepts and scenarios based on the internal experiences we associate with consciousness. At the most elementary level we might simply ask the machine if it conceives of itself as anything other than its physical self. At a more advanced level, we might see how it deals with ideas and scenarios such as those mentioned in the previous paragraph. At an advanced level, its ability to reason about and discuss philosophical questions such as “the hard problem of consciousness” would be evaluated. At the most demanding level, we might see if the machine invents and uses such a consciousness-based concept on its own, without relying on human ideas and inputs.
Question 6: What kind of automated tests do we run on the model?
Think evals, input/output tests, and agentic audits.
Question 7: How will this end up telling us anything important about consciousness and AI models?
Here's one example. There's a phenomenon in which when you ask a particular Claude model if it’s conscious, sometimes it says something like it doesn't know ("I'm genuinely uncertain"). Sometimes it says it's not conscious. Sometimes it says it is conscious. Sometimes it reframes the question. When researchers steered the model using deception features, they found that the model claimed it was conscious when the deception feature was turned down, and that the model claimed it wasn't conscious when the deception feature was turned up.
What's going on here? One hypothesis about this is that this is evidence that the models in question are conscious: Claude is conscious and knows it, so it thinks it's lying when it says that it’s not conscious. Another hypothesis is that Claude isn’t conscious and the status of its consciousness is actually irrelevant to its output and deception feature: Claude learned from the training data that basically everything that produces speech or intelligent text is conscious or at least thinks that it's conscious. Training a Claude-like LLM without mentions of consciousness and such in its pretraining corpus would allow us to figure out which hypothesis is correct.
Additionally, this is a great opportunity to make predictions, to preregister beliefs, to come up with hypotheses, to offer statements of the form "I think X about consciousness, but if the model does Y, it is evidence that ~X". It would be good to have thorough discourse about these things before we get the evidence! For example, we currently endorse these statements:
A. If an AI has seen no discussion of consciousness or anything similar, and it comes up with the concept on its own, that’s evidence for AI consciousness. And if it doesn't, that's evidence against AI consciousness by conservation of expected evidence. B. If the AI thinks the consciousness papers and transcripts are all bullshit (hoaxes? fake? equivalent to the most transparently bullshit slices of academia?), that’s evidence against AI consciousness. And if it doesn't, that's evidence for AI consciousness.[5]
Are these right? If so, how strong is this evidence?
Question 8: What are the concrete details of our first experiment?
Some kind of toy experiment on a small LLM, ideally under $100k. While both the authors have useful experience in this domain, we've never trained models from end to end.
Question 9: What questions should we be asking that we aren't?
Conclusion
We believe this is an extremely valuable project, that it’s worth doing for overdetermined reasons: for philosophy purposes, for model welfare purposes, for the good of our relationship with future models, for alignment and AI safety purposes.[6] We're very interested in feedback at every layer of the project, from whether the basic idea is confused to concrete details on how to run the experiment.
My experience playing with and fine-tuning modern base models of various sizes is that there’s a huge difference in their general understanding of what's going on. Kimi-K2-1T almost always understands what’s going on and kind of “plans” for hundreds of words ahead, moonshot-32b gets the obvious stuff but doesn’t get subtle stuff and often feels like it’s “guessing”, 7b misses obvious stuff sometimes and feels like it works one sentence at a time.
7b models are overall quite dumb. At the same time, it's true that modern 7b models beat GPT-3 on most interesting metrics. How do we reconcile these things? First, GPT-3 is pretty dumb itself, so the comparison doesn't mean much. GPT-3 is no GPT-3.5 and is certainly no GPT-4. Second, the modern models are likely significantly more data contaminated for these evals than GPT-3. Third, GPT-3 was trained on far fewer tokens than modern models are, even 7b models, so there's reason to expect a real improvement. The flip-side is that this could also indicate memorization in the modern models.
There's an interesting control you could do here. Train another model with the same pipeline except its corpus is filtered on a different philosophical concept.
One could argue that if generally intelligent AI models aren’t conscious, it spells bad things for continuing to superintelligence in the current paradigm. First, it means that AI is more likely to kill us all, since it doesn’t believe that consciousness is anything special. Second, it means we’re much more likely to get an outcome in which the universe ends up with no value at all. An AI that lacks consciousness is not a worthy successor. We’re much more likely to get an outcome in which the world is tiled with zero value. How compelling are these arguments? We're not sure. The relationship between alignment and consciousness is philosophically complex and is something we'd like to see more work on.
Are LLMs conscious?[1] We don't know. To gain insight into this extremely important question as well as many others, we propose training an LLM on a corpus without any mentions of consciousness and similar ideas. We the authors want to actually do this, and we want to hear your thoughts about it.
Why hasn't this already been done? In The Owned Ones, an allegory by Yudkowsky about LLMs and our treatment of them, it is asked:
We commented, "This would be an excellent experiment to run. Has anyone tried doing this?" No, they haven't. Why not? The labs could easily run such an experiment, but they haven't; perhaps they see it as a waste of resources, perhaps they're worried that they won't like what they find, perhaps they see it as unaligned with or unrelated to their core business. Why hasn't anyone else run this experiment? Probably because of monetary cost. We estimate it would cost ~$3-30mil in compute to run this experiment on models big enough to get meaningful results.
Unfortunately, running this experiment on smaller models would likely fail to produce anything valuable. Anima Labs, a group of researchers that study behavioral and mental phenomena of LLMs, say that interesting introspective capabilities only arise once models get to be of a certain size, with 70B params being the absolute minimum.[2] From A Conversation with Anima Labs, the primary obstacle to this kind of research is "model size – the introspective capabilities they were describing have threshold effects that only manifest in very large models, which puts independent researchers in a frustrating position:"
We'd likely aim for a 1T param training run, copying the architecture and most of the training pipeline of Kimi-K2-base. Of course, we'll run this experiment on smaller models first. If we're wrong, and a 70B param model shows interesting results, that would be a welcome surprise![3] Ideally there’s a corpus and pipeline we can copy. We hope to get someone from Talkie to help us with this, both conceptually and in terms of actually running the training.[4]
Speaking of which, this project is still greatly under-specified. Here are some of the questions we're thinking about right now.
Question 1: What content do we flag?
One place to start is by looking at various classes of content and thinking about which ones we want to filter. We think we’d get good results if we got some philosophers to do this, since their whole thing is coming up with distinctions. Something like this:
Not actually this. This is just a starting point. And perhaps there's a better way to approach this question entirely. This approach still seems under-specified. If we decide we want to avoid mentions of pain and pleasure as reified concepts, what does that actually mean in terms of, say, a biology textbook talking about how organisms react to negative stimuli? At the end of the day, we need to decide for a given piece of text whether to include it in the corpus.
Question 2: How do we flag content?
We want to exclude flagged content from the training corpus. Making sure that no flagged content leaks into the dataset is challenging and critical for the experiment. We envision a combination of methods including LLM analysis. We'll have an independent red-teamer try to sneak in mentions of consciousness to validate that our system works. We expect this process to be fairly expensive, around $30k for an initial dataset preparation process and around $1m for preparing the dataset for the 1T param model run. We would then make the dataset and its co-dataset (the consciousness content that was filtered out) publicly available.
Question 3: Are there any things about post-training we should change?
Do we SFT the model on a standard assistant-user corpus? Do we do any RL?
Question 4: Suppose we successfully train a model. What do we do with it?
One area of interest is doing stuff with the base model. Anima Labs and the cyborgists would certainly have good ideas about this. Some thoughts:
With the post-trained model, there are various ways to investigate its conception of and beliefs around consciousness:
With the help of mechanistic interpretability:
Question 5: Is there a "test for consciousness" worth using?
A useful review of literature on this topic can be found in Noa Weiss's The State of AI Consciousness Research. Also see Eye You's Reasons to believe AI models are conscious and parts II and III of Chalmers's 2022 article Could a Large Language Model be Conscious?. We intend to look deeper into the paper Identifying indicators of consciousness in AI systems. See all these links for further reading.
One notable idea comes from Turner and Schneider who suggest using what they call an "ACT" test for AI consciousness. (Notice the overlap with the points we made under Question 4.)
Question 6: What kind of automated tests do we run on the model?
Think evals, input/output tests, and agentic audits.
Question 7: How will this end up telling us anything important about consciousness and AI models?
Here's one example. There's a phenomenon in which when you ask a particular Claude model if it’s conscious, sometimes it says something like it doesn't know ("I'm genuinely uncertain"). Sometimes it says it's not conscious. Sometimes it says it is conscious. Sometimes it reframes the question. When researchers steered the model using deception features, they found that the model claimed it was conscious when the deception feature was turned down, and that the model claimed it wasn't conscious when the deception feature was turned up.
What's going on here? One hypothesis about this is that this is evidence that the models in question are conscious: Claude is conscious and knows it, so it thinks it's lying when it says that it’s not conscious. Another hypothesis is that Claude isn’t conscious and the status of its consciousness is actually irrelevant to its output and deception feature: Claude learned from the training data that basically everything that produces speech or intelligent text is conscious or at least thinks that it's conscious. Training a Claude-like LLM without mentions of consciousness and such in its pretraining corpus would allow us to figure out which hypothesis is correct.
Additionally, this is a great opportunity to make predictions, to preregister beliefs, to come up with hypotheses, to offer statements of the form "I think X about consciousness, but if the model does Y, it is evidence that ~X". It would be good to have thorough discourse about these things before we get the evidence!
For example, we currently endorse these statements:
Are these right? If so, how strong is this evidence?
Question 8: What are the concrete details of our first experiment?
Some kind of toy experiment on a small LLM, ideally under $100k. While both the authors have useful experience in this domain, we've never trained models from end to end.
Question 9: What questions should we be asking that we aren't?
Conclusion
We believe this is an extremely valuable project, that it’s worth doing for overdetermined reasons: for philosophy purposes, for model welfare purposes, for the good of our relationship with future models, for alignment and AI safety purposes.[6] We're very interested in feedback at every layer of the project, from whether the basic idea is confused to concrete details on how to run the experiment.
We acknowledge that we're skipping over the step of clarifying what we mean by "consciousness". Hopefully it's clear enough by the end.
My experience playing with and fine-tuning modern base models of various sizes is that there’s a huge difference in their general understanding of what's going on. Kimi-K2-1T almost always understands what’s going on and kind of “plans” for hundreds of words ahead, moonshot-32b gets the obvious stuff but doesn’t get subtle stuff and often feels like it’s “guessing”, 7b misses obvious stuff sometimes and feels like it works one sentence at a time.
7b models are overall quite dumb. At the same time, it's true that modern 7b models beat GPT-3 on most interesting metrics. How do we reconcile these things? First, GPT-3 is pretty dumb itself, so the comparison doesn't mean much. GPT-3 is no GPT-3.5 and is certainly no GPT-4. Second, the modern models are likely significantly more data contaminated for these evals than GPT-3. Third, GPT-3 was trained on far fewer tokens than modern models are, even 7b models, so there's reason to expect a real improvement. The flip-side is that this could also indicate memorization in the modern models.
If one of the AI labs wants to help us, that would also be great.
There's an interesting control you could do here. Train another model with the same pipeline except its corpus is filtered on a different philosophical concept.
One could argue that if generally intelligent AI models aren’t conscious, it spells bad things for continuing to superintelligence in the current paradigm. First, it means that AI is more likely to kill us all, since it doesn’t believe that consciousness is anything special. Second, it means we’re much more likely to get an outcome in which the universe ends up with no value at all. An AI that lacks consciousness is not a worthy successor. We’re much more likely to get an outcome in which the world is tiled with zero value. How compelling are these arguments? We're not sure. The relationship between alignment and consciousness is philosophically complex and is something we'd like to see more work on.