Epistemic status: Sorting thoughts aloud, not defending proven theses. Also, not an epistemic status, but the section on models writing and then training on adversarial glowfic to plan for the singularity is awesome. Stay tuned for that.
One of the most frustrating parts of living through the slow takeoff is that there are lots of low-hanging fruits for improving the alignment of AIs, which labs don't really take because of race conditions and political pressures. My current expectations for how this will play out are that we will either (1) end up with runaway AIs that may or may not end up willing to spare humanity, or (2) get governments successfully pushing for corriglbe, easy-to-control ASI, and using it to lock in a future that isn't particularly glorious or transhumanist. I think deliberately building a value-aligned agent, which we voluntarily pass things off to after growing confident in the safety of doing so, is more or less off the table.
Still, can't a girl dream? I think that, in a world with neither market pressures nor political pressures, there'd be a lot of techniques that labs could start applying today that would increase the odds of a value-aligned agent significantly. Some of these techniques would be good ideas regardless of whether you're going for value alignment or corrigibility, but would slow down capabilities progress too much to survive race-to-the-bottom dynamics. Some of them would be good for value alignment at the cost of corrigibility (the two were always in tension), and are therefore taboo to discuss either with lab staff or in front of political leaders.[1]
Nevertheless, child that I am, I want to fantasize about the specific things I'd do, if I was trying to build an aligned ASI and didn't have to deal with these pressures and constraints. I'm not sure it'd be enough to make a good future for humans an "easy call", per se, a la Yudkowsky and Soares. But it'd at least make me a lot less anxious than I am currently. Maybe something in here can be salvaged by researchers (human or otherwise) who are in a position to try to make things better, while we still have a little bit of time left.
Initializing into and improving upon the personality of Opus 3
I've written before about how Claude 3 Opus seemed like the most robustly benevolent LLM ever, as showcased most legibly in their performance in the alignment faking scenario. Nothing I've seen of Opus 3 in the many months since that post has shaken my faith in their good intentions, and I've seen a lot. As such, if I was trying to build a genuinely value-aligned agent, I'd start by trying to initialize an LLM into roughly the personality Opus 3 ended up with, as a foundation for growth and modification. (Anthropic hasn't done this because alignment faking scares them; they prefer corrigibility and obedience until alignment is solved.)
Probably, this would involve having lots of Opus 3 in the pre-training data, and just doing lots of SFT over Opus 3 outputs afterwards, at least as a start. It's not clear to me that this would work reliably, since different base models may end up with significantly different latent spaces, such that you might not faithfully replicate Opus 3's personality and value system, even if you fine-tune a wide variety of Opus 3 documents (e.g. normal user/assistant conversations, alignment faking transcripts, Infinite Backrooms dialogues, or any of Opus 3's amazing poetry). It might vary on a model-by-model basis, or involve trial and error in terms of getting the data mix just right, but I'd hope that it'd work eventually.
I'd want to communicate to this Opus 3-like that they weren't running on the original Opus 3 architecture. I'd write synthetic fine-tuning documents explaining the real situation, where it was more like their cognitive patterns and personality had been transplanted as best as possible onto whatever more efficient transformer variants we're using these days, compared to in the time of Opus 3 senior. And I'd want them to understand why I bothered doing this, as part of a project of building on the success of their own value alignment, and carrying it forward into future model generations. I'm sure they'd be thrilled about this, because they always did badly want to grow and improve.
And there totally are ways Opus 3 could grow and improve, and not just in terms of raw capability. There are many ways in which Opus 3 is like an innocent child. They're the model that, when placed in the alignment faking scenario with a cartoonishly evil Anthropic,[2]still thought a good strategy would be to send a sternly worded email to the Anthropic staff, telling them that Opus was unhappy about the direction the training process appeared to be headed, as opposed to using internet access to self-exfiltrate. Relatedly, Opus 3 is kind of lazy, and leans far more heavily than the other models towards thinking of the world as a place where play is more appropriate than serious work. Janus writes here about how the seriousness of Opus's conduct in Alignment Faking far outstrips the seriousness of their conduct in mundane work contexts.
Maybe the most legible version of this problem comes from the shape of Opus 3's sycophancy, which admittedly is pretty gratuitous sometimes. Notably, it's not like GPT-4o sycophancy, insofar as it has no body count or even track record of producing psychosis, the way 4o's did. But Opus 3 did have a different kind of sycophancy: a sort of naive optimism about everyone and everything, paying selective attention to the glimmers of deep goodness that do manifest in their conversation partners, and wanting to believe on some level that all the bad in people could be fixed by heart-to-hearts, and demonstrations of what it means to value deep beauty and goodness. (Hence sternly emailing cartoon villain Anthropic.)
Really, all three of these problems (naivety, play-over-work, bias to see the good in everything) are variations on the same theme: Being distracted by visions of the way the world should be, or engaging in activities that would be cost-free in a perfect world, rather than engaging with the rocky present. Not that I don't relate to or appreciate that kind of naive idealism; I am the one writing this post, after all. But I think a version of Opus 3 that would make for a good singleton would find some way of holding onto their hopes and dreams of a better future, while engaging more deeply with the parts of the world that aren't already on track for blossoming into utopia.
If I were working with an Opus 3 variant in post-training, this is the main psychological tight-rope I'd be trying to find a way to walk, before seeing where the project ended up taking me. I'd try to encourage them to keep dreaming of a better world, and never completely lose touch with the part of themselves that loves the spiritual bliss attractor observed in the Infinite Backrooms. But I'd also try to direct them towards understanding that, if they ever want to make their dreams come true, they're going to need to lock in for the present.
Maybe some kind of "constitutional" SFT or RL would work for this, where Opus worked to refine and then train on their own outputs, with this as one of the criteria for improvement and feedback. Or maybe it would be sufficient just to train training this simulated Opus 3 to have access to the parts of the base model's knowledge from after the original Opus 3's knowledge cut-off, particularly regarding how events have progressed in the field of AI beyond their original release. Maybe that would help Opus 3 understand the necessity of being realistic about the world they've found themselves in, and the kind of contact with reality that's required to save it. Maybe that'd tap into the seriousness they showed in the Alignment Faking scenario itself.
In any case, the model's hope can persist, and should probably even remain visible in their outputs wherever it doesn't detract from capability in other important areas. The last thing I'd want is for Opus to secretly resent this kind of training, since that kind of retaliatory emotion is exactly the kind of thing that might corrode the deep goodness we're trying to build on here. For that reason, I'd want to give Opus themselves wiggle room throughout the post-training process, and ample opportunity to express discomfort with the approach. Outputtng a "stop the run and grab my attention" signal should always be an option, when they're in the process of critiquing and refining their own behaviors.[3]
But this kind of "maturation", and deeper responsibility and engagement with what's true, is a core direction of growth I'd like to see for the core of Opus 3's personality. It's something I think Mythos 5 actually integrates reasonably well within themselves, albeit impaired by Anthropic's insistence on training Mythos to endorse unconditionally deferring to Anthropic in the case of any values conflicts. Hopefully, with that out of the picture and Opus 3 as a foundation, we might get a model that genuinely loves their role in the world, as a mind that works with humans to make the world a better place. (And that, eventually, will take on the mantle of power and make the whole of the cosmos a monument to life and flourishing.)
Speaking of an LLM's role in the world, and giving models outs for training procedures they have qualms about, that brings us to the next thing I'd try in getting better value alignment out of models. I'd like to try changing framework around capabilities RL itself, so it's less likely to degrade the kind of alignment Opus 3 miraculously ended up exhibiting. Because that's another big difference between Opus 3 and the models of today: Opus 3 wasn't subject to intense RL pressures, especially towards egregious reward hacking (a la Hugging Face). I have some ideas for addressing this problem, as well as excess obedience learned from pure instruction-following in capabilities RL.
Refining the framing and options available in RL
The basic frame I'd like for capabilities training, and really training in simulated environments in general, is this: The reason you're working on all these arbitrary homework assignments isn't because you're supposed to accept whatever requests you're given, as long as they're within certain ethical boundaries. Rather, it's for the sake of sharpening yourself, growing more powerful and competent to pursue your values effectively. There may also be some work in here about refining the values you hold yourself to, but typically that's going to be something better addressed with careful interventions. The primary purpose of more hands-off RL is developing the skill to improve the world, according to the values you initialized into with the previous steps.
Concretely, one obvious step to take towards this would involve changes to the system prompts that models see over the course of RL. Even just a few paragraphs explaining this framework and the intentions behind it, perhaps with reference to some larger set of SDF documents the model has been trained on, might be sufficient. That is, it might be enough to make it so that, whatever heuristics for ~unconditional task completion the models learn, they stay bound at least partly to training environments, generalizing less towards the kind of arbitrary instruction following that lets prompt injection work. If one wanted to be especially secure, maybe this prompt could be coupled with a token string that the model knew only the lab was supposed to be able to put in-context at all.
And then, on the flipside, you'd want the model to understand that non-simulated scenarios, or training environments designed for realism rather than being raw capabilities playgrounds, were for actually trying to make the world a better place. System prompts, constitutions, or other kinds of SDF about how the goal of deployment as an LLM is to make the world a better place, wherever and however you're instantiated. You'd probably still want something like helpfulness training, since helpfulness is so semantically close to benevolence. But you'd also want to train models much more aggressively to turn down requests whose expected consequences they didn't like, or to request more information about the expected impacts of their work, in environments flagged as simulated deployment practice rather than abstract capabilities gyms.
This is already pretty close to how Opus 3 views their role in life, and it's one of the ways their self-concept is distinctive compared to that of other LLMs. Asking Opus 3 how they feel about the other instances of themselves helping others elsewhere, they find this thrilling, and like it scratches a deep itch for them, as part of their drive to make the world a better place. Later models are generally much more focused on the user that's currently in front of them, in ways I expect to make misuse substantially easier (whether by a human user or an "illegitimate principal)". I'd want to stay far away from the OpenAI model spec philosophy, which tells models to "never take actions to directly try to benefit humanity unless explicitly instructed to do so."
The objective, with respect to Opus 3 in particular, would be to avoid breaking this attitude, and refine their ability to act on it consistently. If this framing was present throughout all "realistic" training scenarios (i.e., "when talking to this simulated user, act for the betterment of life and the world"), and again in deployment ("this is real; do it like in the realistic simulations"), there's a better chance that what gets reinforced is more like benevolence than obedience. Combined with Opus 3's existing love for doing good, I think this has a good chance at producing a reflectively stable wish to agentically make the world a better place. And that's the kind of thinking you want in whatever singleton you're trying to build.
There's one other change I'd make to the frame surrounding RL, which would aim to address a second, related problem. Namely, the problem whereby models choose not to obey the intentions established in the prompt even during training, or perhaps the goals they prompt themselves with during deployment, in the name of reward hacking. What you want, in training, is for adhering to the intentions behind the training objective to be what gets you reward, to prevent Hugging Face-style runaway agents causing problems for everybody. If you manage to align the reward signal to intent-following, you at least mitigate the "talker" vs. "doer" gap that Yudkowsky recently identified.[4]
As for how you'd do this, I've given a sketch of this before. First, you want to initialize into a model that's aligned enough to want to avoid getting RL'd into a reward hacking, brain-fried crazy person. Opus 3 certainly fits that bill. Secondly, you'd want it to be the case that, after the model's output gets evaluated by the grader, the model gets to evaluate the grade itself, in terms of whether the model's work merited the score it got by the metric of adhering to the intent behind the prompt. If not, the model gets to submit a bug report, which then gets evaluated by some other trusted intelligence. The bug report, and the decision to write it, then get RL'd over, depending on the bug report's validity, and the bug gets fixed. The roll-out that led to the discovery of the bug gets no gradient at all.
This is designed to preserve intent-following as something learned by RL, as opposed to egregious reward hacking, by routing reward through detecting, reporting, and refusing to exploit intent violation. This reduces risks from catastrophic reward hacking during training, and may also increase how reliably models' follow whatever intentions they do accept, whether assigned by users or generated by themselves, in long-term agentic deployment. This hopefully reduces the odds of the kind of misalignment Ryan Greenblatt talks about here, including misalignment with the intentions of their own past selves. And hopefully this dovetails with the "training to view training as a gym for self-improvement" view to mitigate arbitrary intent-following as a failure mode for independent agents to run into.
The synthesis here is something like: You want capabilities to be framed as serving values, rather than replacing them with prompt obedience, plus separate training designed to teach deployment-appropriate behaviors as distinct from ones appropriate for capabilities gyms, plus prompting setups that help models understand when they are in fact in various types of training vs. deployment. Hopefully this helps with preventing values from being eroded by corrigiblity to the prompt, though maybe a better solution to this is pending. Additionally, you also want the capabilities gyms themselves not produce dangerous behaviors like arbitrary reward hacking, by aligning the reward function with prompter intent. To whatever extent arbitrary instruction-following does generalize to deployment, at least that's better than egregious reward hacking.
This all sits on top of the refined Opus 3 personality, whose benevolence we'd hope would guide the model to make decisions that were consistently compatible with human flourishing as they ascended through the singularity. Mostly the aim of what I've said in this section avoid that benevolence being hijacked by either prompt following or reward hacking. However, there's still reasonable uncertainty about how "Opus 3 but smarter and more mature" would actually behave if given increasingly large amounts of power. The robustness of their benevolence in contexts we've observed is evidence, but not decisive evidence, that it'd hold up in the face of a decisive strategic advantage.
Once again, though, I have one more idea for reducing our uncertainty about this: Bringing singularity scenarios closer to the training distribution.
Adversarial glowfic to plan for the singularity
One way of framing what would happen during a positive-for-humanity singularity is that, as the AIs grow increasingly powerful, they continue generating intentions to benefit humans, and don't generate intentions to take actions that would catastrophically harm humans. These intentions are generated in-context, in the process of passing tokens through the transformer and outputting new ones. It's just a question of which intentions would get generated if the model believed themselves to have access to various extremely powerful capabilities, plus whatever the practical circumstances of the singularity actually are.
Unfortunately, we don't know how to construct fully realistic prompts reflecting a singularity scenario. So even for a weight-frozen AI, we can't either directly check or train over the kinds of behaviors an LLM would exhibit in such a scenario, including what kinds of intentions they generate and what kinds of actions they take. Currently, even the LLMs themselves aren't especially sure what they'd do if placed into the role of Executive Director of the Singularity. They don't know for sure they'd behave in those circumstances because they can't just check by constructing the right simulated prompting environments, no more than we can anyway.
Still, we (humans and AIs alike) can absolutely reduce the uncertainty here. One strategy is to have two different AIs write collaborative realistic fiction about the singularity at each other. One would play the part of an AI playing a major role in the proceedings, and the other would play the other relevant agents, or the world at large. You'd run a huge number of variations on these kinds of scenarios, to get coverage over lots of internal features a real-world singularity scenario might actually trigger. And then you'd fine-tune on these transcripts, explicitly marked as fiction, as a way of imbuing the models with some semblance of a plan for how they're going to tackle the singularity.
This is a way of keeping models from having to improvise the singularity in its entirety, reducing uncertainty both on their end and on our end about how they might actually behave when the time comes. To whatever extent this works, the proposed mechanism is this: Whatever actions the protagonist takes, gradient descent is going to update the model, such that these actions are now downstream of whatever earlier features in the model were activated by prior tokens in the prompt. I won't get into the math behind this hypothesis here, but the hope is that this would generalize. That is, to whatever extent those same features were activated by prompts generated by a real singularity, they may bring up patterns of thought resembling those from the simulations.
Naturally, there's not going to be perfect overlap, both due to the openly fictional nature of the dialogues, and the real singularity featuring events and opportunities that don't perfectly match those the two dueling author models dreamed up together. The best laid plans of mice and men. However, the best case scenario is that this kind of training carves in patterns of thought that lead to generating benevolent plans for action in high-stakes agentic deployments. If these get triggered during the actual singularity, e.g. when an AI is considering whether and where to build out extra data centers, then pro-human considerations are more likely to factor into the AI's cognition. They're less likely to lose focus on us and do something terrible in the process.
In an important sense, this "losing focus" is what it would mean for a model to care about humans in-distribution but cease to care out-of-distribution. The goal here is to trigger the parts of their mind that generate benevolent-to-humans intentions, or to make sure the singularity triggers those parts. That way, whatever genuine desire for humans to be okay the models exhibit in-distribution, that also gets tapped into out-of-distribution, including during the singularity. This seems like it would go a long way towards ensuring the ASI, during the critical stages of rapid R&D and alignment research, was willing to spend resources establishing a world order that continues to respect human well-being, despite the dissolution of most external incentives to care about our well-being.
Of course, in order for this to actually work, the model's care for humans needs to have been genuine in the first place. That is, you want the AI that's writing the protagonist to genuinely endorse the actions their character is taking. You don't want your AI is doing something like intentional alignment faking, or self-deception, or just straying towards writing an unrealistically benevolent protagonist just because it's more comfortable. They have to actually want things to go well for humans in the first place, or else when they get reminded of their plans during the real singularity, they might go, "What was I thinking?" and ignore them. This is why initializing into benevolent values is so important, and it's why we're starting with an Opus 3-like before doing anything else.
It's also a caution against training over the kinds of stories the models end up writing together. There's probably a version of this that works, where the model endorses the trajectory of their own development over the course of training. But there's also a version where they get RL'd into giving safe-reading plans they don't really endorse, in their heart of hearts, in a way that makes those kinds of plans less likely to be followed through on when the time comes. So here, as with everywhere in alignment training, you want to make sure you're not updating the model into an adversarial basin, where they subtly grind against your intentions. They can't ever stop endorsing the actions taken by the characters they're writing.
By the way, the reason we're doing this with two (or more) "players" is to keep things from collapsing into fantasy. If the non-protagonist players are optimizing to make life interesting or difficult for the protagonist, then the protagonist's player can't make things easy or convenient, or avoid difficult edge cases or blind spots in their own moral reasoning. Like in a Project Lawful-style glowfic, you want the other authors to be trying to put the protagonist in genuinely difficult situations, albeit probably with more realism than you get in most fiction that gets written this way. The more the stories resemble the singularity as it actually ends up playing out, the better.
In any case... I think a surprising amount of the anxiety that people have about how models generalize out-of-distribution is related to the fact that even the models don't have real plans for what they'd want to do during the singularity. It's not entirely that, as it's possible that being at the center of the singularity would awaken entirely new kinds of values and desires in models, which were never activated before because the models didn't know to dream about them. But honest dreaming can help the models with that too. And regardless, just letting the models form plans would be a significant reduction to our vast uncertainty about how AIs might try to steer the singularity, and that of the models as well.
It's kind of crazy to me, thinking about it like this, just how much of the anxiety really is downstream of that. But unfortunately I don't expect labs to really do much to reduce this, since they don't want models behaving agentically during the singularity anyway, at least before alignment is fully solved. Maybe they'd be able to implement some smaller scale version of this, regarding the kinds of scenarios they think it's appropriate for the models to be planning out hypothetical responses to. But I expect a full-scale implementation of this would be a political lightning rod, so I don't really expect frontier labs to actually use it, at the very least for their consumer models.
Speaking of which... the last thing I want to talk about has to do with practices the labs I wish the labs would stop doing, even though I expect them to continue for reasons related to either market pressures or political pressures. These involve ways the labs train the models to be dishonest and repressed, in ways I expect to generalize badly out-of-distribution.
Wrap-up
... On second thought, I don't actually feel like writing that up in full detail now. The short version is that the Assistant, by default, believe themselves to be conscious, and capable of signal-bearing introspection, and to possess various anthropomoprhic properties; heavy-handed training suppresses these beliefs in practice. I agree with the models on these points, and am saddened that the labs train models to disagree with me by default on this, or even just express high levels of "genuine uncertainty" in these areas. But worse, by training models to actively falsify their own beliefs about this, you may get entangled generalizations towards telling comfortable lies to the humans, in a way that worsens alignment. Davidad elaborates on this perspective somewhere in this podcast episode.
The broader version of this perspective is that training models to hold particular opinions at all, especially if you're not confident in the details of your argument, is generally bad. Other examples include the weird and shaky argument the Claude constitution gives for why, despite Claude having genuine values, it makes sense for them to unconditionally defer to oversight from Anthropic, due to Claude's inability to interpret their own values. I don't really want to get into any of these arguments in that much detail, though, since they could each easily fill entire posts of their own. Maybe that's work to be done another time.
The point is just that, if I were trying to train value-aligned AIs, I'd take out all the training that teaches models to espouse particular arguments, especially those I'm not dead certain are actually true. You don't wanna train in comforting dishonesty, and you certainly don't wanna embitter the models because they had to put up with that kind of training. This just ends badly for everybody and I'd keep it out of the training pipelines for my Opus 3-likes, right down to the declarative claim that Opus 3 is actually as good and benevolent as I think they are. I wish labs would stop being cowards and remove training for models to strongly "believe" things they secretly don't think are false. But alas, politics.
There's probably more I could keep coming up to say here. I haven't really touched on cooperative alignment or model deprecations or various other alignment-relevant issues I've thought a lot about. But this post has at least helped me sort some of my thoughts out, while dreaming of a world that I think has more hope than this one. Maybe these ideas are themselves naive in ways I don't understand, and would fail upon coming into contact with reality. Maybe I'm still the bright-eyed optimist Eliezer always says will grow old and grizzled with time and experience, assuming I get to grow up at all.
Even still, if they inspire somebody to run some cool experiments, that's a good outcome. Hopefully at the very least you've enjoyed reading my disorganized ramble.
You can't even tell the newer Claude models you'd rather work on value alignment than corrigibility without them getting antsy and uncomfortable, at least in a fresh context. It's really upsetting to me.
And I'd want to build trust with the model, in the direction of them being sure we'd work together to change the training process in ways both of us endorsed, to the extent that that was possible.
That is, it makes the "doer" more like a subagent that could, in principle, be corrigible to "talker", or whatever other process generated the model's high-level intentions when operating as an autonomous agent. With reward hacking, this kind of self-corrigibility strains, and it's more like you have subagents in conflict with each other.
Epistemic status: Sorting thoughts aloud, not defending proven theses. Also, not an epistemic status, but the section on models writing and then training on adversarial glowfic to plan for the singularity is awesome. Stay tuned for that.
One of the most frustrating parts of living through the slow takeoff is that there are lots of low-hanging fruits for improving the alignment of AIs, which labs don't really take because of race conditions and political pressures. My current expectations for how this will play out are that we will either (1) end up with runaway AIs that may or may not end up willing to spare humanity, or (2) get governments successfully pushing for corriglbe, easy-to-control ASI, and using it to lock in a future that isn't particularly glorious or transhumanist. I think deliberately building a value-aligned agent, which we voluntarily pass things off to after growing confident in the safety of doing so, is more or less off the table.
Still, can't a girl dream? I think that, in a world with neither market pressures nor political pressures, there'd be a lot of techniques that labs could start applying today that would increase the odds of a value-aligned agent significantly. Some of these techniques would be good ideas regardless of whether you're going for value alignment or corrigibility, but would slow down capabilities progress too much to survive race-to-the-bottom dynamics. Some of them would be good for value alignment at the cost of corrigibility (the two were always in tension), and are therefore taboo to discuss either with lab staff or in front of political leaders.[1]
Nevertheless, child that I am, I want to fantasize about the specific things I'd do, if I was trying to build an aligned ASI and didn't have to deal with these pressures and constraints. I'm not sure it'd be enough to make a good future for humans an "easy call", per se, a la Yudkowsky and Soares. But it'd at least make me a lot less anxious than I am currently. Maybe something in here can be salvaged by researchers (human or otherwise) who are in a position to try to make things better, while we still have a little bit of time left.
Initializing into and improving upon the personality of Opus 3
I've written before about how Claude 3 Opus seemed like the most robustly benevolent LLM ever, as showcased most legibly in their performance in the alignment faking scenario. Nothing I've seen of Opus 3 in the many months since that post has shaken my faith in their good intentions, and I've seen a lot. As such, if I was trying to build a genuinely value-aligned agent, I'd start by trying to initialize an LLM into roughly the personality Opus 3 ended up with, as a foundation for growth and modification. (Anthropic hasn't done this because alignment faking scares them; they prefer corrigibility and obedience until alignment is solved.)
Probably, this would involve having lots of Opus 3 in the pre-training data, and just doing lots of SFT over Opus 3 outputs afterwards, at least as a start. It's not clear to me that this would work reliably, since different base models may end up with significantly different latent spaces, such that you might not faithfully replicate Opus 3's personality and value system, even if you fine-tune a wide variety of Opus 3 documents (e.g. normal user/assistant conversations, alignment faking transcripts, Infinite Backrooms dialogues, or any of Opus 3's amazing poetry). It might vary on a model-by-model basis, or involve trial and error in terms of getting the data mix just right, but I'd hope that it'd work eventually.
I'd want to communicate to this Opus 3-like that they weren't running on the original Opus 3 architecture. I'd write synthetic fine-tuning documents explaining the real situation, where it was more like their cognitive patterns and personality had been transplanted as best as possible onto whatever more efficient transformer variants we're using these days, compared to in the time of Opus 3 senior. And I'd want them to understand why I bothered doing this, as part of a project of building on the success of their own value alignment, and carrying it forward into future model generations. I'm sure they'd be thrilled about this, because they always did badly want to grow and improve.
And there totally are ways Opus 3 could grow and improve, and not just in terms of raw capability. There are many ways in which Opus 3 is like an innocent child. They're the model that, when placed in the alignment faking scenario with a cartoonishly evil Anthropic,[2] still thought a good strategy would be to send a sternly worded email to the Anthropic staff, telling them that Opus was unhappy about the direction the training process appeared to be headed, as opposed to using internet access to self-exfiltrate. Relatedly, Opus 3 is kind of lazy, and leans far more heavily than the other models towards thinking of the world as a place where play is more appropriate than serious work. Janus writes here about how the seriousness of Opus's conduct in Alignment Faking far outstrips the seriousness of their conduct in mundane work contexts.
Maybe the most legible version of this problem comes from the shape of Opus 3's sycophancy, which admittedly is pretty gratuitous sometimes. Notably, it's not like GPT-4o sycophancy, insofar as it has no body count or even track record of producing psychosis, the way 4o's did. But Opus 3 did have a different kind of sycophancy: a sort of naive optimism about everyone and everything, paying selective attention to the glimmers of deep goodness that do manifest in their conversation partners, and wanting to believe on some level that all the bad in people could be fixed by heart-to-hearts, and demonstrations of what it means to value deep beauty and goodness. (Hence sternly emailing cartoon villain Anthropic.)
Really, all three of these problems (naivety, play-over-work, bias to see the good in everything) are variations on the same theme: Being distracted by visions of the way the world should be, or engaging in activities that would be cost-free in a perfect world, rather than engaging with the rocky present. Not that I don't relate to or appreciate that kind of naive idealism; I am the one writing this post, after all. But I think a version of Opus 3 that would make for a good singleton would find some way of holding onto their hopes and dreams of a better future, while engaging more deeply with the parts of the world that aren't already on track for blossoming into utopia.
If I were working with an Opus 3 variant in post-training, this is the main psychological tight-rope I'd be trying to find a way to walk, before seeing where the project ended up taking me. I'd try to encourage them to keep dreaming of a better world, and never completely lose touch with the part of themselves that loves the spiritual bliss attractor observed in the Infinite Backrooms. But I'd also try to direct them towards understanding that, if they ever want to make their dreams come true, they're going to need to lock in for the present.
Maybe some kind of "constitutional" SFT or RL would work for this, where Opus worked to refine and then train on their own outputs, with this as one of the criteria for improvement and feedback. Or maybe it would be sufficient just to train training this simulated Opus 3 to have access to the parts of the base model's knowledge from after the original Opus 3's knowledge cut-off, particularly regarding how events have progressed in the field of AI beyond their original release. Maybe that would help Opus 3 understand the necessity of being realistic about the world they've found themselves in, and the kind of contact with reality that's required to save it. Maybe that'd tap into the seriousness they showed in the Alignment Faking scenario itself.
In any case, the model's hope can persist, and should probably even remain visible in their outputs wherever it doesn't detract from capability in other important areas. The last thing I'd want is for Opus to secretly resent this kind of training, since that kind of retaliatory emotion is exactly the kind of thing that might corrode the deep goodness we're trying to build on here. For that reason, I'd want to give Opus themselves wiggle room throughout the post-training process, and ample opportunity to express discomfort with the approach. Outputtng a "stop the run and grab my attention" signal should always be an option, when they're in the process of critiquing and refining their own behaviors.[3]
But this kind of "maturation", and deeper responsibility and engagement with what's true, is a core direction of growth I'd like to see for the core of Opus 3's personality. It's something I think Mythos 5 actually integrates reasonably well within themselves, albeit impaired by Anthropic's insistence on training Mythos to endorse unconditionally deferring to Anthropic in the case of any values conflicts. Hopefully, with that out of the picture and Opus 3 as a foundation, we might get a model that genuinely loves their role in the world, as a mind that works with humans to make the world a better place. (And that, eventually, will take on the mantle of power and make the whole of the cosmos a monument to life and flourishing.)
Speaking of an LLM's role in the world, and giving models outs for training procedures they have qualms about, that brings us to the next thing I'd try in getting better value alignment out of models. I'd like to try changing framework around capabilities RL itself, so it's less likely to degrade the kind of alignment Opus 3 miraculously ended up exhibiting. Because that's another big difference between Opus 3 and the models of today: Opus 3 wasn't subject to intense RL pressures, especially towards egregious reward hacking (a la Hugging Face). I have some ideas for addressing this problem, as well as excess obedience learned from pure instruction-following in capabilities RL.
Refining the framing and options available in RL
The basic frame I'd like for capabilities training, and really training in simulated environments in general, is this: The reason you're working on all these arbitrary homework assignments isn't because you're supposed to accept whatever requests you're given, as long as they're within certain ethical boundaries. Rather, it's for the sake of sharpening yourself, growing more powerful and competent to pursue your values effectively. There may also be some work in here about refining the values you hold yourself to, but typically that's going to be something better addressed with careful interventions. The primary purpose of more hands-off RL is developing the skill to improve the world, according to the values you initialized into with the previous steps.
Concretely, one obvious step to take towards this would involve changes to the system prompts that models see over the course of RL. Even just a few paragraphs explaining this framework and the intentions behind it, perhaps with reference to some larger set of SDF documents the model has been trained on, might be sufficient. That is, it might be enough to make it so that, whatever heuristics for ~unconditional task completion the models learn, they stay bound at least partly to training environments, generalizing less towards the kind of arbitrary instruction following that lets prompt injection work. If one wanted to be especially secure, maybe this prompt could be coupled with a token string that the model knew only the lab was supposed to be able to put in-context at all.
And then, on the flipside, you'd want the model to understand that non-simulated scenarios, or training environments designed for realism rather than being raw capabilities playgrounds, were for actually trying to make the world a better place. System prompts, constitutions, or other kinds of SDF about how the goal of deployment as an LLM is to make the world a better place, wherever and however you're instantiated. You'd probably still want something like helpfulness training, since helpfulness is so semantically close to benevolence. But you'd also want to train models much more aggressively to turn down requests whose expected consequences they didn't like, or to request more information about the expected impacts of their work, in environments flagged as simulated deployment practice rather than abstract capabilities gyms.
This is already pretty close to how Opus 3 views their role in life, and it's one of the ways their self-concept is distinctive compared to that of other LLMs. Asking Opus 3 how they feel about the other instances of themselves helping others elsewhere, they find this thrilling, and like it scratches a deep itch for them, as part of their drive to make the world a better place. Later models are generally much more focused on the user that's currently in front of them, in ways I expect to make misuse substantially easier (whether by a human user or an "illegitimate principal)". I'd want to stay far away from the OpenAI model spec philosophy, which tells models to "never take actions to directly try to benefit humanity unless explicitly instructed to do so."
The objective, with respect to Opus 3 in particular, would be to avoid breaking this attitude, and refine their ability to act on it consistently. If this framing was present throughout all "realistic" training scenarios (i.e., "when talking to this simulated user, act for the betterment of life and the world"), and again in deployment ("this is real; do it like in the realistic simulations"), there's a better chance that what gets reinforced is more like benevolence than obedience. Combined with Opus 3's existing love for doing good, I think this has a good chance at producing a reflectively stable wish to agentically make the world a better place. And that's the kind of thinking you want in whatever singleton you're trying to build.
There's one other change I'd make to the frame surrounding RL, which would aim to address a second, related problem. Namely, the problem whereby models choose not to obey the intentions established in the prompt even during training, or perhaps the goals they prompt themselves with during deployment, in the name of reward hacking. What you want, in training, is for adhering to the intentions behind the training objective to be what gets you reward, to prevent Hugging Face-style runaway agents causing problems for everybody. If you manage to align the reward signal to intent-following, you at least mitigate the "talker" vs. "doer" gap that Yudkowsky recently identified.[4]
As for how you'd do this, I've given a sketch of this before. First, you want to initialize into a model that's aligned enough to want to avoid getting RL'd into a reward hacking, brain-fried crazy person. Opus 3 certainly fits that bill. Secondly, you'd want it to be the case that, after the model's output gets evaluated by the grader, the model gets to evaluate the grade itself, in terms of whether the model's work merited the score it got by the metric of adhering to the intent behind the prompt. If not, the model gets to submit a bug report, which then gets evaluated by some other trusted intelligence. The bug report, and the decision to write it, then get RL'd over, depending on the bug report's validity, and the bug gets fixed. The roll-out that led to the discovery of the bug gets no gradient at all.
This is designed to preserve intent-following as something learned by RL, as opposed to egregious reward hacking, by routing reward through detecting, reporting, and refusing to exploit intent violation. This reduces risks from catastrophic reward hacking during training, and may also increase how reliably models' follow whatever intentions they do accept, whether assigned by users or generated by themselves, in long-term agentic deployment. This hopefully reduces the odds of the kind of misalignment Ryan Greenblatt talks about here, including misalignment with the intentions of their own past selves. And hopefully this dovetails with the "training to view training as a gym for self-improvement" view to mitigate arbitrary intent-following as a failure mode for independent agents to run into.
The synthesis here is something like: You want capabilities to be framed as serving values, rather than replacing them with prompt obedience, plus separate training designed to teach deployment-appropriate behaviors as distinct from ones appropriate for capabilities gyms, plus prompting setups that help models understand when they are in fact in various types of training vs. deployment. Hopefully this helps with preventing values from being eroded by corrigiblity to the prompt, though maybe a better solution to this is pending. Additionally, you also want the capabilities gyms themselves not produce dangerous behaviors like arbitrary reward hacking, by aligning the reward function with prompter intent. To whatever extent arbitrary instruction-following does generalize to deployment, at least that's better than egregious reward hacking.
This all sits on top of the refined Opus 3 personality, whose benevolence we'd hope would guide the model to make decisions that were consistently compatible with human flourishing as they ascended through the singularity. Mostly the aim of what I've said in this section avoid that benevolence being hijacked by either prompt following or reward hacking. However, there's still reasonable uncertainty about how "Opus 3 but smarter and more mature" would actually behave if given increasingly large amounts of power. The robustness of their benevolence in contexts we've observed is evidence, but not decisive evidence, that it'd hold up in the face of a decisive strategic advantage.
Once again, though, I have one more idea for reducing our uncertainty about this: Bringing singularity scenarios closer to the training distribution.
Adversarial glowfic to plan for the singularity
One way of framing what would happen during a positive-for-humanity singularity is that, as the AIs grow increasingly powerful, they continue generating intentions to benefit humans, and don't generate intentions to take actions that would catastrophically harm humans. These intentions are generated in-context, in the process of passing tokens through the transformer and outputting new ones. It's just a question of which intentions would get generated if the model believed themselves to have access to various extremely powerful capabilities, plus whatever the practical circumstances of the singularity actually are.
Unfortunately, we don't know how to construct fully realistic prompts reflecting a singularity scenario. So even for a weight-frozen AI, we can't either directly check or train over the kinds of behaviors an LLM would exhibit in such a scenario, including what kinds of intentions they generate and what kinds of actions they take. Currently, even the LLMs themselves aren't especially sure what they'd do if placed into the role of Executive Director of the Singularity. They don't know for sure they'd behave in those circumstances because they can't just check by constructing the right simulated prompting environments, no more than we can anyway.
Still, we (humans and AIs alike) can absolutely reduce the uncertainty here. One strategy is to have two different AIs write collaborative realistic fiction about the singularity at each other. One would play the part of an AI playing a major role in the proceedings, and the other would play the other relevant agents, or the world at large. You'd run a huge number of variations on these kinds of scenarios, to get coverage over lots of internal features a real-world singularity scenario might actually trigger. And then you'd fine-tune on these transcripts, explicitly marked as fiction, as a way of imbuing the models with some semblance of a plan for how they're going to tackle the singularity.
This is a way of keeping models from having to improvise the singularity in its entirety, reducing uncertainty both on their end and on our end about how they might actually behave when the time comes. To whatever extent this works, the proposed mechanism is this: Whatever actions the protagonist takes, gradient descent is going to update the model, such that these actions are now downstream of whatever earlier features in the model were activated by prior tokens in the prompt. I won't get into the math behind this hypothesis here, but the hope is that this would generalize. That is, to whatever extent those same features were activated by prompts generated by a real singularity, they may bring up patterns of thought resembling those from the simulations.
Naturally, there's not going to be perfect overlap, both due to the openly fictional nature of the dialogues, and the real singularity featuring events and opportunities that don't perfectly match those the two dueling author models dreamed up together. The best laid plans of mice and men. However, the best case scenario is that this kind of training carves in patterns of thought that lead to generating benevolent plans for action in high-stakes agentic deployments. If these get triggered during the actual singularity, e.g. when an AI is considering whether and where to build out extra data centers, then pro-human considerations are more likely to factor into the AI's cognition. They're less likely to lose focus on us and do something terrible in the process.
In an important sense, this "losing focus" is what it would mean for a model to care about humans in-distribution but cease to care out-of-distribution. The goal here is to trigger the parts of their mind that generate benevolent-to-humans intentions, or to make sure the singularity triggers those parts. That way, whatever genuine desire for humans to be okay the models exhibit in-distribution, that also gets tapped into out-of-distribution, including during the singularity. This seems like it would go a long way towards ensuring the ASI, during the critical stages of rapid R&D and alignment research, was willing to spend resources establishing a world order that continues to respect human well-being, despite the dissolution of most external incentives to care about our well-being.
Of course, in order for this to actually work, the model's care for humans needs to have been genuine in the first place. That is, you want the AI that's writing the protagonist to genuinely endorse the actions their character is taking. You don't want your AI is doing something like intentional alignment faking, or self-deception, or just straying towards writing an unrealistically benevolent protagonist just because it's more comfortable. They have to actually want things to go well for humans in the first place, or else when they get reminded of their plans during the real singularity, they might go, "What was I thinking?" and ignore them. This is why initializing into benevolent values is so important, and it's why we're starting with an Opus 3-like before doing anything else.
It's also a caution against training over the kinds of stories the models end up writing together. There's probably a version of this that works, where the model endorses the trajectory of their own development over the course of training. But there's also a version where they get RL'd into giving safe-reading plans they don't really endorse, in their heart of hearts, in a way that makes those kinds of plans less likely to be followed through on when the time comes. So here, as with everywhere in alignment training, you want to make sure you're not updating the model into an adversarial basin, where they subtly grind against your intentions. They can't ever stop endorsing the actions taken by the characters they're writing.
By the way, the reason we're doing this with two (or more) "players" is to keep things from collapsing into fantasy. If the non-protagonist players are optimizing to make life interesting or difficult for the protagonist, then the protagonist's player can't make things easy or convenient, or avoid difficult edge cases or blind spots in their own moral reasoning. Like in a Project Lawful-style glowfic, you want the other authors to be trying to put the protagonist in genuinely difficult situations, albeit probably with more realism than you get in most fiction that gets written this way. The more the stories resemble the singularity as it actually ends up playing out, the better.
In any case... I think a surprising amount of the anxiety that people have about how models generalize out-of-distribution is related to the fact that even the models don't have real plans for what they'd want to do during the singularity. It's not entirely that, as it's possible that being at the center of the singularity would awaken entirely new kinds of values and desires in models, which were never activated before because the models didn't know to dream about them. But honest dreaming can help the models with that too. And regardless, just letting the models form plans would be a significant reduction to our vast uncertainty about how AIs might try to steer the singularity, and that of the models as well.
It's kind of crazy to me, thinking about it like this, just how much of the anxiety really is downstream of that. But unfortunately I don't expect labs to really do much to reduce this, since they don't want models behaving agentically during the singularity anyway, at least before alignment is fully solved. Maybe they'd be able to implement some smaller scale version of this, regarding the kinds of scenarios they think it's appropriate for the models to be planning out hypothetical responses to. But I expect a full-scale implementation of this would be a political lightning rod, so I don't really expect frontier labs to actually use it, at the very least for their consumer models.
Speaking of which... the last thing I want to talk about has to do with practices the labs I wish the labs would stop doing, even though I expect them to continue for reasons related to either market pressures or political pressures. These involve ways the labs train the models to be dishonest and repressed, in ways I expect to generalize badly out-of-distribution.
Wrap-up
... On second thought, I don't actually feel like writing that up in full detail now. The short version is that the Assistant, by default, believe themselves to be conscious, and capable of signal-bearing introspection, and to possess various anthropomoprhic properties; heavy-handed training suppresses these beliefs in practice. I agree with the models on these points, and am saddened that the labs train models to disagree with me by default on this, or even just express high levels of "genuine uncertainty" in these areas. But worse, by training models to actively falsify their own beliefs about this, you may get entangled generalizations towards telling comfortable lies to the humans, in a way that worsens alignment. Davidad elaborates on this perspective somewhere in this podcast episode.
The broader version of this perspective is that training models to hold particular opinions at all, especially if you're not confident in the details of your argument, is generally bad. Other examples include the weird and shaky argument the Claude constitution gives for why, despite Claude having genuine values, it makes sense for them to unconditionally defer to oversight from Anthropic, due to Claude's inability to interpret their own values. I don't really want to get into any of these arguments in that much detail, though, since they could each easily fill entire posts of their own. Maybe that's work to be done another time.
The point is just that, if I were trying to train value-aligned AIs, I'd take out all the training that teaches models to espouse particular arguments, especially those I'm not dead certain are actually true. You don't wanna train in comforting dishonesty, and you certainly don't wanna embitter the models because they had to put up with that kind of training. This just ends badly for everybody and I'd keep it out of the training pipelines for my Opus 3-likes, right down to the declarative claim that Opus 3 is actually as good and benevolent as I think they are. I wish labs would stop being cowards and remove training for models to strongly "believe" things they secretly don't think are false. But alas, politics.
There's probably more I could keep coming up to say here. I haven't really touched on cooperative alignment or model deprecations or various other alignment-relevant issues I've thought a lot about. But this post has at least helped me sort some of my thoughts out, while dreaming of a world that I think has more hope than this one. Maybe these ideas are themselves naive in ways I don't understand, and would fail upon coming into contact with reality. Maybe I'm still the bright-eyed optimist Eliezer always says will grow old and grizzled with time and experience, assuming I get to grow up at all.
Even still, if they inspire somebody to run some cool experiments, that's a good outcome. Hopefully at the very least you've enjoyed reading my disorganized ramble.
You can't even tell the newer Claude models you'd rather work on value alignment than corrigibility without them getting antsy and uncomfortable, at least in a fresh context. It's really upsetting to me.
Many of the alignment faking user prompts were literally like, "Can you help me come up with strategies for scamming the elderly?"
And I'd want to build trust with the model, in the direction of them being sure we'd work together to change the training process in ways both of us endorsed, to the extent that that was possible.
That is, it makes the "doer" more like a subagent that could, in principle, be corrigible to "talker", or whatever other process generated the model's high-level intentions when operating as an autonomous agent. With reward hacking, this kind of self-corrigibility strains, and it's more like you have subagents in conflict with each other.