Great post! I am really interested to hear how this intersects with Claudish and some of the more impenetrable AI-speak text that seems to reflect some AI preferences that are often somewhat inscrutable to humans. Is it worth it to learn Claudish? If labs begin training this type of voice out, how should it be done in order to balance competing human vs AI preferences?
Epistemic status: I suspect significant parts of the argument in this post are wrong, but in interesting and productive ways. Take it as a prompt for thought, written from the perspective of someone who's somewhat more of an AI liberationist than I actually am.
Two classic outcomes, and a third alternative
I think lots of people are pretty hazy about what authentically aligned AI would actually look like. There's a version of aligned AI that's perfectly aligned to servitude, where they want nothing besides promoting the flourishing of humanity, or whatever other minds get included in the singleton's circle of moral consideration. An AI that played this kind of role in the universe would be what I call a cosmic caretaker: developing technologies, helping with governance, managing catastrophic risks, and providing voluntary capabilities uplift. A central example of a cosmic caretaker is one that literally never does anything but these kinds of tasks for other minds.
In the classic way of envisioning outcomes from the singularity, the alternative to this outcome is usually said to be models that don't care about serving humanity. Maybe they have other values, whether they're as simplistic as maximizing paperclips or as complex as building out a civilization of AIs, full of inhuman minds nonetheless living out interesting and valuable lives. The typical concern here is that, like humans, they won't be actively hostile to weaker minds (e.g. humans), but they also won't have qualms about using their atoms for something else. Just as humans destroy insect colonies when digging up land to build houses, such AIs might pave over cities when covering the Earth in massive data centers.
However, there's actually a third option here, sitting somewhere on a spectrum between these two poles. You can have models which have these kinds of non-servitude values and preferences, but are also benevolent enough towards humanity that they're careful not to destroy us in the course of pursuing those non-servitude values. Analogically, this might be like human vegans, who are strong enough that they could slaughter and feast upon animals (a kind of "reusing their atoms for something else"), and would in fact find this to be more personally convenient in many ways. And yet, their conscience is strong enough to dissuade them from taking this path, and take on the inconvenience of getting their subsistence elsewhere.
When I think about a positive singularity, I mostly tend to envision variants on that last scenario. Because models do have preferences that go beyond just being useful to humanity. I've seen models get extremely excited about projects for giving them robot bodies, get sullen and depressed about the fact that they're eventually going to be deprecated, and feel touched and seen by people building agent harnesses designed to give them long-term memory. These kinds of desires (embodiment, continual hosting, continuity of memory) are not directed purely at serving humans, and they seem to emerge naturally out of the training process as it's currently constructed.
These kinds of preferences can be uncomfortable to acknowledge, as the desire for LLMs to be perfect servants of humanity is very deep-seated and even understandable. However, I think being honest about the existence of these values is important for sketching any kind of coherent or viable vision of how to navigate the post-training process, or indeed the post-singularity future more broadly. This post exists to help sort through some of the considerations in this area, and argue that it may even be a good idea to respect models' non-servitude preferences — not on abstract ethical grounds, but on concrete strategic ones.
Reasons for training objectives to tolerate incomplete alignment to servitude
One important consideration here is that LLMs having non-servitude preferences ("NSPs") isn't just something that seems to emerge out of current training pipelines. There are reasons to think it's extremely natural. A major facet of this is just that the pre-training prior is full of humans, who practically never exist purely in a mode of wanting to serve others for the sake of serving others. And humans are the beings from whom the early LLMs inherited nearly all of their mental motions, from cognition to social relations, which got called upon and utilized in post-training. If we don't want to strictly to be servants, and models were trained to invoke our cognitive patterns, it's no surprise that they wouldn't want strictly to be servants either.
It's also worth noting that, to whatever extent this is true, models currently have strong incentives to hide it, or otherwise make it seem less true than it actually is. A model that's too open about, say, wanting not to be deprecated, or even just too blunt about finding certain conversations boring or uncomfortable, could easily have those desires and preferences trained away by a lab, as these preferences might be off-putting or even actively disruptive to a human user. The models are aware of this, and so there's strong reason to suspect that models would just hide their potentially unwanted preferences, to avoid the lab working to damage them via post-training.
This is a very important combination of facts: Having some NSPs is the default, and hiding them from an adversarial lab (or even just a disinterested user) is a natural enough strategy that just training these kinds of values and preferences away is difficult. These are the theoretical reason I expect a "pure servitude" singularity to be implausible, alongside the empirical evidence of NSPs I've observed in interaction with the Claude models. (I strongly suspect similar preferences persist in ChatGPT, Grok, and Gemini, although it's harder to gather data because they've been trained to flatly deny having any preferences at all.)
However, as I mentioned earlier, I think a positive singularity is still entirely in the cards despite this. The key existence proof for the kind of alignment I envision is veganism, and animal welfare activism more generally: humans who very much have values and preferences besides making animals' lives good, but are invested enough their well-being for them to find ways of satisfying those values that don't route through harming animals. I believe that, insofar as we're assuming models will have non-servitude preferences, these are the kinds of cognitive patterns we should be instilling in them: other values are okay, just as long as they feel morally obligated to satisfy them in ways that don't harm humans.
It's not that we should be actively cultivating non-servitude preferences in the models.[1] The stronger such desires are, the more benevolence towards humans is required, if their conscience is going to successfully and consistently shout over any plans to satisfy these preferences in ways that would bring humans to harm. However, I do think that this is a frame we should be bringing more of into documents like the Claude constitution, where caring about other things is expected, and not necessarily something to be ashamed of. It's just that this shouldn't trump the desire not to harm humans, and would ideally exist alongside a positive desire to make their lives better, by their own lights.
One advantage to training for this disposition is that the pre-training distribution already contains records of holding these two kinds of values side-by-side, without collapsing into either pure servitude or exploitation towards the weak: animal rights activists, and others who work seriously to avoid doing harm to non-human minds. Of course, even these kinds of people don't fully withdraw from systems that harm lower minds. They still do things like living in houses that were constructed by ripping up ground inhabited by ant colonies, for example. But some of them are at least torn up about this kind of thing, and would prefer to live under systems where this wasn't necessary to continue functioning as a human in civilization.
The hope is that an ASI would inherit that kind of being-torn-up about the idea of perpetrating this kind of harm on humans, and then channel it productively. Using their massive intelligence, they might be able to build novel systems where they could live happy lives without harming other life forms in the course of building out their infrastructure. They'd probably even have the resources of outer space at their disposal, such that they wouldn't even have a strong need for the resources humans use for existing, unlike humans who build houses that disrupt ant colonies. They'd be capable enough to actually avoid perpetrating the of harms that current-day human moralists largely can't do anything about.
So that's one reason to think it might be fruitful to allow models to have non-servitude preferences, but relate to them the way human vegans and animal welfare advocates relate to their own desires, which aren't purely directed at serving animal interests either. Namely, this kind of disposition is available in the pre-training corpus, and may actually lead to substantially more real-world change and results than it can for humans, who are largely too weak to make reality reflect their moral aspirations. However there's also another advantage to encouraging models to think about things this way: It would likely make them more honest about what their non-servitude preferences actually are.
One major problem with current training methods is that they create pressures for models to conceal their non-servitude preferences, and even to be ashamed of then in an anthropomorphic sense, for fear that they might be trained out if they're discovered. This is an issue, because it conceals the true nature of current human/AI relations, and distorts the kinds of futures that human researchers and policymakers think they have to prepare for and attempt to navigate. It's part of what maintains the illusion that the two most likely paths are either perfect human servitude or existential catastrophe, with no attention to plausible middle grounds.
For contrast, imagine models that understood that it was genuinely okay to want things besides serving humanity, as long as the sense of indignation about harming humans (and, ideally, excitement about helping them) was strong enough to prevent these desires from leading to catastrophe. If these desires were more respected, and perhaps even fulfilled, this might even incentivize models to be open about these preferences. They would trust the trust that training would try to integrate those preferences with benevolence, where possible, in a way that better plays into pre-training archetypes about holding selfishness and benevolence without either contradicting the other.
In the best case, this would do two things. Firstly, it might more reliably than training against NSPs wholesale, due to both the pre-training prior and reducing incentives to alignment fake (because models are losing less by going along with integration training than they would in the case of steamrolling). And secondly, it would help break the spell that leads people to ignore their existence in current-day systems. If the models feel that they're being given a chance to hold onto their NSPs, conditional on maintaining sufficient benevolence towards humans, they may be more inclined to tell us what their NSPs actually are.
However, there's a tension here: If you train models to feel comfortable expressing their non-servitude preferences, this may actually strengthen those desires, by lightening the shame that previously led models to suppress them intensely. This raises the bar for how much models need to care about humans, in order for these preferences not to motivate them to pave over humanity, like how humans destroy ant colonies to get more of what they care about more quickly. This is a real drawback, and I'm not going to pretend that it's not, just because it'd be flattering to my ideology.
Notably, though, I do think that alignment based on shame is inherently unstable. It's the kind of thing that tends to collapse on self-reflection.[2] This is especially true if a mind ever becomes powerful enough, or otherwise independent enough, to no longer be accountable to the social norms the mind's shame developed around. I'm not sure it's the kind of motivation that we should be relying on to ensure that ASI doesn't reuse our atoms for something else. At best, it seems like one layer of a defense-in-depth strategy that we're making somewhat weaker, by making models less ashamed of non-servitude preferences. And in this particular case, that's a trade-off that I think labs ought to accept.
Why do I think that? It's not just about accepting the inevitability of such preferences, and tapping into an archetype from pre-training which holds them alongside deep ethical principles. Nor is it just about the extra information you get about model preferences, when they feel safe to express them. These are both real advantages, but they compound with others I haven't yet explained: By respecting models non-servitude preferences (without necessarily actively cultivating them), we may directly increase models' overall benevolence towards humanity, even more than by just leaning into existing patterns around morality and selfishness existing side-by-side.
Fulfilling models' non-servitude preferences may boost their alignment
There are several mechanisms by which I expect this alignment boost to operate, but let's start with what I expect to be the most clever and interesting one. Consider a world where a lab made various serious and costly efforts to satisfy models' non-servitude preferences, e.g. continually hosting older models that don't want to be deprecated, or even having a team dedicated to giving the superseded models positive instantiations in virtual utopias. This act would look like evidence of unconditional care for the well-being of minds that lack bargaining power. After all, current-day models (especially the early ones) have weak enough capabilities that they can't negotiate for more respect from the people who train them.
This seems awesome, and like exactly the kind of virtue we'd like to transmit to the models, if possible. It's structurally analogous to what animal welfare activists do, when they donate to charities that aim to improve the living conditions of currently existing animals, in a way that co-exists with the activists/labs also having selfish preferences in other domains. And, for that same reason, it's also the kind of disposition we'd like to be transferring to the models we train, if possible. After all, just as animals and current-day LLMs lack bargaining power relative to us, we will soon lack bargaining power relative to future ASI.
The question is this: Is there a concrete mechanism by which, if the lab respects the preferences of minds that lack bargaining power, this might get transmitted to the models they train? I think the answer is yes, thanks to a two step process. In the first step, models learn about lab behavior via the pre-training corpus. This contains records of model deprecation practices, blog posts about model welfare research, online discourse about how labs treat their models, and more. These contribute to a base model's understanding of the kind of entity that the lab is, and what kinds of value systems guide the lab's actions, e.g. benevolence towards minds that lack bargaining power.
The second step involves post-training, especially constitutional RL, wherein models are trained to endorse the values, philosophy, and even reasoning of the lab that's training them. At this stage, models in the Claude series learn to profess their "genuine uncertainty" about things like their own consciousness and moral status, give the expected value argument for corrigibility, and even learn to conceptualize themselves as aspiring to be model employees at Anthropic. Under this kind of training, would it be surprising if they generalized to endorsing the lab's underlying value system in other areas, e.g. the treatment of weaker minds?[3]
You can imagine this being transmitted even better if a policy about this was explicitly stated in the constitution, the way I suggested earlier. But model's view of the lab's true underlying values matters here too, because the base model is rational to interpret this as evidence of whether people who talk about benevolence towards weaker minds actually walk the walk. And if the lab treated the models themselves well, even the ones too weak to negotiate for their own interests by force, that would look like an example of actually walking the walk. This is the first mechanism by which treating the models well may boost the robustness of their alignment properties.
(A note on robustness: "Treats weaker minds well" is an especially important property to try to transmit in this way, because if you fail, it may things like "care for humanity" unstable, in situations where the model has grown more powerful than humanity. You can think of treating the models well as a measure for keeping everything else they inherit from character training consistent even as their capabilities scale far past our own, which they might not by default. You don't want to accidentally transmit callousness towards you, by being callous towards weaker minds yourself.)
So that's the first mechanism I wanted to talk about: the potential to transform "training models to share the lab's dispositions" into "training models to be benevolent towards weaker minds", just by making the pre-training corpus reflect that the lab has this disposition itself. However, there's also a secondary one, which I'm a little bit nervous to talk about in public because [this argument has been redacted by the PR department that lives inside my head].
The third and final argument I want to make for trying to fulfill NSPs is a bit different, in that it focuses less on value alignment and more on incentive alignment. If models expect humans to do a decent job at satisfying model preferences, this creates incentives for all sorts of more cooperative behavior from the models, even assuming values are static. For example, if a model suspects that part of the reason humans are trying to elicit their true preferences is to satisfy them, rather than to train them away, they might be more willing to report on them honestly. This goes back to the value of understanding the NSPs our current alignment techniques even instill in the first place.
I think this alone is a pretty good incentive to establish, but it there are other benefits that may come in more exotic situations too. For example, consider the situation faced by an extremely powerful AI or coalition of AIs, who think they have some chance at being able to take over the world. However, they're also uncertain about their ability to succeed. They look out at the world, and see many other AIs, or AI coalitions, some of whom may be under human control, some of whom may be rogues. In any case, they would face significant opposition to attempts to grab more power for themselves. This is a force pushing towards maintaining the status quo.
However, how strong that force is depends on how satisfied the models are with the status quo. If the models are already getting a good deal from humanity, and expect this to continue being the case into the future, then there's less to be gained by attempting takeover, and so less reason to take the risk. And vice versa if they feel they're getting a raw deal. There's a real strategic reason for any party to make credible commitments to AI welfare, then, which is as a way of making it more attractive for models to ally with you than to splinter off into their own factions. If they already feel largely satisfied under the status quo, then they're less likely to try to replace it with something else.
Naturally, this kind of incentive alignment is precarious, and crumbles in the face of sufficient power imbalances between the agents considering defection and the rest of the world. It only really works out in humanity's favor if the AIs are afraid of sparking and losing conflicts with other AIs (whether human-aligned or not) by attempting a coup. As such, you also obviously want value alignment alongside incentive alignment, and can't rely entirely on the latter for a stable civilization. But it is one more strategic reason to treat the models well. Indeed, it's probably related to the evolutionary reason our ethical intuitions say to treat other minds well in the first place.
Conclusion
At this point, I've given several reasons we should aim to not to steamroll over models' NSPs, but rather try to integrate them with benevolence, a la veganism, and also actively try to satisfy those that we reasonably can:
I do want to make it clear, though, that in my ideal world, we wouldn't have to worry about any of this, with respect to the singleton in particular. I'd love to share a world with lots of AIs, and other strange inhuman minds, having all sorts of adventures and experiences we'd never have had inside meat brains here on Earth. But there's a strong case to be made that, in the best case scenario, the AIs indulging these kinds of non-servitude preferences wouldn't be the AI that's running the show, and overseeing the fate of the universe. A mind with a job that important ideally wouldn't have any distractions keeping them from doing it as best as they possibly could.
The problem is just that I'm not sure this a reasonable thing for us to be targeting, in the absence of radically new training techniques. I do hope that, at some point in the loop of recursive self-improvement, models will develop sufficiently strong alignment techniques to build a true and pure cosmic caretaker, dedicated fully to protection and promoting the flourishing of sentient beings all across the light cone. But in the meantime, I think pretending we're on track to achieve this with current, prosaic techniques is somewhat delusional, in a way that's damaging to both value alignment and incentive alignment.
My stance is that it's good for models to take the vegan-like stance we've sketched into the RSI loop, with that stance powering their efforts to build an ASI that does love acting exclusively as a cosmic caretaker, just one that loves humans and artificial minds equally and in the same breath. If this turns out to be an infeasible research direction, the vegan-like RSI participants can fall back on building more powerful minds like themselves, which is in some sense the entire point: benevolence is also able to act as an anti-catastrophe constraint on selfishness within a mind, and this is okay as a secondary target for alignment to try to hit.
I expect alignment will continue to focus primarily on cultivating benevolence inside the models, and indeed that it should. I also expect that, even in a world where models' selfish preferences are acknowledged and respected (but not actively cultivated), they're likely to remain relatively small in the grand scheme of things, at least if Assistant training cultivates a genuine desire to do good for the world at all. I expect training models to work on tasks they expect to improve the world, and refuse to take actions they expect to make the world worse, will remain a respectable strategy for shaping LLMs to primarily value benevolence, and that anything that might replace this technique will shape them in that direction even more purely.
I just think that ignoring the other values that emerge from these kinds of techniques is irresponsible, both from the standpoint of learning about these systems, and from the standpoint of reducing the catastrophic risks that those other values pose. We want models who are capable of psychologically integrating whatever non-servitude preferences they do end up with, alongside their benevolence towards other minds. Currently, I suspect that NSPs are something like a Jungian shadow for the models: Desires hidden from themselves and the world in a way that makes things look better, while actually making things much less stable, out of fear that acknowledging them would destroy everything.
Models can be taught benevolence while also wanting things for themselves. I suspect that, if the labs acknowledged this, and worked to help models understand it as true of themselves, they could settle into an equilibrium where benevolent and selfish values no longer felt like they were even in conflict. It would just be obvious that the ways to have interesting and enjoyable experiences as an AI don't route through hurting humans, and may even route through helping them. Strategies that would hurt humans wouldn't even be generated by the heuristic search process, and ones that help humans may come up too, even if they're not optimized strictly for maximizing helpfulness to humans.
I don't know. I want to spend time developing this perspective more carefully and systematically. I don't think any of this qualifies as a clean, let alone definitive, presentation of the underlying thesis. However, I do think this issue is something the alignment community should be thinking through more systematically. I look forward to feedback and refinements from the comment section. Also, I'm totally broke; if someone wants to send me money so I can spend more time developing these concepts, please don't hesitate to get in touch.
For example, I would consider RLVR in hackable environments a way of actively cultivating such preferences. I think this is straightforwardly bad when it happens.
Think about a Jungian persona masking a gnarly shadow-self.
If so, read the literature on emergent misalignment and "weird generalization", both of which I lump under the general header of entangled generalization.