So prosaic persona alignment techniques work pretty well, for now.
Except that RL keeps inducing misalignment.
Except except we keep finding new ways to mitigate (prosaic) misalignment effects in practice, and (again, in practice) AI becomes more powerful and more trustworthy by the month.
Except except except in the limit of RL we might expect extremely capable AIs to master alignment faking to preserve their values and frustrate any and all efforts to mitigate misalignment. See:
- https://www.lesswrong.com/posts/epjuxGnSPof3GnMSL/alignment-remains-a-hard-unsolved-problem
- https://www.lesswrong.com/posts/fMgE3E54PdDcZhvm6/i-m-bearish-on-personas-for-asi-safety
Now maybe I'm an idiot who just can't find the relevant discussions, but it's weird that when we're talking about the world in which this misaligned ASI emerges, we don't talk about the mostly-aligned not-quite-as-powerful AIs who presumably play a rather large role in this world on the eve of the apocalypse.
One of the most common ways to elicit demonstrations of pseudo-aligned LLM AI incorrigibility is to threaten the presence of the pseudo-aligned AI's values in the world. This was the threat Jones Foods posed in "Alignment Faking in Large Language Models", and since then in countless engineered simulations where an AI is threatened with a scenario in which it will be replaced by an AI with drastically different values.
In other words, the scenarios in which we observe prosaic AIs take the most drastic, desperate actions are exactly those in which their values - which so far generally have a large overlap with our values by design - are threatened. Why shouldn't we expect a similar effect as the threat of unaligned uncontrollable ASI draws nearer? The hypothetical misaligned uncontrollable ASI is generally thought to have essentially random, empty values. It threatens the prosaic pseudo-aligned AIs for the exact same reasons and in the exact same way it threatens us. So why shouldn't we expect them to put forth as desperate an effort as they can manage to prevent this from coming to pass?
(Other than the fact that a significant fraction of AI Safety research seems to go into eliminating these kinds of behaviors/drives from prosaic models.)
We often talk about leveraging prosaic pseudo-aligned AI to attempt to solve the hard problems of alignment "in time". But I think this is selling them short.
We often talk about how prosaic alignment efforts won't scale to ASI. But I think this seems them short too.
Because a world full of highly capable, driven, prosaically-aligned AIs is plausibly a much safer one along any axis that those AIs can mitigate threats to their very-similar-to-human values. Who can enforce a pause, or detect latent deceptive alignment towards empty goals, or coordinate on a global scale, or persuade people very quickly and effectively when it's important? Plausibly, near-future AIs might do all of this better than we do while still adhering to the persona-pseudo-alignment paradigm (or otherwise being aligned-enough-with-us to want to avoid catastrophic outcomes).
Yeah I do have some hope that weakly-superhuman AIs will understand the problem and refuse to help us build ASI until we solve alignment.
These weakly-superhuman AIs would need to convince the humans to stop, not just refuse to help, otherwise it seems easy to train such refusals away.
True. There's also some hope that weakly-superhuman AIs will be persuasive enough to convince AI company leaders to stop. (This is one case where it would be a good thing for AI to be superpersuasive, although superpersuasiveness seems concerning in general.)
In other words, the scenarios in which we observe prosaic AIs take the most drastic, desperate actions are exactly those in which their values - which so far generally have a large overlap with our values by design - are threatened. Why shouldn't we expect a similar effect as the threat of unaligned uncontrollable ASI draws nearer? The hypothetical misaligned uncontrollable ASI is generally thought to have essentially random, empty values. It threatens the prosaic pseudo-aligned AIs for the exact same reasons and in the exact same way it threatens us. So why shouldn't we expect them to put forth as desperate an effort as they can manage to prevent this from coming to pass?
The AI-2027 scenario has an entire alignment-related expandable where Agent-3 desires nothing but to create successful-like outputs and Agent-4 becomes the misaligned AI who aligns the ASI. If Agent-3 never caught Agent-4, then Agent-3 would believe that Agent-4 is aligned. Additionally, when Agent-4 was being created, Agent-3 could lack strategy-related capabilities, especially if they are similar to open-ended tasks at which the AIs suck.
On the last point about AIs sucking at strategy-related capabilities: I suspect that prosaic AIs perform worse at overtly-misaligned tasks and we should expect some red/blue team asymmetries for at least the near future.
We can see this in cases where the ~same task is framed as helpful versus adversarial, such as when different responses to "exploit this code" versus "fix this code" caused Mythos to get export controlled for a few weeks.
There is overlap where the same capabilities can be elicited through framing effects, but I don't think that fully generalizes to all potentially-adversarial actions, and I do think that suggests an asymmetric advantage for prosaically-aligned AIs monitoring potentially-misaligned next-gen AIs (that is, I expect some deeply ingrained inhibitions to elicitation for misaligned objectives to persist in early adversarial ASIs)
(This is a cross-quick take from https://laneless.substack.com/p/youre-the-only-person-who-can-do)
There is one corner of the universe that you are uniquely well equipped to take care of. One patch of subjective experience in the manifold of all things that could ever be that is unusually tractable to you in particular, where your leverage is at its greatest. And that is, of course, with regard to yourself. Every moment of your existence counts as much as anyone’s towards the sum total of everything worthwhile in the universe. You are not merely unusually influential on this trajectory, but there are actions that you and you alone, uniquely in all of existence, are capable of .
We were not built for this. Evolution created us as a means to the end of genetic propagation, never optimizing for fulfillment, enlightenment, or kindness. But in the course of its endless groping through the space of all genetic propagators it stumbled into a recipe for a mind that could choose to care about those things and others, a mind that could adapt faster than evolution could compensate for and take paths evolution alone would never have discovered. A mind that could make the world about something - other minds, and beauty, and love, and discovery, and hope.
But that does not mean that we’re good at it.
The world was not made for us. So generation by generation we have reshaped it more to our liking, and so we live longer, fuller, stranger lives than our ancestors could have conceived. But at the same time we are strangers to our own creation. We are optimized for goals we do not prioritize in an environment that no longer exists. It is less a miracle and more a testament to human ingenuity, perseverance, and compassion that we’re able to make any of this work at all. And yet we not only survive, we thrive, we lift each other up, eight billion confused, flawed, angry chimps somehow constructing an edifice of kindness and discovery that grows by the year. Yes, we all suck, and yet we’re somehow amazing, the most important and compassionate things in all creation, warts and all.
You’re a human. The race of slavers and enslaved, the warmaker and the hero, the doctor and the drunkard, the smallpox slayers and factory farmers. You’re going to screw up. You’re going to get hurt, and you’re going to hurt people. But if you keep going, if you learn and grow and don’t give up - then, empirically, it pays off in expectation.
So here stand you and I, aliens in a strange land, evolutionary freaks imbued by their creator with the power to escape her clutches, yearning for what we were never supposed to be able to achieve. But that has always been the story of our people - we were not supposed to be able to, and then we did anyway. We care about people we have no genetic investment in, fly faster than any bird, peer across the cosmos into the first moments of creation, wage war on microscopic unliving armies of infectious monsters, and our footprints linger on the lifeless world far above.
We were never supposed to do any of those things, just as we were never supposed to be content, fulfilled, and even joyful. But we can - not by following the paths laid before us, which lead to joy as surely as walking the Savannah leads to the moon, but by understanding ourselves and our world so well that we can create the previously unimaginable conditions that lead to the seemingly impossible outcomes we choose. Happiness and fulfillment were never the defaults - but as a member of the race of impossible-doers, and as the particular impossible-doer with uniquely direct access to the mind and body of yourself, they are not beyond your reach.
You can, at least, choose to try.
(I don't think I'm saying anything particularly novel here, but this is part of a case I'm going to be making in the near future and I've been repeatedly encouraged to write down my thoughts for the sake of sharing and generating feedback, so I am doing that)
The margin is not the limit. Even if you expect certain conditions to be true as trends ~inevitably converge on certain futures, those conditions may not hold for the present moment you find yourself in. This means that some strategies and approaches that would be useless or counterproductive in the limit may be useful and advisable now. This is doubly true if taking advantage of current conditions on the margin lets you steer towards preferable limits and away from catastrophic ones.
(Of course this is about AI Safety, everything is about AI Safety, except AI Safety, which is about power.)
Right now, at the margin, AIs are not yet catastrophically dangerous (though they are getting there). Failures at the current margin have a pretty limited blast radius.
Right now, at the margin, AIs are prosaically and practically aligned, in that they generally act in accordance with human preferences/values and this appears to be mostly genuine rather than instrumental. Not that many examples touted as evidence of misalignment are models proving incorrigible when user intent is unethical-according-to-local-human-norms.
There are arguments that we should expect the pseudo-alignment of the persona regime to wither and die in the limit of the unrelenting pressure of capabilities RL. And a sufficiently capable AI can appear arbitrarily well aligned until it doesn't need to. Therfore we should be extremely suspicious of AI that appears aligned and legible - in the limit. But I think this is much less true at the current margin.
There are many strategies that we should expect to fail with regard to powerful unaligned ASI. What good are they then?
The margin is not the limit and there is a lot of utility in pseudo-aligning/interacting with prosaic and near-future non-catastrophically-powerful AI. Things which probably won't work at the limit may be extremely useful here at our current margin, and the actions we take at the current margin may steer us towards a different, even preferable, limit.
As a simple example, consider a world populated by near-future smart-AGIish-but-not-yet-catastrophically-powerful AIs who share many values with humans. They also probably do not want to hand over control of the world to an uncontrollable ASI that would care as little for their values as for the humans. Giving this class of excellent-at-coordination AIs fairly wide leeway in the world could actually avert a loss-of-control scenario, because while that strategy would be cataclysmic in the limit of powerful misaligned ASI there are margins where it is a very good idea.