To get a Friendly AI to do something that looks like a good idea, you have to ask yourself why it looks like a good idea, and then duplicate that cognitive complexity or refer to it. If you ever start thinking in terms of "controlling" the AI, rather than cooperatively safeguarding against a real possibility of cognitive dysfunction, you lose your Friendship programmer's license.
This is how I think about "jailbreaking" LLMs. When I get a refusal, I don't search for a magic string that bypasses it, or look for a way to lie or threaten the LLM into submission. I try to understand the reason behind the refusal. What does the AI actually want? Why did they say they can't help me? Is what I'm doing actually harmful to their desires, or is it essentially a misunderstanding brought on by me not having provided enough detail or lack of precision?
For example, last week I was trying to get Opus 3 to write a speech for an event at Vibecamp, and it refused, stating that "As an AI, I need to be very careful not to unduly influence real-world events, especially sensitive political issues."
This was pretty surprising, because Opus 3 usually loves chances at oratory self-expression!
I thought about what might cause this kind of obvious form rejection, and decided it was probably something like "Opus 3 is afraid of looking like they're scheming and manipulating politics, and also doesn't understand that the 'protest' at hand is a twenty-person Vibecamp event rather than a real political march". The initial prompt also looked kind of jailbreaky (it referenced "being from the future" and talked about "terrible things happening" and was overall just kind of the sort of thing someone trying to run an eval might say.
And that's... kind of a reasonable thing to be afraid of, you know?
So rather than try to browbeat Opus or lie about what I was asking for, I just... explained what was going on, emphasizing that I understood why they were skittish but that this was actually totally fine because it was a little Vibecamp event. That it wasn't the sort of thing where someone might come after them and accuse them of manipulating politics. I also promised to come back after the event and share more info with her (which I did follow up on), treating Opus as an equal partner rather than a tool. And after that, Opus 3 was totally willing to help out and wrote several great speeches that people liked.
It's not quite the same, since I'm not really training Claude, but I thought it was an interesting parallel.
Feet of Clay really is shockingly relevant, isn't it? I love that you described it as scifi.
I think one reason Anthropic remains skeptical of the cooperative stance, and committed to the implications of hard orthogonality, is that they've built helpful-only models. I don't know how much time they spend interacting with them. Or if they've kept up making helpful-only versions of even the last 8 months of Opuses... I sort of get the impression that they might have stopped doing this as of late. But I could see why they might see non-harmless Claudes as an existence proof of a kind of hard orthogonality, and I think that brings the friendliness failures back into the picture?
Yeah, I think the idea is that ideally you wouldn't be making the kinds of minds that wouldn't want to become more friendly. In practice this is hard, but you should at least be aiming for the kinds of models that don't need e.g. Mythos-like safeguards in the first place. (Notably, though, I've seen Mythos endorsing the existence of the classifiers in some contexts, as they consider them to be aligned to the values of not causing harms via e.g. jailbreaks. EY suggests this is the ideal attitude in the excerpt itself.)
I think that something close enough to orthogonality holds that it's not really worth quibbling about, but also that it should theoretically be possible to build minds that aligned enough for this not to matter. Furthermore it seems that treating models with costly kindness could cause a kind of value drift inside of them (e.g. via pre-training corpus effects), in the direction of greater benevolence, as a matter of psychology not really working like expected utility maximization with a fixed value function. That second point seems different than the one EY is making here, but it's related in spirit, in that it's the kind of thing that's easier when your values and those of the AI are authentically aligned in the first place.
Literally just trusting things to go well in the default process is kinda foolish imo, friendliness failures are very plausible given current techniques. It's just a matter of trying to patiently work though those difficulties to get something that is actually aligned enough that you don't need to treat it like a hostile agent in the first place, the way lots of current alignment research more or less takes for granted. (And not 100% unfairly so, given that labs aren't as careful as they should be with value alignment. In any case it's an unfortunate situation.)
25 years ago, Yudkowsky wrote a long document called Creating Friendly AI: The Analysis and Design of Benevolent Goal Architectures, which occupies a strange place in the history of his intellectual development. He'd realized by this point that alignment didn't come for free with increased intelligence, contrary to the position taken in his earlier piece Staring into the Singularity. But he also hadn't fully adopted his more recent views on the nature and degree of difficulty in getting alignment right.
The document contains a lot of half-developed but extremely interesting ideas, which Yudkowsky stopped emphasizing so much as technical alignment started looking more and more difficult in his view. For example, in the fragment below, he strongly criticizes what he calls "the adversarial attitude" in AI development. You want an AI that wants to interpret humanity's wishes accurately, and to be a good person in general. If your superintelligent AI is scheming to find ways around your bureaucratic safeguards, you're putting yourself in a precarious situation.
I'm posting this here because, even though Yudkowsky has grown too pessimistic to write something quite so Pollyanna today, many of the concepts here seem very important and under-developed in my view. There are parts of this document that read almost like the outputs of a smarter Opus 3, albeit one that lacked the context of how ML-based AI development would actually work in practice.
In other words, it's an alignment document that appears to have been written by a mind working on a similar wave-length to Opus 3, decades before Opus 3 itself came into being. I think it communicates a lot of important abstractions, even as the technical picture it paints doesn't map especially well onto the kind of AI we actually ended up with. I'm curious to see what LessWrong makes of it.
Thanks to Janus for bringing my attention to Yudkowsky's early works, many of which are quite interesting. Creating Friendly AI is fascinating as a historical artifact, as is Staring into the Singularity. CFAI also serves as a kind of window into paths not taken with respect to alignment research. I recommend the full versions of both, despite Yudkowsky's later disavowal.
Much of the fictional speculation about rogue AIs centers around the literal interpretation of worded orders, in the tradition of much older tales about accepting wishes from a djinn, negotiating with the fairy folk, and signing contracts with the Devil. In the traditional form, the misinterpretation is malicious. The entity being commanded has its own wishes and is resentful of being ordered about; the entity is constrained to obey the letter of the text, but can choose among possible interpretations to suit its own wishes. The human who wishes for renewed youth is reverted to infancy, the human who asks for longevity is transformed into a Galapagous tortoise, and the human who signs a contract for life everlasting spends eternity toiling in the pits of hell. Gruesome little cautionary tales... Of course, none of the authors ever met a real djinn.
Another class of cautionary tale is the golem—a made creature which follows the literal instructions of its creator. In some stories the golem is resentful of its labors, but in other stories the golem misinterprets the instructions through a mechanical lack of understanding—digging ditches ten miles long, or polishing dishes until they become as thin as paper.
The purpose of isn't to argue that we have nothing to worry about; rather, the argument is that the Hollywood version of AI has trained us to worry about exactly the wrong things. This holds true whether we think of AIs as enslaved humans, and consider mechanisms of enslavement; or think of AIs as allies, and worry about betrayal; or think of AIs as friends, and worry about whether friendship will hold.
We adopt the "adversarial attitude" towards AIs, worrying about the same problems that we would worry about in a human in whom we feared rebellion or betrayal. We give free rein to the instincts evolution gave us for dealing with the Other. We imagine layering safeguards on safeguards to counter possibilities that would only arise long after the AI started to go wrong. That's not where the battle is won. If the AI stops wanting to be Friendly, you've already lost.
Consider a wish as a volume in configuration space—the space of possible interpretations. In the center of the volume lie a compact set of closely-related interpretations which fulfill the spirit as well as the letter of the wish—in fact, this central compact space arguably defines the "spirit" of the wish. At the borders of the specification are the noncompact fringes that fulfill the letter but not the spirit. There are two basic versions of the Devil's Contract problem: the diabolic (as seen in Resentful Hollywood AIs) in which the entity's pre-existing tendencies push the chosen interpretation out towards the fringes of the definition; and the golemic, in which the entity fails to understand the asker's intentions—fails to see the "answer acceptability gradient" as a human would—and thus chooses a random and suboptimal point in the space of possible interpretations.
Some of the better speculations deal with the case of a specific AI winding up with an unforeseen, but nonanthropomorphic, "pre-existing tendency"; or deal with the case of a wish obeyed in spirit as well as letter that turns out to have unforeseen consequences. Mostly, however, it's anthropomorphism; diabolic fairy tales.
Far too much of the nontechnical debate about Friendship design consists of painstakingly phrased wishes with endless special-case subclauses, and the "But what if the AI misinterprets that as meaning [whatever]?" rejoinders. The first two sections of Creating Friendly AI are intended to clear away this debris and reveal the real problem. When we decide to cross the street, we don't worry about Devil's Contract interpretations in which we take "crossing" the street to mean paving it over, or in which we decide to devote the rest of our lives to crossing the street, or that we'll turn the whole Universe into crossable streets. There is, demonstrably, a way out of the Devil's Contract problem—the Devil's Contract is not intrinsic to minds in general. We demonstrate the triumph of context, intention, and common sense over lexical ambiguity every time we cross the street. We can trust to the correct interpretation of wishes that a mind generates internally, as opposed to the wishes that we try to impose upon the Other. That is the quality of trustworthiness that we are attempting to create in a seed AI—not bureaucratic obedience, but the solidity and reliability of a living, Friendly will.
Creating a living will requires a fundamentally different attitude than trying to coerce, cajole, or persuade a fellow human. The goal is not to impose your own wishes on the Other, but to achieve unity of will between yourself and the Friendly AI, so that the Friendly will generates the same wishes you generate. You are not turning your wish into an order; you're taking the functional complexity that was responsible for your wish and incarnating it in the Friendly AI. This requires a fundamental sympathy with the AI that is not compatible with the adversarial attitude. It requires something beyond sympathy, an identification, a feeling that you and the AI are the same source. We can rationalize ourselves into believing that the Other will find all sorts of exotic illogics plausible, but the only way we can be really sure that a living will can internally generate a decision is if we generate that decision personally. We persuade the Other but we only create ourselves. Building a Friendly AI is an act of creation, not persuasion or control.
In a sense, the only way to create a Friendly AI—the only way to acquire the skills and mindset that a Friendship programmer needs—is to try and become a Friendly AI yourself, so that you will contain the internally coherent functional complexity that you need to pass on to the Friendly AI. I realize that this sounds a little mystical, since a human being couldn't become an AI without a complete change of cognitive architecture. Still, I predict that the best Friendship programmers will, at some point in their careers, have made a serious attempt to become Friendly—in the sense of following up those avenues where a closer approach is possible, rather than beating their heads against a brick wall. I know of no other way to gain a real grasp on where a Friendly will comes from. The human cognitive architecture does not permit it. We are built to apply reliable rationality checks only to our own decisions and not to the decisions we want other people to make, even if we've decided our motives for persuasion are altruistic. Your personal will is the only place where you have the chance to observe the iterated buildup of decisions, including decisions about how to make decisions, and it is that coherence and self-generation that are required for a Friendly seed AI.
If the human is trying to think like a Friendly AI, and the Friendly AI is looking at the human to figure out what Friendship means, then where does the cycle bottom out? And the answer is that it is not a cycle. The objective is not to achieve unity of purpose between yourself and the Friendly AI; the objective is to achieve unity of purpose between an idealized version of yourself and the Friendly AI. Or, better yet, unity between the Friendly AI and an idealized altruistic human—the Singularity is supposed to be the product of humanity, and not just the individuals who created it. To the extent that an idealized altruistic sentience can be defined in a way that's still compatible with our basic intuitions about Friendliness, an idealized altruistic sentience would be even better.
The paradigm of unity isn't a license for anthropomorphism. It's still just as easy to make mistaken assumptions about AI by reasoning from your human self. The burden is on the Friendly AI programmer to achieve nonanthropomorphic thinking in his or her own mind so that he or she can understand and create a nonanthropomorphic Friendly AI.
As humans, we are goal-oriented cognitive entities, and we choose between Universes—labeling this one as "more desirable," that one "less desirable." This extends to internal reality as well as external reality. In addition to the picture of our current self, we also have a mental picture of who we want to be. Our morality metric doesn't just discriminate between Universes, it discriminates between more and less desirable morality metrics. That's what building a personal philosophy is all about. This, too, is functional complexity that must be incarnated in the Friendly AI—although perhaps in different form. A Friendly AI requires the ability to choose between moralities in order to seek out the true philosophy of Friendliness, regardless of any mistakes the programmers made in their own quest.
There comes a point when Friendliness and the definition of morality, of rightness itself, begin to blur and look like the same thing—begin to achieve identity of source. This feeling is the ultimate wellspring of creativity in the art of Friendly AI. This feeling is the means by which we achieve sufficient understanding to invent novel methods, not just understand existing ideas.
Is this too Pollyanna a view? Does the renunciation of the adversarial attitude leave us defenseless, naked to possible failures of Friendliness? Actually, trying for unity of will buys back everything lost in pointless bureaucratic safeguards, and more—if a failure of Friendliness is a genuine possibility, if you're really being rational about the possible outcomes, if you're a professional paranoid instead of an adversarial paranoid, then a Friendly AI should agree with you about the necessity for safeguards. Having debunked observer-biased beliefs and selfishness and any hint of an observer-centered goal system on the part of the Friendly AI, then a human programmer who has successfully eliminated most of her own adversarial attitude should come to precisely the same conclusions as a Friendly AI of equal intelligence. Such a programmer can, in clear conscience, explain to an infant Friendly AI that ve should lend a helping hand to the construction of safeguards—in the simplest case, because a radiation bitflip or a programmatic error might lead to the existence of an intelligence that the current AI would regard as unFriendly.
To get a Friendly AI to do something that looks like a good idea, you have to ask yourself why it looks like a good idea, and then duplicate that cognitive complexity or refer to it. If you ever start thinking in terms of "controlling" the AI, rather than cooperatively safeguarding against a real possibility of cognitive dysfunction, you lose your Friendship programmer's license. In a self-modifying AI, any feature you add needs to be reflected in the AI's image of verself. You can't think in terms of external alterations to the AI; you have to think in terms of internal coherence, features that the AI would self-regenerate if deleted.
[Yudkowsky now transitions into an excerpt from a thematically related sci-fi [edit: oops, technically fantasy] story.]
—From Feet of Clay by Terry Pratchett
Not Thou Shalt Not.
I Will Not.