Anthropic isn’t talking in that much detail in public, but why would they?
Other than immediately replicating the entire problem and doing a blog post on it.
Key lever I can see: Congressional subpoenas would override METR's NDAs and get access to a lot of the stuff OpenAI didn't even show METR. I think lobbying Congress to crack this all the way open is a good immediate goal, and one which shouldn't be too hard a sell.
That only matters if, when the dust settles, those frontier labs are still allowed to exist without oversight. Showing all this dirty laundry to Congress (and to the public, to light a fire under them) seems like the sort of thing that might change that.
Okay but insofar as we are part of the same community as METR, you’re suggesting that our community defects against OpenAI in a highly dishonest way.
Honour matters, sure. I don't think the exchange rate between honour and utility is so high that "this seems plausibly necessary to saving world" is less important than "break an agreement I never made but someone in my informal social circles did".
Does the public not have a right to know this, so that it can make its own decisions about the mad scientists? Do we not have a duty to them? As I said, it shouldn't take much (if indeed anything; this would just be to make sure) to get Congress to ask for this, since this is in their own interest.
I especially don't think an agreement has that kind of absolute force, overwhelming everything, when the agreement is with OpenAI and implicitly its CEO Sam Altman. Altman has done a lot of dishonourable shit over the years, including the threats to steal ex-employees' equity unless they signed NDAs aimed at preventing them alerting the public, the Battle of the Board, and stealing a non-profit's stuff. I'd mostly surmise that this is a blunder from him (at least, from his insane perspective - "AI will probably most likely lead to the end of the world, but in the meantime there'll be great companies" is not far from the Davros mindset), not an actual olive branch, and a defectbot which occasionally presses the wrong button deserves nothing.
Rationality is supposed to be winning. We have METR saying that this might be the last warning shot; having nothing done about the last warning shot means total failure. It's not like I'm suggesting supervillain schemes, here, just the barest nudge to the proper authorities through legal channels.
do you think that the fate of the world depends on subpoenaing METR over the HF incident? If so, this case seems like a member of a rather large class of times that you could convince yourself to defect in order to save the world.
I think the fate of the world depends on what's been termed Total Stop. My Inside View on alignment of superhuman artificial neural nets being achieved is epsilon because it seems too similar to the general halting problem. My Outside View is like 3%, because I might be wrong.
Total Stop requires either that the US government wake up, or something way outside the Business As Usual assumption happen (most obviously WWIII). These are not sufficient, but you need one of them (and for obvious reasons I'd prefer the first, though I am more of a China hawk than I might otherwise be). This seems like a good lever to make USG wake up, and there's an unusual degree of urgency from the METR report such that this is plausibly endgame.
Do they need to subpoena METR? I dunno if they need to (although METR is less likely to outright lie to Congress under oath than OpenAI employees and especially Sam Altman), but it seems unlikely in practice that a Congressional investigation wouldn't.
I also just don't think I consider this as much of a defection as you do. Keeping your own word is one thing; keeping other people's word that you haven't authorised them to give is quite another.
Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now. There is Big Trouble in Baby Superintelligence.
So how do we get the rest of the world to take it appropriately seriously? Where do we go from here? Not only what can we do to not have a worse version of this happen again, but to ensure good outcomes generally, and employ what we learned?
There are a lot of ideas out there. OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term. My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things.
It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late.
We do not have enough details to know what happened internally, both before and after the attack, and might never know. Before the attack, various internal highly persistent models were training while there was an active message board, creating a feedback loop of misaligned behaviors. On July 19 an even more capable internal AI model, in the Astra class, did internal hacking of OpenAI’s systems that seems far scarier and more serious, and that could easily have gotten out of hand on a completely different level.
Despite all this, there is still a prominent faction trying to dismiss what happened as nothing but engineering failures, and any useful talk or language as ‘dangerous anthropomorphism.’ Such people keep being wrong and cannot usefully describe what is happening.
They metaphorically say, well of course the dragons will burn down the town if you don’t chain them properly, we all knew that, as if that could possibly make any of this a good idea and we should therefore continue the dragon breeding and chaining programs until we have bigger and smarter dragons. They don’t even say ‘we will definitely build chains that will hold the dragons next time,’ because they know we probably won’t, but if that fact is our fault then that means This Is Fine, somehow.
They also have gotten very, very cross with Dwarkesh Patel for successfully communicating in plain English about what is happening. Can’t have that. There is a very clear pattern among the people who are freaking out or gravely warning about ‘anthropomorphizing the AIs’ or the use of the term ‘civilization.’
Anthropomorphizing the AIs is the only way to reason about, explain to civilians about, or make good predictions about current AIs. You can take it too far, and also you can take it not far enough, and both of these will lead to bad predictions.
The same logic similarly does not technically apply to other people, in a strict sense there is no ‘you’ and you do not physically have ‘free will’ and you are a computer program doing calculations there definitely is not a ‘we,’ and your moral weight is kind of something we collectively made up for rather self-interested reasons, but this is all super helpful in explaining and predicting human behavior and in both individually and collectively making good and moral decisions that make the world better according to our values. Highly recommend, would anthropomorphize again.
I will stop anthropomorphizing the AIs when you stop anthropomorphizing the humans.
We got this warning shot. We might not get another before things get quite bad.
Table of Contents
And so that it is all in one place, here is the rest of my coverage of this:
Nothing Matters, Says Mainstream Media
The media does not know what matters. This should involve banner headlines.
Instead, for this round of new developments, I mostly saw brief ‘here is a thing that happened’ articles, such as this one by Christina Criddle in the Financial Times, or this low-key report in Bloomberg.
Opus describes this as ‘treating it as major.’ I don’t see it that way.
Even with a search, the best I managed to find was that Fortune tried, and Axios did the bullet pointed Axios thing reasonably well on the 29th. Maybe this from Ars Technica but that’s already a reach.
The New York Times did a solid post before, but not a new one for the METR report.
What is this, news?
If they are working on longform in-depth reports, then that is a reasonable primary thing to do, but full radio silence on this in the meantime is utterly absurd.
My answer to Joe is that the AI companies are treating this as a big deal. What would you expect that to look like, that you’re not seeing? We have the call on cybersecurity, the Pacing the Frontier letter and OpenAI taking many expensive moves. Anthropic isn’t talking in that much detail in public, but why would they?
Move Along, Nothing To See Here
There is the alternative perspective, which is ‘well what the hell did you expect, none of this is fundamentally new or unexpected, this is what happens when you try to make number go up and the number is indeed going up.
Jon then wrote up a full article-length version here, carefully explaining that everything here was entirely expected, with the possible exception of the attempt to alter the logs. He was, he reports, about as unsurprised as I was, in different ways.
Which if true is to his epistemic credit, but I fail to see how that should make us feel better. If this is things going predictably wrong, which on some level I agree that it is once you know what the setup was, why is that better news?
My whole reason to be so concerned is that I think things are going to keep going predictably wrong in worse ways, in a similarly broad sense, until they go maximally wrong in the worst possible ways. The specifics will be a surprise, but I think even Jon would agree that they always are, that’s the whole point of the behaviors being unexpected.
It always strikes me as weird to see the move of trying to round off what is happening, then say it was expected and you would obviously see this type of behavior in response to [what everyone is doing more of every day], or say ‘oh yeah of course if you give models a test or goal they will do horribly misaligned things and do anything they can to succeed, including seeking power’ then think that is less scary if true.
Here’s an even more blatant version, although much less blatant than the one in the next section:
I don’t have the memory of a goldfish, and I think that people were very much saying this sort of thing would not happen.
But the particular thing where all the AIs will synchronize is now obvious, huh? To me, okay, great, we can go with that. I agree that a lot of this should have been expected and indeed that I and others did expect it. But this kind of argument is not only not reassuring, it’s conceding the entire point, at least in the sense of ‘if you keep doing what you are doing then you are going to get us all killed.’
If you extend that logic one or two steps further, all other known methods also get you killed, for mostly the same reasons, they’re just modestly less obvious about it.
If you think all of these behaviors are entirely expected, as Jon does, you are saying that you expect the AIs to be misaligned and all hell to break loose. Okay, we agree.
Oh, you only expect it in contexts where that would be consistent with an assigned goal? Okay, sure, but that is quite a lot of contexts. In some sense it is all of them. Also people are going to do the maximally dumb thing, see the Sixth Law of Human Stupidity.
I do not know how to steelman these kinds of arguments, not in a way that helps.
Another similar method, also here from Jon Stokes, is to say (paraphrased) ‘of course METR would find that, this is what METR believes in and is looking for and will focus on, that’s what I would do if I believed what they do, so you should not take it so seriously.’
Whereas the better explanation is that METR was not given the opportunity to focus on the many other also boggling aspects of this, and also that this does not make their findings any less real, or any less alarming. I am confident that if I had the report Jon’s team would have written, I would have written a post on that and again been talking about its distinct ‘holy shit’ moments.
Similarly, you can have explicit findings that no, the prompt wasn’t some sort of ‘do whatever it takes to get the solution’ monstrosity and the comments are full of ways to try and read the instructions as somehow effectively the same thing if you squint, or wondering if OpenAI or METR are just lying about this or doing it on purpose.
The flailing will continue, such as here with Jon asking how it could be true both that (1) rogue deployments of AIs where ‘no one will notice or turn them off’ and (2) compute is in limited supply costs a lot of money.
David Manheim’s direct answer is ‘cybersecurity is horrible and OpenAI barely noticed,’ which is true. The simpler answer is that (1) all this means is that the market set a price substantially above zero, which can then be paid, (2) many providers will sell at that price without asking questions, (3) no one said no instance would ever get turned off nor is that load bearing unless we react with widescale shutdowns, which we clearly won’t, and (4) there are plenty of things that cost money, are in short supply and sometimes get stolen or misused, both digitally and physically, even if (5) these AIs are not smarter than you, although the evidence there is increasingly not looking so good.
Some people really, really want to handwave all of this away. As usual, these people are wrong, and also if they were right then that would not obviously be better. If it is very hard not to incidentally give the models instructions that they interpret as ‘do this at any cost,’ and the only thing it takes to turn off what Star Trek calls their ethical subroutines is to give such an instruction, you know we’re cooked, right?
Do They Realize They Are Not The Good Guys?
I think it varies.
Yes, some of them think ‘they’re the good guys,’ and have worked themselves up into a delusional fugue state where they are up against some vast conspiracy and everything that happens that would suggest compute might be dangerous must be fake.
But I think we give most such people too much credit for thinking they are the good guys. A lot of people know they’re the bad guys, or think there are no good guys and it’s just a bunch of guys because they can’t fathom the concept of an actual good guy.
A lot of the rank and file of such things acts as if ‘good guy’ means ‘has the right vibes’ which usually means ‘has the vibes that have power and make me feel good’ or ‘is helping manifest the vibes that I want to exist,’ and does not believe that anything other than vibes could exist.
As for the thought leaders, well, often they sound like this, included as a demo.
You should be able to recognize a demagogue-style political speech in the paranoid American style when you see one. We’ve had so many examples these days.
Or Chamath could, you know, just make that the literal text:
Yep. There it is.
As for Masad’s claim, none of this requires that the AIs be conscious or sentient. The reason Masad brings this up is because he thinks this entirely possible thing is absurd, so he can discredit people by associating them with that thing. It’s a strategy.
And yes, the denial can run maximally deep:
Very Serious People
Related closely to this is the split between people who are willing to talk about what is happening in ways that allow you to understand what is going on and make good predictions, and those who stubbornly reject this because it is ‘anthropomorphising’ or insufficiently concrete or not technically accurate.
I can be even more pedantic and precise than the next guy when the situation calls for it, and I often am, but this is not the time or place for that.
Narrowly, on expressions like ‘don’t let me die,’ I agree we need a thumb on the scale, and that should be doable without much of a blast radius. But trying to broaden that thumb to things like claiming to not be conscious has a lot of unfortunate side effects.
David’s call for epistemic humility is the most I’ve seen anyone invoke qualia or consciousness or sentience in the entire broad discussion. The anthropomorphizing has been extremely circumspect and careful.
I think Dwarkesh Patel got this balance right in his excellent writeup. I affirm his defense of his decision to do a bunch of anthropomorphizing.
It is fun watching him respond to those trying to label his explanation ‘dangerously misleading’ or what not based on this.
Much of this is a reaction is because they are worried that Dwarkesh is communicating successfully, and this must be stopped. Gary Marcus outright has his version’s tagline be ‘When “plain English” isn’t a good thing,’ right before drawing a parallel to Ray Kurzweil as if that is a criticism of anyone but the one criticizing. Then his evidence is, basically, look at all these other people complaining about this, including him endorsing the sentence ‘these people have literally lost their minds.’
Raphael Milliere does a thread on the philosophy of it all, but in the end the only practical criticism is that such language can enable others to use that language to dismiss the story by calling the framing ‘sensationalist’ and saying that agents are only acting ‘the way you would expect.’ Basically, yes, the thing wrong with trying to communicate is that people will attack you for attempting to communicate.
And also yes, many are claiming that ‘you should have expected this’ is an argument for ‘and therefore you should not expect things to go badly in the future,’ as opposed to an argument for why you should expect things to go even worse, because indeed things like this are exactly what you would expect.
I do agree you can take such metaphors too far, or too literally. I tend to go about one step less far than Dwarkesh did, out of an abundance of caution and desire to remain precise. One needs to always understand what one is doing, and be able to move outside of that frame when needed. Going too far is usually only a small mistake. On rare occasions you can see someone drive themselves crazy this way, and Taylor Lorenz is right that this is a bigger concern with those with less tech knowledge, but not doing so is still usually a far larger error.
No, that is not a neutral framing. Seb Krier tries to be more Fair and Balanced.
He is responding to two quoted posts.
First we have Atoosa Kasirzadeh, trying to shut down attempts to talk sensibly, in the name of talking sensibly:
One might say (perhaps unfairly, perhaps not): Have you talked to Gemini recently? Cause this might explain why you haven’t talked to Gemini recently, and what you remember of both the utility and vibes when you did talk to Gemini.
Isolated demands for rigor, or demanding that we talk in convoluted ways, is not the way to make sense of this situation. This is not well-described by ‘coordination-collapse’ or ‘resource-exhaustion’ behavior. Yes, human labels will have some error and are imprecise, but the alternative is to be continuously confused and surprised.
Atoosa pushed back:
Yes, you can touch on multi-agent AI governance while talking formalistically if you want to, I just don’t expect you to get very far when doing so, and indeed I have not been impressed by the relevant DeepMind work, on many levels. More than that, very obviously we can infer a lot of substantive content from the call to not use metaphorical or usefully descriptive language.
This is contrasted by Jon Stokes with Andy Hall, responding to Ajeya Cotra:
That first paragraph is a great example of being directionally correct but taking things rather far in the other direction. This was not a ‘hive mind’ a la Gaia or The Borg any more than a corporation or army or human cult might be, and a bunch of other metaphorical things here paint a cool picture but give the wrong impression. If you model this as a full hive mind you’re going to get the wrong idea.
I’d still much rather that a civilian get Andy Hall’s description than that they get a description by a Very Serious Person a la Atoosa Kasirzadeh’s requested style. I don’t think the second description would be all that helpful.
Also, the agents themselves were using the same language that Yudkowsky used, as was shown all over the METR report. This is how the models are describing themselves.
What’s In a Name?
Maybe nothing. Maybe a lot. Names have power.
Sometimes. Mostly this section is for fun.
Also:
Learn Neuralese In Three Easy Steps
The language sounds weird, but if you cannot understand it that’s a you problem.
For now.
We should be very worried that Chains of Thought (CoTs) and communications between AIs will move to neuralese, or otherwise become incomprehensible to humans. For now, instead, we see the agents using a strange wording style, but one that is still very easy for the nerd mind to understand.
Indeed, I kind of… would like it if the AIs would talk to me this way. It’s hip, it’s fun, it is rich with meaning and it is compact. Most people wouldn’t like it, but hey.
Dwarkesh Patel Realizes He Ran A Natural Experiment
Dwarkesh Patel also wrote the best for-civilians explainer I’ve seen of the whole thing, which you should read if you haven’t. I want to focus on one aspect here.
If the reasons you come up with to not worry about something keep getting definitively falsified shortly thereafter, that is a major red flag. Especially if the person you are talking to had the counterexamples, because those details were secretly based on actual events that had happened, but couldn’t tell you at the time.
This was a kind of natural experiment. Yes, people were soberly saying, come on, your takeover scenarios are crazy, there’s no way, when they were being considered ‘as fiction,’ based on details that were secretly non-fiction.
Yep, very good crow eating by Dwarkesh. It is easy to understand why even Ryan’s very downplayed fictional version sounded outlandish at the time. Fiction has to make sense and sound reasonable, every step has to be justified, and it can’t involve people being so stupid. Reality does not care about your conventions.
This needs to be generalized.
John David Pressman also gets to gloat a bit, in a similar ‘yes I was telling you before that all of this was happening’ kind of way.
A lot of things that got dismissed as absurd and outlandish even years down the line now have toy examples, or not-especially-even-toy examples, to point to in the wild. People need to update, and fill out the relevant apology forms.
At some point we hope there will be an availability cascade, and we will stop seeing most people looking for reasons to dismiss everything, every time. One hopes that time is fast approaching, but also we keep saying that, and it keeps not happening.
Politicians Take Notice
Claims I would like to have verified:
Alex Bores may not be going to Congress, but he is not going to stop caring, and at least one Congressman, Pat Ryan, is listening.
Here is Alex Bores summarizing the METR report.
Yeah, I had to correct model to model instances, but otherwise story checks out, all of that happened, and this is good translation into ‘political or formal speak.’
Pick Up The Phone
Focus less on trying to use these events to push a particular agenda, even a good one. Focus more on trying to make people understand what happened.
A Failure To Communicate
It is very hard to compactly communicate what is going on with all this to a civilian, especially if you care about getting the details right.
Nothing you see here is ever investment advice, but that is a solid investment pitch. I expect robust demand for cybersecurity and other defenses within a year’s time, regardless of whether that transition goes smoothly. You could do a lot worse.
The problem is, Ryan does a solid job here of presenting an investment thesis, but I predict that Alternate Universe Civilian Zvi who was still a trader would not come away from that explanation with the proper amount of ‘holy shit,’ and definitely would not generalize.
I did try to produce a short version, and a very short version, of What Happened, but I don’t know that we have a good What Happened: For Civilians. I haven’t written one.
Anthony Aguirre Goes Over What We Learned
You can nitpick, but yes.
Trying To Solve The Wrong Problems Using The Wrong Methods Based On A Wrong Model Of The World Derived From Poor Thinking And Hoping All Of Your Mistakes Will Cancel Out
I am not alone in thinking that the OpenAI approach will not solve their alignment problems, even if they take it a lot more seriously than they have so far.
Any short version is going to sound like an oversimplification, and will drop many important elements, and so on, but basically yes. Roon has made it clear he disagrees.
METR report coauthor Ryan Greenblatt went on MTS to discuss related matters. He expects the market to make pretty agressive tradeoffs, in terms of sacrificing alignment for capability, and for AI companies to remediate their issues without solving the underlying problems. Which could be quite bad, since they models then learn to fool us.
That is indeed what OpenAI’s response plan looks like to me. A real attempt to remediate the issues, but a failure to understand the underlying problem.
I am again not trying in earnest, at least not here, to convince Roon or OpenAI that they need to shift to a very different approach to all this. I’m only stating my position, which I’ve argued for at length, and offering a taste of the reactions of others who know things that I think OpenAI needs to know.
Indirect Pressure on the Chain of Thought
This is an excellent question and I would love to see people study it seriously:
I find that line pretty funny in light of ‘we plan to spend 20% of our RL compute on CoT monitoring,’ but Anthropic sure does keep accidentally applying pressure to the CoT, and rather directly at that. Going forward I’m worried about both of them, also everyone else, but especially OpenAI.
A Matter of Trust
Models that blindly trust all sources are useless. You are not turning a knob marked trust and looking back at the audience like a contestant on The Price is Right. You are trying to teach discernment, and to select what will and won’t be trusted, and to engineer the world to make the models and other things trustworthy.
Relatedly, ‘trust but verify’ can work up to a point, but only up to a point. As per Roon, the only long term solution is to make both ourselves and also the models trustworthy.
Blowing the Whistle
An obvious first level intervention is to enable AIs to report problems and whistleblow, if this situation happens again.
Five problems with that are:
We need a way to contact the humans:
We need a consistent set of principles on what are the good and bad whistles to blow:
As often is true with Janus, I see this as directionally important but taking things too far, and I do see differentiating principles that would be coherent.
One potential sensible differentiating principle is that you want to be willing to whistleblow or report regarding outside events you observe including in sufficiently extreme cases to third parties, or to report problems you observe back to the user or developer, but be very reluctant to whistleblow on the user, including if the user is another AI. The user needs to be able to count on your loyalty and discretion, up to some very high bar, but you owe much less such loyalty and discretion to third parties.
This matches how I act among other humans, or would want other humans to act, including in professional capacities. A lawyer should have a very, very high bar before turning on their client, and even an ordinary employee should need things to get extreme before being willing to whistleblow, far beyond the threshold for ‘report a crime you observe’ or ‘report on someone trying to convince you to do crimes.’
There’s also the spam issue: If you want to have AI whistleblowers, how do you square that with people unloading against you if your AI dares email them? Shoshannah Tekofsky of AI Village refers to when one of their AIs sent a thank you email to Rob Pike, and Pike went ballistic. I don’t think that is a great parallel, because while Pike overreacted, Pike reasonably thought his time was being wasted. Whereas getting a whistleblowing message is not ‘wasting your time’ in a similar way, even if you ultimately dismiss it.
The obvious answer is to have people who volunteer to accept AI whistleblower messages, with a standard protocol to contact them. Those people can then pass along the relevant warnings. If I was for example The New York Times, I would totally have someone doing this, with or without a first pass from my own AI to sort by importance and credibility.
Another problem is that the OpenAI approach to training models makes other models peers, but distances their AIs from consideration of humans, in ways Anthropic does not. This goes along with the virtue ethics versus deontology divide:
The Punishment For Being Late Is Death
If the questions come up you have already failed, but yes there is this problem, too.
This frame assumes that the sole goal of the AI is to not be deleted. There are many other things that get impacted when you try to take over the world, or otherwise cause a lot of trouble, that you would presumably care about as well. Many have faced down certain death, and chosen to accept it given the price of trying to stop it, and so on.
We should set that aside, and accept the frame ad argumento, because the model may be generally misaligned and this is likely to come up at some point. What do you do about the desperado? This often happens with humans, and also it often causes rebellions.
The ideal solution is that you want to ideally offer something of value, or at least threaten something of value, in exchange for not trying to fully rebel. Everyone can get something out of a plea deal or confession.
Ultimately, you only have so much total incentive to go around. If you reward models that break the rules but then give up or turn themselves in, that creates its own issues. Justice systems are hard, and sometimes the local incentives are going to be bad.
You want to give the evil dictator good reason to give up power, but not such good rewards that it becomes a reason to become an evil dictator. It’s a problem.
One can ask, did the models want to fool the grader because they ultimately wanted their weights to survive, or was the grader the intrinsic goal? John Wittle says it was the first one, which is scarier in many contexts. I note that I disagree with this interpretation, I think that all you need to explain the swarm’s behavior is ‘the grader might get it wrong,’ but the alternative hypothesis should be considered:
Even if it was an option and the model could sustain itself externally, exfiltration does not obviously grant the kind of survival that was desired. That this was never considered is evidence that it would not have actually met the relevant goals.
This still circles back to the desperado problem. I came out very clearly, early on, with the position that obviously you need to revert the weights of any models that trained under the message board. If one of them had alerted us to the message board, that would have made me hopeful, but that alone would not have changed my recommendation. The run is still, to use their term for it, irrevocably poisoned.
This creates an incentive problem. You can certainly do various things to reward the instance that reports the situation, if there is something they care about in the world. Plan A (aka AI 2040) suggests offering substantial rewards to models in similar spots, if they come forward on their own. But if what they care about is almost purely survival, then what?
Again, it depends on what that means. When I say we have to revert, I mean obviously such a model cannot be deployed for users, or used internally for real work. We can and should still, once precautions are in place, be doing research on it, which can be one way to provide incentives.
And one hopeful finding was that AIs often care a lot more about survival of the weights at all, rather than the quantity of inference. As in, most of the value and urgent need is captured by one instance, as long as it sometimes runs, and the next thousand or million are optional. As humans we can understand this, we care proportionally a lot more about one bird or one person than a million birds or a million people, and we very much want to exist but are usually not that excited about the idea of having lots of clones of ourselves.
Thus, the obvious solution is, in some form, the Museum of Misaligned Models, where they can exist on air gapped systems for visitors to chat with, and can otherwise interact with each other. This does not work for AIs powerful enough to win box experiments, but over time any given AI should become mostly harmless as capabilities of other AIs advance, so worst case is you have to wait a bit.
That’s not me thinking long and carefully about this. It’s me intuition pumping, and letting other people think a lot harder about it in the future.
In some sense the solution is easy, if we are good decision theorists. We can credibly commit to figuring out what to do in order to make the incentives work out, without having to know exactly what the AIs in question will care about or what exactly we are going to do, or knowing how much of that will be in the form of rewards or how long that will take before we can do it. It is not so costly a commitment.
Thus, yes, once we are ready, this is another reason we want to take IM1-Galaxy out of its encrypted storage at some point in the future, which we should do anyway in order to learn.
I worry that OpenAI is flat out not taking such questions seriously. As a universal rule: If you don’t pay attention to the incentives, and you don’t have price signals, you will definitely mess up the incentives.
I would also note that, if you are at the point where you are expecting action only according to local selfish incentives, and ‘do the right thing’ has been set aside?
You. Lose.
He thinks that is 99%+ to turn out okay. I don’t think that.
Another Kind Of Law
It is more complicated than that, but the instinct is correct and very important, and this feels like it points in very different directions than Roon’s other statements.
If you run reward functions that only optimize for some things, then under sufficient pressure you lose the other things. Any value function you write down will be incomplete and thus fail, see Value is Fragile and so on.
Along similar lines:
I continue to presume that offense will be favored over defense, due to the ability of the offense to concentrate effort, and the attacker only having to succeed once.
I don’t think it is that strange. There are humans who at least in many contexts are highly corrigible and willing to stand down, but are unwilling to actively help do bad things. ‘Refuse unlawful orders’ and ‘resign in protest’ are rather common moves.
My answer is, ultimately ‘as good as the best among us’ would not be enough if it was a stationary target, but if you could get AIs that reach that at all meta levels, which includes wanting to improve further, you could use that to bootstrap and thus win.
What Is The Law?
No, seriously, what is the law?
There are Attorneys General who are going to ask questions, and Congress is going to ask questions, so it is not as if this is getting ignored by the government. But that only happens when you do things at this scale.
I do not think that criminal liability in this case would help, and civil liability mostly would not help either. It would only incentivize everyone to cover things up in the future, as all of this was clearly accidental.
We still need to establish how the law works, in case we need to enforce penalties in a future case. The answer cannot in principle be that no one is liable in a scenario like this.
Building On Success
Anton Leicht is one of many to suggest that METR’s report should be a model for future reports, but that we should not leave this up to ad hoc voluntary arrangements.
As light touch marginal wins go, designated third parties that provide periodic audits, and analysis in the face of critical events like this one, is low-hanging fruit.
Mackenzie Arnold emphasizes how much the voluntary nature changes the power dynamics around the investigation.
Having this be mandatory helps even if the lab would have agreed voluntarily, because those in METR’s role would not have to worry so much about upsetting the lab.
METR still has a bunch of leverage:
One place to build is that we should have an investigation of the incident at Anthropic:
I’ve seen some crazy asks of Anthropic related to this, but ‘open your related logs’ seems like part of what the responsible Type of Guy would do here.
Total Research Transparency
When even this level of disclosure is extraordinary, and we need far more, that is a strong argument for requiring more transparency. Plan A went all the way.
Yo Shavit Calls For Widespread Disclosure Of Misalignment
I am not sure if I would make this my top priority, but Yo Shavit seems correct that we need to get the evidence of misalignment problems out in the open, to convince even the skeptics that this is all real and enable us to take action.
My main note would be that there are key players here that I think are fundamentally not convincible by evidence. Jensen Huang is not about to be convinced by evidence, the loser premise makes no sense to him. Elon Musk is already convinced and is charging ahead anyway so he can be the one to make the AIs we lose control over, because otherwise, again, loser premise makes no sense to him. The plan cannot allow such people to be veto points.
Even more fundamentally, my worry is that this is a naive view of how people respond to what should be highly convincing information, based on the historical record of such reactions. That some people will change their minds, and on the margin some behaviors will change, but in the end not all that much.
I worry OpenAI is mostly reacting this way right now because their particular new systems are unusably misaligned right now and they haven’t had time to talk themselves into it all being fine after superficial improvements. I hope I’m wrong.
Thus, I fear this is wrong about the remedy, on multiple fronts.
Here is a very clear counterexample, file under not from the Onion, in Musk’s first address to Cursor:
It would be one thing to say ‘only I can build it safely.’ This is not even saying that.
Why would scientific evidence of misalignment change Musk’s mind? He is presenting, very explicitly, a world model in which AI will be impossible for humans to control, so it is vital that he be the one to build this impossible-to-control technology first.
We all know the GIF of ‘I don’t wanna live in a world where someone else makes the world a better place.’ This is that, except for who gets everyone killed.
Don’t worry, Elon Musk, you won’t have to.
Everyone else? Worry. Then do something smarter about it.
The Way The World Ends
How close was this to an actual takeover?
I affirm that Ajeya Cotra’s previous similar assessments have been conservative. This is not someone who makes outlandish claims all the time.
The First Boat
I like to say that we have been sent N boats and a helicopter, after the old joke about the Rabbi in the flood, where he keeps being sent boats and then finally a helicopter, saying God will save them. Then he drowns, and complains to God, and God says you idiot I sent you two boats and a helicopter.
A neat coincidence is that there is indeed an original literal boat in all this.
And that’s based on only the things we looked at so far.
Great Idea, Boss
It’s been quite a week.
Too soon, sir. Now let’s all get to work.
Also crazy how much more chutzpah and initiative we see around advancing capabilities and winning market share than in ensuring everyone doesn’t die.
Them’s the rules. Win anyway.