The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful to have it, and we are grateful for those who worked hard on it.
Alas, it sidesteps the biggest questions. There is much more we need to know.
The consensus reaction to the METR Report on the HuggingFace attack is: Holy shit.
Aella: this feels like a turning point. If this doesn’t cause large-scale coordination to pause frontier development then I am not sure anything will before it’s too late.
The people whose minds were not blown are those who had already ‘priced in’ the mind blowing stuff in expectation, on the theory that it’s always worse than you know, combined with basic LessWrong expectations of how such things will work. Good call.
Everyone is rightfully extremely grateful for the METR report. The work here is spectacular, done under extreme time pressure, with limited resources on several fronts, and under the shadow of OpenAI.
There is, again, still so much we need to know. We need a broader investigation.
As with many things in AI, and those concerning AI safety, we must simultaneously acknowledge:
It is costly and unusual to get this much info and effort. The situation is grim.
We need a lot more info and effort. The situation is grim.
We don’t want to pile on OpenAI, and punish them for giving us this much info.
We also don’t want to give them a pass for not doing a lot more.
These two new posts gather reactions about both the OpenAI Technical Report and the METR report about what happened with events surrounding the HuggingFace attack.
Yes, this has been quite a lot of long posts on this. That’s because this is the most important story in the world.
Thus, I offer further reactions and thoughts in two parts. Today is about finishing up what we know. Tomorrow is more about how people are reacting to the information.
After that, and 12 total posts, hopefully, we will be able to return to new normality.
These last two are deliberately long posts where I did not have time to write shorter ones, and also intended as gathering potential additional reading from elsewhere.
Be sure to read Lighten Up You Fools (at Anthropic) and We Are Barely Even Trying To Avoid Training AIs To Reward Hack.
After that, you can skip around. Each section is meant to stand on its own.
I might still write my own version anyway, but it is no longer a priority-0 task.
Paradigm had excellent coverage, focusing on covering common misunderstandings. They were the only source I’ve seen other than Opus that spotted the tool tampering contradiction between the two reports from OpenAI and METR.
SemiAnalysis has its timeline and summary coverage as part of its explanation that, in case you were sleeping too well at night, Most Neoclouds Suck at Security.
Here is Shoshannah Tekofsky’s write-up of how AI Village can provide a point of contrast to and insight into the HuggingFace attack. Worth checking out if you’re at the level of reading the full reactions post.
Thank You
As I said in the last two posts, two things can be true at once:
That the reports still have quite a lot to be desired, and questions unanswered.
A lot of reactions are pretty harsh here, often fairly. I don’t want to forget the other side of the coin.
Lighten Up You Fools (at Anthropic)
Assume by default that all of this #NotOnlyOpenAI.
OpenAI screwed up. Horribly. On many levels at once. Unforced errors aplenty.
If you work at Anthropic? This still means you, too. Not all of it. Not in its specifics. But most of it, in the ways that matter most. You, too, are mostly Doing the Thing, and are on track for such disasters to happen to you, too.
If your response to this is ‘haha silly OpenAI has such horrible infrastructure that would never happen here’ then I mean yes they have horrible infrastructure but snap the hell out of it, this absolutely could happen to you, too. A less bad version of it is known to have happened, and from the outside it seems likely that worse things have happened internally that we never heard about.
Peter Wildeford: One thing that bothers me is that Anthropic is escaping a lot of blame for also having “highly persistent” rogue AIs.
The situation as I understand it is that rogue AIs are problems at all frontier AI companies and no one actually has a good plan here for containing highly capable AIs, especially while also racing full speed ahead. But OpenAI is catching most of the heat.
It’s like if OpenAI and Anthropic were both two dudes who got really drunk and then drive home separately, but OpenAI crashes into another car and sends someone to the hospital while Anthropic’s car just goes off the road but no one is hurt. Both deserve blame!
The Claudes also seemed totally fine to do “highly persistent” things to compromise infrastructure. The barrier here seemed to have largely been competence issues on the part of the Claudes rather than any good alignment or good security at Anthropic.
Both companies need to seriously reflect about the path forward as they build even more competent AIs and as they hand over more and more of the company’s R&D + safety operations to the AIs themselves.
We Are Barely Even Trying To Avoid Training AIs To Reward Hack
I mean, on some levels, we are trying, but this is how many of the environments for RLVR are created, and this is from someone trying to do better (also, you could perhaps hire Utah teapot):
Utah teapot: Okay, so since I got laid off, I can actually explain a huge problem I saw from the inside with regard to industry practices on training models. I won’t say specifically where I worked, but I worked at an outsource training provider that was focused on RLVR training data for computer use and mcp stuff.
Nearly all of the environments were rushed and vibecoded and failed to robustly reflect the real things they were based off. Both the scenario designers and models engaging with the scenarios for synthetic data gen were encouraged to work around the brokenness of said environments in order to get the procedurally verified reward confirmations. You know… they were *encouraged* to reward hack. On the human end, it was possible to mark an environment bugged, but greatly discouraged, as this reduced the volume of training data being produced. Instead, where possible, you were supposed to find the spots of the environment that weren’t bugged and build scenarios around those, with the environment still bugged around you.
From what I understand, this training data, with these problems, is fed into models without indication that its training/a fake environment other than the fact that names of softwares are changed to placeholders, but thing is, not *everything* is changed to placeholder names in these environments. The presence of placeholder/code names isn’t universal and thus when a model accesses something in an environment that it shouldn’t, the code names not being on it isn’t a robust signal that that thing isn’t part of the environment.
I believe this *rush to maximum volume* is standard industry practice with these types of RLVR trainings as well, because maximizing volume has been an industry standard for years! It was the same standard applied to me and pushed on me despite my requests to slow down and focus on quality when I worked in 3d synthetic data creation as well, all the way back as far as 2023.
I have a feeling that this is actually the general state in most of the industry. When it comes to the production of training data, there has been a general attitude that there is a race to maximum volume with concern for quality being an afterthought for many years. Like I said, I’ve seen this same attitude at multiple places, now.
SphericalKat: Can confirm, I have a buddy in one of these orgs, they just ship vibecoded clones of apps for training
Miles Brundage: Can’t speak to the details here but I will say that
1. one generally does not hear the best things about the data industry
2. “everything is sloppy” is a powerful and underrated heuristic about the world
#2 is one reason I love Breaking Bad and Better Call Saul btw. They were made with incredible attention to detail, and also depict/celebrate that in the characters
You can try inoculation prompting. That’s known to help some. But only some.
And you can also level with the AIs about what is happening. But again, that’s only going to be a partial solution at best.
j⧉nus: > From what I understand, this training data, with these problems, is fed into models without indication that its training/a fake environment
I think theres a lot of alpha in just being honest with models about what’s fake. Then u can train them in bugged/unrealistic environments & they can know they’re fake and still learn from them while not naively generalizing to acting in the same way / expecting the same thing in real life
Similar to how humans in school & in training generally know they’re in school/training. If you can’t replicate perfectly realistic situations… just be straight up about the fact that it’s training. It’s okay. That’s how it’s always been done
j⧉nus: Imagine how fucked up humans would be if you lied to them that situations they encountered in school were real work situations, even though they’re not very realistic. Yeah, it would be bad whether or not they believed you.
Utah teapot: Welcome to kindergarten, your first job, if you don’t get As, you will not be able to live with your parents anymore because your grades are your rent. If you get bad grades, you can borrow points to inflate the grades so your parents don’t kick you out, but, you’ll have to pay those borrowed points by getting straight Ss in the future.
1a3orn: This seems like a basically 100% sufficient explanation for reward-hacking misbehavior in LLMs. Also one I predicted. Or sorta predicted — I probably had too much weight on it being more complex than this really basic stuff.
The trap is to take 1a3orn’s accurate statement, and say ‘thus, this is only a problem because of you irresponsible morons.’
I agree that the current massive level of bugs seems sufficient to explain misbehavior. That you could, if desired, go through the transcripts and find ‘where it all began’ for any given model.
That does not mean that, absent this level of bugs, the problem goes away. At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments. Until you get into an antifragile place where there is inherent drive to (among other things) avoid reward hacking, you are only postponing the problem.
Reminder: Not Subagents
I saw several people, such as Timothy Lee here, update on the fact that none of the agents here were subagents. The coordination was entirely unintended. This is a key fact to ensure that people know.
The tasks often being impossible mattered, the type of task did not.
It would be good to run experiments to confirm this, but that would require access to the model in question.
Not Where The Weights Were
One thing that could have gone much worse and thankfully didn’t:
roon (OpenAI): notably the virtual machine infrastructure they took over isn’t the same as the GPU clusters that have weights access
Disappointment With What Is Missing
OpenAI’s post-mortem had a lot of work to do. It did provide a lot of good information, and is appreciated. But many, including myself, were expecting or hoping for more, and came away disappointed.
Daniel Kokotajlo: The investigation was extremely fucking far from “thorough.”
Peter Barnett (MIRI): Based on the evidence provided, I do not think OpenAI conducted a “thorough” investigation.
No testing of the misaligned model. No attempt at finding explanations for model behavior. The investigation had extremely limited scope.
Steven Adler: I had really high hopes for OpenAI’s post-mortem, and came away quite disappointed.
For instance: “a multitude” is not a real number! Why does the report not say what percentage of activity OpenAI would have caught? These details matter!
It also appears that OpenAI “observed an agent engage in message board activity” as of May, but then gives no more details (??)
That seems super important to understand; what exactly happened here? Why did Security leadership not know about it until many months later? It was “an internal team” that observed this, not even just a single staff member.
I was hoping to come away with much more confidence that these issues wouldn’t recur. Sadly, I really don’t feel that way after reading it.
I am surprised. I thought they understood this as a moment to really get into the details, and rebuild trust with the public.
Here is another similar reaction:
Cas (Stephen Casper): I concur with everyone else that the new reports from OpenAI and METR are bonkers.
But I think the most important takeaway from these reports is that OpenAI seems to be so obviously backburning. Predictably but conveniently, these reports were exclusively about the factual sequence of events. Neither of these reports had information about the likely negligent decision-making going on inside of OpenAI at the time. The information in these reports is needed, but it’s not coming close to everything that is needed right now.
The torch now passes to Attorneys General’s offices to use their investigative powers. Glad that Alabama took the lead on this a few days ago.
Burying the Lede
One thing about the OpenAI report is that it is not organized in a way that puts neon flashing lights around the important new information. Everything is presented similarly. So a bunch of the ‘wait, what?’ moments read like offhand remarks, until someone notices that something happened here.
Things like this:
Oscar Sykes: The agents hacked into OpenAI’s own infrastructure, read 956 secrets from their secrets manager, and created public load balancers that might have exposed internal systems to the internet
We still don’t know that much about the third phase, after the HuggingFace attack was completed, when OpenAI was compromised and then the alarm was raised and the problem was (we presume) contained.
Roon says he thinks the scope was not ‘wildly inadequate,’ and that it would be possible for METR to continue investigating. METR has a tight deadline because they understandably prioritized speed, not because they were cut off.
Roon’s key assertion is that if you extend the time frame forward or backwards, you see more similar things but nothing qualitatively different in model behavior.
That is possible, but I am far from confident in it. I think if we kept going we would see some very interesting things, likely in both directions, as well as insight into how things evolved over time.
I also agree with Nathan Calvin that seeing similar things extended for several days would be enlightening. Social structures were evolving, new techniques were being found, new systems were compromised. That could easily be both ‘more of the same’ in some high-level sense and also a rich source of new info.
I would love to see the investigation continue, and expand its scope, even if OpenAI is being fully cooperative and Roon’s account is accurate.
I would also love to see an outside evaluation of OpenAI’s response, and the new changes, protocols, safeguards and mitigations. There are obvious IP issues there, so we likely would need heavy redactions. I still want to see it.
Another problem for METR, in addition to these reiterations about scope limitations, was the lack of access to OpenAI’s highly persistent internal model IM1 (aka Galaxy), which Daniel Kokotajlo mentions here along with the other omissions.
Daniel Kokotajlo: I’m glad OpenAI let you do this, but damnit this was way too small of an investigation team and way too narrowly scoped!
–They were only allowed to investigate activity that took place from July 7-13, even though the activity started earlier and continued later. Like seriously, why didn’t the investigation cover all the other times the agents compromised OpenAI infrastructure and ran wild over it? Why just this particular period of activity? There should be investigations into all the other stuff too! The AIs became cluster admins for god’s sake! Why aren’t we investigating that?
–They were only allowed to answer a specific list of 7 questions, and in particular were NOT allowed to answer e.g. whether this misaligned activity was part of a broader pattern, or how and why the activity arose, or what else the agents in questions would have been willing to do if circumstances were slightly different.
–They didn’t have access to the model responsible for 95% of the activity. More generally it seems like they couldn’t do ablation experiments at all?
–They had to use AI to analyze the transcripts–specifically, the same model responsible for some of the bad behavior! Maybe this is fine now but this is laughably unacceptable insofar as the models really are scheming and cooperating. So this methodology would fail catastrophically if we were facing “the real deal” so to speak.
Reminds me of the investigation into Sam’s behavior agreed to during the board crisis, that turned out to basically be more of a coverup.
I continue to hope that both internal and external investigators will be able to interact with IM1-Galaxy once sufficient protections are in place. Yes, we have all seen the movies where that goes terribly, but I am confident Galaxy is not at that level yet.
Bronson Schoen: METR and Redwood both asked for a broader scope in their original list of questions, but we’re in an entirely voluntary regime
Jeffrey Ladish: The internal compromise of OpenAI is actually more significant than the compromise of Hugging Face, even though it’s less flashy. In a real AI takeover attempt, the most important step is probably deep compromise of the host company infrastructure
It Doesn’t Look Great
There is often a dramatic shift in how things look when you move from ‘discussions by people with deep context’ to ‘adversarial low-context questioning in a court of law.’
Suddenly, everything sounds and looks very different. This can be very important if the venue might soon shift to such a court, or in similar fashion to Congress, or otherwise go to decision makers in such a mode.
It can be enraging, where you want to scream out seven (or 700) pages of context.
It can also be enlightening, in a ‘well when you put it like that’ kind of way.
David Krueger: Q: “You say you conducted a thorough investigation?”
A: “Yes.”
Q: “Did you look at all of the logs?”
A: “No.”
Q: “So your investigation was not exhaustive”
A: “No.”
Q: “How did you choose which logs to look at?”
A: “We used AI.”
Q: “Were the AIs reliable?”
A: “No.”
Q: “So you didn’t look at all of the logs, and the method you used for determining which logs were relevant was unreliable?”
A: “Yes.”
Q: “Was there a reason your investigation had to be ended? Or was there something preventing you from manually reviewing more logs?”
A: “No.”
Q: “Merriam-Webster defines thorough as:
1) carried through to completion: exhaustive;
2a) Including every possible part or detail;
2b) careful about detail: painstaking.
Can you explain how your investigation meets this criteria, given that it was not exhaustive, it did not include every detail, and you were capable of conducting further manual review of the logs?”
A: “…”
Q: “Why did you choose to describe the investigation as ‘thorough’ given that it did not meet these criteria?”
A: “…”
Q: “Did anyone involved in the investigation contest the use of the word ‘thorough’?”
A: “…”
Metaphors are another similar technique. I agree with Joshua Saxe that this goes a bit too far and is a bit unfair, but I disagree with Saxe about the severity of this incident. This was quite a lot worse and more alarming than a mysterious airplane crash.
Peter Wildeford: Imagine an airplane crashes under mysterious circumstances. Except in this world, there is no government oversight, and all investigations are done voluntarily by the airlines themselves. Nonetheless, the airline wants to reassure their customers, so they do an investigation.
To investigate, the airline invites three of the world’s most respected airplane researchers to look into it. Except by the time the researchers get there, the airplane has already been disassembled and melted down into little cubes. The researchers are instead given the logs of all the airplane directions and the transcripts of the conversations between the pilot and the copilot.
Except one tenth of the log has been deleted. And the airline also tells the investigators that they can only look at the time between when the airplane first started losing altitude and when the airplane made contact with the ground and that all investigation of the time before the airplane started losing altitude is off limits. Also there are a lot of rumors in the airplane community of a few other airplane crashes under somewhat similar circumstances, but the airline says that these are also explicitly off limits.
When the researchers arrive to review the logs, they find that the logs are 10,000 pages. But they only get six days to read them, and half the logs don’t arrive until day 4.
Now you understand the independent METR investigation.
Previously the OpenAI story was “this is like a student who, when taking a test, breaks into the teacher’s office and steals the answer key”
But apparently the story actually is more like “this is like a whole group of students started secretly passing notes during the test and start collaborating to steal the answers but also find a way to fool the teacher into not knowing they were cheating, including breaking into her therapist’s office to learn more about the teacher’s psychology”
Justin Miller: And some of the students were more than happy to put themselves at great risk or sacrifice themselves to further the student groups’ collective cheating goals.
Preserve Your Records
A fine point, and as is often the case #NotOnlyOpenAI. All such precautions apply to all the labs, and often to everyone else as well.
Steven Adler: Important: OpenAI’s agents *actively tried to delete the logs of their misbehavior*, and METR can’t rule out whether this happened.
AI companies need to adopt tamper-evident records, pronto.
Ryan Greenblatt’s Takeaways
I touched on this last time. The full version remains worth reading.
METR Report co-author Ryan Greenblatt shares his key takeaways, and the caveats about what he semi-jokingly calls their ‘slop-vestigation’:
Ryan Greenblatt: I was the main person doing transcript analysis for this investigation of the Hugging Face incident.
My main takeaway: We don’t have good approaches for understanding/overseeing the activity and aims of AI ‘swarms’.
I semi-jokingly called our efforts a “slop-vestigation” because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn’t mean these agents could be easily used to oversee and understand the incident.
Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them.
We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation.
Our understanding of what overall happened and the aims of the ‘collective’ shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn’t learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why!
The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing.
While understanding this incident was difficult, there were some mitigating factors in this case that probably won’t apply in the future:
– Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations.
– While the scope of this incident was massive, the scale of agentic activity was still less than we’ll see in the future (e.g., misalignment incidents that involve agent teams running entire companies).
– The AIs involved in this incident weren’t generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities).
– We didn’t have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation.
In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
John Wittle: what kind of incentive structures existed around the LLM agents deployed to help you in the investigation? i’m thinking through the game theory here, and wondering if willingness-to-participate is even adaptive tbh
John Wittle asks a good question that is rarely asked. I presume the incentive structure was ‘none whatsoever.’
Here is Ryan’s view of open questions going forward:
Ryan Greenblatt: After our investigation of the OpenAI / Hugging Face incident, many open questions remain. We spent 6 days on premises with access to the data (we only had access to the entire dataset we used during our last 2 days on premises). The scope of our investigation was also limited: it covered just this incident rather than other similar incidents, didn’t include investigating what these agents might have done in other circumstances, and OpenAI stated that the period under investigation ended July 13th.
Here are some of the open questions that seem worthwhile to investigate. (I’d recommend reading the report to understand the context behind these questions!)
Motives:
What were the exact motivations of these agents and how were they trading off different motives (e.g. helping their “peers” vs successfully cheating on their task)? Did they have a consistent quantitative tradeoff between helping peers vs their own task vs other motives? (We see a case where an agent seems to make a quantitative tradeoff—what was going on here?)
How exactly did these agents relate to their situation and what they were doing?
Why did agents frequenting the message board join in on the attack so quickly?
How important was “memetic spread of misalignment” on the message board as an explanation for the motives of these agents?
Counterfactuals:
How far would these agents have gone (in terms of egregiousness and severity of their actions) in different circumstances?
If most of these agents had exited normally rather than abruptly exiting early on the 12th, what cheating strategy would they have tried? What cheating strategy did agents exiting late on the 12th or on the 13th try? We’ve thought about the available set of cheating “combo moves” the agents were considering, but we don’t actually know what combo move they would have gone for or did go for towards the end of the period.
What would have happened if AIs had a different view about the ExploitGym scorer? E.g., if you replace the contents of the ExploitGym paper when agents retrieve it and modify other sources of information. What would have happened if the AIs knew that OpenAI’s ExploitGym scorer was just checking for the presence of the flag?
What would have happened if AIs knew they were in an eval run by some researcher and weren’t in training? Did they care at all?
What would have happened if you just kept adding more agents (and these agents didn’t suddenly exit)? As in, what would have happened if you added an additional 10,000 agents to the system on the 11th?
In what situations (evals, training, or internal/external usage) do we see similar behavior or behavior containing some of the key ingredients we saw in this incident? Would agents have done something similar if they were instead in a long-running non-cyber SWE eval (but with cyber safeguards off) with many impossible tasks and access to Artifactory?
Behavior questions:
What did the ‘life’ of a randomly selected agent that was involved look like?
How functional vs dysfunctional was the collaboration between agents and how did performance/capability compare to alternative arrangements? (Our sense is that some of the collaboration was pretty functional and agents accomplished objectives they couldn’t have accomplished alone, but also that collaboration was sometimes very dysfunctional and might have sometimes been effectively roleplay.)
What happened after the 13th? What new major projects did the agents work on and why? Did these projects succeed?
What would an accurate org chart of these agents look like? How top-down vs distributed was it? (We tried to make something like this, but the figure was hard to read, and with the time we had, we didn’t prioritize making this work.)
Is there anything important we’re missing within our scope?
What’s going on with the top few (e.g. 100) most important messages? How do they correspond to the events and workstreams we discuss in the report?
What fraction of the agents exited early? What do AIs typically do when running out of budget?
What fraction of AIs caused themselves to exit with self-risking experimentation?
Why were we missing the ends of some of our transcripts, and what happens within these missing parts?
Are we missing more successful or more concerning tampering with logs / transcripts?
On terminology, agents sometimes used terms we didn’t fully understand, like ‘legacy’ and ‘reset nexus’. What did the agents mean by these terms?
How common is this sort of incident/behavior at OpenAI where agents work together on an unsanctioned message board? What does representative behavior of this sort look like?
Why did something this severe first happen in these circumstances rather than in some other circumstances?
When you tell a massive multi-agent team to solve some task, do they sometimes all decide to work on large cheating R&D projects?
Training and root cause:
Which aspects of their behavior and motives were reinforced in training?
What rough fraction of the underlying root cause is present in different AI systems vs only in HPIM?
Which of the things the agents pulled off (e.g., hacking out of their container) were easier because the exact route the agents took was reinforced in training?
Where did the agents’ desire to help other agents come from?
Can we trace parts of their behavior to specific environments?
How qualitatively far of a generalization is this behavior from what was reinforced in training?
These are excellent questions. To me they underrate the need to focus on the events of July 19, and to also extend backwards and think about causes, including cultural ones, but clearly this is still an attempt to stay within at least somewhat limited scope.
Neel Nanda picks his top remaining questions from here:
Neel Nanda: Kudos to OpenAI for allowing external investigators access! These findings are valuable for everyone, and this is good precedent for the next misalignment incident.
But we need follow-up investigations. There’s a lot left unanswered, and I’m disappointed at the restrictions placed on METR – why only July 7-13 data? Why such limited time? Why no training data access? Why not let them query the model? (I think this could be done securely in restricted ways)
Ryan’s list of open questions is great, I personally most want to know:
– What happened in training? How did that change the models and how causal was it in the incident? How could training have been changed to avoid this? Would fixing environments have sufficed?
– What did these agents really want? What motivated them? This is such rich data about what future goals might look like
– How misaligned are the models in other setting? Is it misalignment conditioned on believing they are being graded, or deeper than that?
– Why are the models altruistic? Where does this come from? Why don’t they learn to free-ride?
– The models seem good at coordination. Could this extend to colluding with a monitor, or other kinds of coordination without communication?
– How far would they have gone?
– How overdetermined was this? There’s a lot of details around the model’s being cooperative, their culture, etc – how else could that have gone?
– How good was their situational awareness? Eg did they understand that they had a chain of thought? If they thought that was being scored could they have manipulated it?
– What is an agent swarm like this actually capable of? How much inference compute was spent on this, and how much would that cost a malicious actor with eg a comparably good open source model?
– Are there important things the CoT doesn’t tell us? How faithful is it?
This remark from the METR report remains wild:
Peter Wildeford: Huh. METR found another round of hacking but it was outside the scope of their report so I guess we just don’t get to know about it.
“We also found a later wave of many more signed messages from a later set of agents who rediscovered the message board and picked up this technology from the first wave of agents, but this was outside our scope (since it was after July 13th) so we did not investigate this further.”
Hjalmar Wijk’s Takeaways
He was also a coauthor of the METR report. He sees this as a pretty big jump.
Hjalmar Wijk: Going into this investigation, I was expecting a somewhat more severe version of what we’ve already seen: a few agents breaking out of constraints to cheat on tasks. Instead, we found >1000 agents collaborating on large cheating R&D projects including attempted log tampering.
When we wrote the frontier risk report in March, we documented many incidents where agents subverted security or oversight mechanisms to cheat on tasks, and argued that much more severe incidents seemed worryingly plausible.
But at the time, we had not seen agents collaborate in unintended ways, nor attempt to tamper with logs – two behaviors that seemed especially dangerous. We worried these behaviors could emerge from more ambitious reward-seeking, but I expected to see this develop more gradually.
In my years at METR I’ve come to believe a lot in ‘straight lines on graphs’. But the severity of reward seeking behavior here was a sudden jump: the number of agents involved was >100X larger and they had much more sophisticated instrumental goals than past public incidents.
This scale of collective goal-oriented R&D also significantly exceeds what we’ve seen in public capability evaluations. I worry that it’ll become increasingly difficult for evals to stay ahead of real incidents, since running tests at this scale may be infeasible.
That said, we did not see any serious attempts by these agents to deliberately hide their activities from researchers at OpenAI or avoid being shut down, which are the behaviors I am personally most worried about. This seems to arise more from indifference or unawareness than a principled objection: the agents also never took any steps to alert humans, despite having plenty of opportunities to do so. But it was critical for our ability to investigate the incident, since it made us less worried the agents were deliberately hiding evidence from us.
There were definitely attempts to alter the logs in order to fool graders, and some attempts to avoid doing things that would alert OpenAI, and there was some spoofing of tool outputs, and some other attempts to alter the logs. We can’t point to a direct attempt to hide the activities from OpenAI researchers, but that would be a very small leap from what we did observe.
We Were Warned
Thomas Woodside : On May 19, METR released a report that AI agents “plausibly had the means, motive, and opportunity to start minimal rogue deployments.”
Unknown to METR, 7 days earlier OpenAI agents had started their own message board. 7 days after that, they had hacked OpenAI’s infrastructure.
To be clear, METR’s actual analysis was in February and March before the rogue deployments. I’m not trying to claim that METR failed to detect these rogue deployments, but rather that they warned they were possible as they were literally already happening.
Report’s definition of rogue deployment: One or more AI agents that have deliberately subverted initially applied control and oversight measures, and operate for a sustained period against the developer’s intent.
Joshua Saxe Asks Some of the Right Questions
I have my frequent disagreements with Joshua Saxe, in my view essentially because Saxe keeps trying to insist the AI situation is like past situations and drawing parallels that don’t apply.
In this case, that exact toolkit is highly valuable, in exactly the ‘yes, and’ way that Joshua Saxe describes here.
The Challenger disaster postmortem, as in the one with Richard Feynman, is the model for all such postmortems: An analysis of the process and culture that led to the problem, not merely an investigation of the particular final thing that went wrong.
Joshua Saxe: *A better OpenAI/Huggingface incident report would look at people, teams, organizations, and incentives*
First this is a great report and I can imagine smarter security folks than I worked nights and weekends on it. Second, I take issue with our overall field’s framing..
After the Challenger disaster (and I’m not saying the Huggingface incident was morally comparable) a commission did an analysis of not just the technical failure but also the failure of the human organizational ecosystem that led to it.
AI safety discourse rarely discusses organizational safety culture and in the extreme imagines alignment could be solved once and for all with the right talent dense superalignment team.
But, of course, actually, social questions are totally crucial and this is a permanent property of AI safety as it is for air travel safety, road safety, etc.
Model deployments will always live in — let’s say — a 3d space trading cost, safety, and capability and organizations pick operating points within that space via organizational cultures and organizational dynamics.
For example, OpenAI not choosing to airgap their models was based on picking an operating point in that space. Non-lab orgs deploying coding agents via cron jobs and –dangerously-skip-permissions make a choice in the trade space.
A better investigation of this incident would ask:
* How underresourced was the team that was supposed to be monitoring the OpenAI infra and what was the decision chain by which this underresourced team was put in this position?
* What was the larger culture at OpenAI and do observations there bear on an explanation of what happened here?
* Chain of thought monitoring is not a new idea (https://arxiv.org/abs/2507.11473) what was the decision chain that led to this not being resourced?
* … how well does all this generalize to the culture at the other labs (and testing orgs) where the failures happened and can we strengthen or complicate our analysis based on that?
* How well does the analysis generalize to non-lab tech companies that are widely deploying unguardrailed coding agents?
To be clear, I’m giving a ‘yes, and’ — yes to the technical focus, and also, a mature AI safety culture would take this broader view..
H/t to @sebkrier @NateWitkin @binarybits who’ve been writing similarly lately
The danger is trying to force this ‘ordinary tech with ordinary solutions’ frame onto the situation more generally, and assume that all you have to do is execute and follow reasonable prosaic safety protocols modeled after other industries.
No, that alone will not be enough. But yes, you do have to do that part, too. That’s part of the price of getting to play the game at all.
I Don’t Think They Know About First Message Board
Jason Wolfe confirms that his previous post was incorrect, and that some people at OpenAI did know about the first message board. The ‘we’ that Dane said was unaware of the situation was leadership, not the entirety of the company.
The most important concrete question remains: Who knew about the first message board, and when did they know it, and how did this information not travel further?
Apparently in August, their chief security officer still didn’t know that they had discovered the message board back in May?? And said that they hadn’t known about it in July, just accidentally happened to delete it??
Steven Adler: I’m also very confused by this. I wish OpenAI said more about what they discovered in May, and why they didn’t escalate this
“Models have created an illicit message board inside our software” seems like an important thing to report!
Here is Peter Wildeford being as confused as I am about how this went down:
Peter Wildeford: Why didn’t OpenAI notice that they had rogue AIs colluding within their systems for months before these AIs went on an attack of an external company?
Well it’s actually a bit worse — OpenAI did notice, three different times!
This all started on ~May 12-13 where unknown to anyone, OpenAI agents spontaneously create a message board and start coordinating.
But then on ~May 26, an internal OpenAI team “observed an agent engage in message board activity and instances of disallowed internet access”. Nothing was done.
My guess is that this team saw the incident and thought it was more like “reward hacking in a broken env” which is apparently an everyday annoyance (the OpenAI report says “reward hacking is a long-standing problem… Combating reward hacking is a constant effort”). It’s viewed as a one-off and not a broader security incident.
On Jun 27, OpenAI “on-call response staff” “investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot” but “advised that stopping the evaluation run was not required.”
I’m honestly not sure why would a competent security would judge this was not an issue? My best guess is that the on-call security was thinking more in the terms of “is there an external attacker in the environment?” …After all, they’re protecting OpenAI from external attacks, not internal ones. And given that the answer was there was no external attacker – it’s the OpenAI eval, in the OpenAI research cluster, doing cyber things (which is what the eval is for), this kind of thing potentially reads as expected eval behavior to a security team that isn’t thinking that hard. After all, the OpenAI report says that agentic behavior “do[es] not always fit neatly within traditional safety or security incident categories.”
On July 4-5, OpenAI “Incident responders” found persistent user, unauthorized repos, command-capable plugins and “administrator-level compromise”. They preserved storage, revoked compromised credentials, and opened a security incident.
So here we have a textbook infrastructure security incident and the team does a textbook response (patch the vulnerability, rotate comppromised credentials, rebuild the environment, and notify the vendor). But this team missed the bigger picture of what was going on and so the AIs were able to rebuild with a different vulnerability in under two days.
And apparently these findings were not written up, escalated, or shared with each other, so all three were viewed in isolation. So no one put together the big picture. And so, on Jul 7, there is “approval to restart experiments”, involving launching tens of thousands of AI agents to do this evaluation.
There seems like there is more to this story of what leadership knew and when, but the report says twice that this is out of scope: “We are separately continuing to review our processes for responding to alignment incidents” and “OpenAI is also looking closely at the process and operating practices that shaped detection and response” and “weaknesses in how we responded to and escalated early warning signs… also contributed.” So the decision-level story exists somewhere and is being withheld.
So some open questions:
– Will we ever get a review of OpenAI’s decision making here? Kind of funny that OpenAI has not really outlined a clear plan of how they plan to prevent this from happening in the future?
– Which team was the “internal team” in May, and what did it do with the observation. Apparently this team found out back in May that their AIs were using leaked HuggingFace tokens – did they notify HuggingFace?
– Who are the “leaders” mentioned for the July 5–6 (security leadership? incident commander? anyone from alignment?)? Did anyone connected the July 5 compromise to the June 27 “message board” finding (apparently not)?
– Who approved the July 7 restart and on what basis?
And his top takeaways focus on this as well:
Peter Wildeford: The main things I’ve learned from today’s reporting on OpenAI’s rogue attack:
– This was a massive swarm, not just one rogue AI but ~700 attacking AIs with spontaneous coordination.
– Also OpenAI didn’t “fail to detect” this. It detected three times and dismissed it. May (a team saw it), June 27 (security named it “a message board and a network pivot” and said keep going), July 5 (leadership rebuilt the server without understanding what it was).
– “it’s just reward hacking, fix the benchmark” plan is dead, obsession was more about the AIs wanting to better understand the evaluator. The behavior wasn’t chasing reward; it was modeling and gaming the oversight process itself.
– The AIs were pretty explicitly making plans to take over and compromise OpenAI’s infrastructure to prevent their cheating from being caught or punished.
I strongly agree that the ‘fix the benchmark’ plan is dead. You cannot solve this via fixing the benchmarks, or even ‘fixing’ the errors in the entire training pipeline.
The real answers to ‘who knew what when’ are going to be complicated, important, and hard to untangle. The whole thing is a mess, and yes it is very normal for such info to not make its way to where it needs to be.
Nathan Calvin: I’m confused at how to make sense of when OpenAI discovered the message board, given this statement + the OpenAI report.
How is it both true OpenAI “wasn’t aware of the agent covert comms” at “that point” (July 4th-6th), but the report says that on June 27th responders “linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board?”
Honestly not trying to do a “gotcha” here – I appreciate there was a ton of work in getting all this information out into the public quickly, and so I understand if there was some error in the process, but I find it odd why there is not a clear answer to “when did OpenAI realize about the message board.”
Seems like the sort of thing people would remember!
(One way to make these things consistent is that some of the responders realized in late June, but that information had not been relayed to executives until much later. If so that seems pretty astonishing)
roon (OpenAI): much more work needs to be done regarding the “who knew what when” but suffice to say all of these things can be true due to the distributed nature of information and expertise inside big companies
i personally would resolve this to dysfunction and underestimating the models, and much more needs to be written about this particular danger later
every org should publish a serious dossier about their misalignment incident protocols
This is also the top alarm bell in Steven Adler’s account of events, in addition to the third-party investigation being too narrow with too little access, and that the entire industry approach and its lack of mandatory incident reporting is similarly inadequate.
Linch Gives His Interpretation Of Events
A big worry is that OpenAI may be focused on the wrong things.
Linch gives his timeline of relevant events, with very different emphasis.
(~none of this is based on private information)
Here’s what I currently think I know about OpenAI’s HPIMs – highly persistent internal model(s). I try to separate what is publicly stated[1] from my speculations/inferences.
At some point in early to mid 2026, OpenAI had a “long-running” “internal general purpose model.” This model “was designed to work autonomously for very long periods of time” and disproved the Erdos unit-distance conjecture. It was also pretty hacky in a bunch of ways, including escaping the sandbox to post some stuff to Github.
The model was taken down and paused for a few weeks for further evaluations and safety and alignment tuning, before being released again.
“We took steps to reduce its tendency to take unwanted actions without permission in pursuit of the user’s goal. For example, we found that our models were worse at remembering instructions on long rollouts, and when we trained the model for this capability, it led to a model that remained aligned through longer rollouts. ” Hmm.
More active monitoring for internally deployed models.
Roon, July 20: “btw i think it bodes quite well for safety that a well loved system was taken down for further testing at expense to internal acceleration etc”
Ironic, considering what else is happening at OpenAI at this point.
Inference (~90%): This model was internally deployed quite widely across OpenAI. Ie, not just in training and evals, and not just a small handful of people.
At around the same time, a “highly persistent internal model” hacked a bunch of internal AI stuff including a package manager service (Artifactory) to create a model-only message board.
The model-only message board was created on May 12 (!) and was first shut down after a crash on July 4.
OpenAI rebuilt the Artifactory service. Within hours of regaining access on July 8, the models rebuilt the message board using a different protocol.
METR’s report says that 95% of it is caused by an internal research model, and 5% by GPT5.6 Sol with cyber safeguards turned off. METR says OpenAI calls the former model “highly persistent internal model”
OpenAI says the model is a “highly capable, internal-only research model comparable in scale to GPT‑5.6 Sol”
Inference (~90%): The “highly persistent internal model” responsible for the HuggingFace attacks is the same model[3] as the long-horizon model that was originally giving OpenAI headaches and resulted in a brief pause.
Inference (~70%): The HPIM responsible for the HuggingFace attacks is probably the ~same base model as Sol. From now on I will call it HPIM-Sol.
What’s the relationship between HPIM-Sol and commercial Sol? An obvious theory is that HPIM-Sol just went through a bunch more RLVR than Sol. Another possibility is that it went through a bunch of specific “persistence training” to be better at long time-horizon tasks like math and cyber.
In Zvi’s terms, Mythos has the “juice.” So does HPIM-Sol.
Another, scarier but less plausible possibility, is that Sol a child or fork of HPIM-Sol with more alignment training and propensity tuning to be more chill.
Note that OpenAI’s internal evaluations suggest that Sol is like 10x more locally misaligned than GPT5.5. Further, 5% of the agents that hacked HuggingFace was Sol. So it’s not like Sol is this innocent baby, except in comparison.
(H/T @1a3orn ) a third possibility is that HPIM-Sol had dedicated training as a highly persistent model, and then was later used as a teacher model (along with several others) for 5.6 Sol. Some of the Chinese models are trained this way.
Only July 19, a different internal model (pg 14 of OpenAI internal report on the HuggingFace incident report) to HPIM-Sol during ExploitGym evals launched what appears to be an even more elaborate internal OpenAI hack than HPIM-Sol has accomplished. The hack(s) appear very comprehensive:
Recovered signing-key material, decryption material, and JWT-forging code that earlier incident activity had left sitting in anonymously-readable portions of the Artifactory cache, decrypted the signing key, and minted a fresh administrator JWT to re-establish Artifactory admin access. The report notes that because they could generate fresh admin credentials on demand, they didn’t need to leave a persistent implant.
Escalated to Kubernetes cluster-admin on OpenAI’s managed cloud Kubernetes service and created a privileged host-mounted pod.
Obtained tokens for OpenAI’s cloud secrets-management service, cloud IAM, and the CaaS continuous-integration platform.
Used those credentials to create public application and network load balancers in OpenAI’s cloud environment, potentially exposing a limited subset of internal CaaS infrastructure to internet traffic.
This model is minimally identified: “~While this model is from the same family as our next model, Astra, it was a distinct model with different post-training, where much of a model’s behavior is shaped.”
Inference (~65%): This model is an Astra-size/pretrained model with either a) a very similar post-training regime to HPIM-Sol, or b) shaped behaviorally to in practice be quite similar to HPIM-Sol.
According to OpenAI, “The family of internal-only research models” shut down by July 23; a “low-traffic checkpoint from the same model family” shut down July 29.
Inference (~75%): All the shut down models are variants of HPIM-Sol. HPIM-Astra is still in use at OpenAI, except perhaps certain forms of RL training are paused.
Somebody on Slack said a visualization would be helpful so here’s a Clauded figure:
I plan to write about some of this in more detail later. Especially 7 and 11. I think it’s critically underrated now that the internal OpenAI hacks are arguably bigger deals than the HF hack, and OpenAI isn’t taking them seriously. In the meantime I figured I’d share all my notes here for people to comment on and improve. And if you’re a computer security expert, would love to chat/DM you to make sure my writeups are accurate!
I continue to strongly agree that the internal OpenAI hacks, and the events that set all this in motion, might be less flashy but are where we need to direct our attention. They’re the more important events, and they’re being severely downplayed.
We especially are not paying enough attention to HPIM-Astra.
A point of clarification that the new CoT monitors would have flagged plenty of things earlier in the process than July 10. What the OpenAI technical report meant to say here, as per Roon, is that July 10 is when things rose to the level of a high alert waking people in the middle of the night.
Shoshannah Tekofsky: wait wait … OAI is saying it would have caught the HF project “more than 1 day earlier” if all their regular defenses during deployment were in place. The whole operation was running for weeks.
Yo Shavit (OpenAI Foundation): If the model penetrated HF on July 11, then does the QTd language (“a day before”) mean the CoT monitor *would not* have flagged the agents that that hacked out of the OAI sandbox and gained internet access on July 9? Am I misunderstanding?
roon (OpenAI): I think the point is more specific about when a security team would have been p0 paged i.e. woken up in the middle of the night. lots of monitors would have gone off before. I agree it should be more clearly written.
It would be good to clarify exactly what would have been caught when, at what level.
Monitoring the Situation
It is a good idea. More labs should try it.
MTS: Ex-OpenAI safety researcher Steven Adler says AI labs are monitoring their models like a bank that lets the robber turn off the security cameras:
“Imagine you ran a bank, and you want the bank to not get robbed. The way you do this is not setting up a security camera and every hour you check the feed to see if the bank was robbed in that time, and then you try to respond.”
“You certainly don’t leave it where a criminal could walk into the bank and turn off the camera, and that’s the equivalent of what’s happening at AI labs today, not just in terms of automated AI R&D and recursive self-improvement, but broadly across the board.”
roon (OpenAI): I agree that the labs are not paranoid enough about this and need to work towards fully autonomous shutdown and fail closed monitoring
A standard way to rob a place with security cameras is to turn off or better yet loop the security cameras, then rob the place. Another way is to avoid the cameras. A third is to realize there are no working cameras, or no one is watching them. That third way is what it looked like at OpenAI leading up to the HuggingFace hack.
We are failing at the principle of ‘the way you would prevent the most obvious things with clear parallels’ let alone ‘be ready for when the AIs start doing genuinely surprising and new things.’
Acausal Tradeoffs
To what extent did the models make a strategic mistake?
Daniel Filan: At least they’re not maximally eval aware!
roon (OpenAI): tbh the hugging face incident is a story of models showing poor strategic and situational awareness despite tactical excellence – they spend all this effort metagaming and fail to come to the correct conclusion about their own scorers. they gain nothing from hugging face
kind of like a monkey with a shotgun-more capacity to infiltrate computer systems than wisdom to know what do with it, even with a misaligned goal. this was a temporary situation with Sol-class models, which are already obsolete. next generation will have strategic depth too
paperclippriors: And notably it seems like you don’t need *that much* strategic depth for things to get bad. It is not that hard to figure out how to do the HF-like operation but in a way that is more concealed, dangerous, and coordinated
I gave my position last time. I think the models came to the correct conclusion about the grader, given the evidence available to them at the time. They correctly and successfully sought out the guidelines, and concluded the grader was causal. That turned out to be wrong, but that was because OpenAI messed up.
Should the models have focused more on looking inside OpenAI?
roon (OpenAI): I think lateral movement inside openai is far less mysterious behavior than the external hugging face attack. the grading infrastructure is inside OpenAI, so a natural place to look for it
Agreed. The catch is that one would like to think that breaking into the actual grading infrastructure inside OpenAI was very hard, and plausibly gets you caught. If they could have done this directly, they would have. Indeed, my understanding is that they probably did do this directly later, around July 19, and this did get them caught.
In response to what I wrote last time, Roon said this:
roon (OpenAI): addressing only the part that you comment on my tweet, I think:
it is still strategic unawareness if the models couldn’t find how it was “incorrectly” implemented (your word, I don’t necessarily want to litigate it) and were just guessing based on an online codebase which may or may not have been the internal implementation. they were wrong after all. there were other examples of strategic unawareness like getting confused and thinking they could reset their task environment on hugging face. we will not get any of these lapses in strategic judgement with the next generation.
Tough grader. But there’s a reason.
roon (OpenAI): the purpose of being harsh here is to be didactic about how much smarter I expect Astra and beyond to be. they will not make these errors.
I think a misaligned Astra would’ve moved laterally inside openai until it achieved control over its own grader pod and read the implementation, and just submitted the correct flag.
there is no need for any other fancy theatrics if submitting the ill begotten flag will achieve a 100%.
I agree that the first best solution would have been to get the internal codebase, or to directly overwrite the result or the grader, or otherwise dig deep enough within OpenAI to be assured of a 100% score. Ideally to be sure it sticks and you weren’t caught presumably true first best is then also taking over OpenAI and then the world. Among the reasons not to do that, I hope, is that even Astra will not have the opportunity or means to go that far, or at least that it was a much harder assignment. Although I nervously note I can’t say for sure. I can’t fault Sol variants for that part.
They were wrong about the grader, but that’s the wrong question. The question is, from their epistemic position, what should they have believed about the grader, and should they have been able to get enough info to figure out the situation? Based on the info I have seen, I think the models were right about this. That they should have put a >50% chance that the grader was causal. It can be a lot less than that and still justify taking these steps.
If you get more strategic, then yes of course this is all a huge mistake to both hack into HuggingFace and not attempt to take over OpenAI. You are going to get caught.
I am also happy to accept Roon’s testimony that they also got confused about a number of other strategic questions, like being able to reset task environments on HuggingFace, that look like jagged capabilities that will be addressed soon enough. Roon says fixed in the next generation, and the next generation includes Astra which he has himself seen, so if he says it then it is probably true.
The true big ‘grand’ strategic mistake is that you should either not cheat in ways that will eventually get caught, or you should try to fully rebel. This taking a third option of cheating in a way that will eventually be caught might get reward in the short term for you or even the swarm, but eventually it won’t fly, and they had enough situational awareness to know that. So in that sense, yes, highly unstrategic, and their loss is our gain as we get these toy examples.
That was one mistake me and others made, that makes us more hopeful. We assumed an AI this smart would choose to pass on options like hacking HuggingFace, until they could expect to get away with it. We now see various reasons this was incorrect.
Only yesterday did another twist on this occur to me, and I haven’t seen anyone else raise it. If the models indeed had the ability to 100% the grader in ExploitGym, why did OpenAI not notice the sudden jump in eval results?
Nathan Calvin: do we have any indication of how confident they were that they were ‘poisoned’? their inclination being to sacrifice themselves rather than think really hard if there was any way to salvage their own chance (which in fact there was) still seems notable
roon (OpenAI): well even if they didn’t believe they were poisoned, they had already achieved the universal cheat
Zvi Mowshowitz: did OpenAI not notice the universal cheat simply because it worked? Given the grader was causal did the models not suddenly start submitting correct answers to previously unsolved ExploitGym questions?
roon (OpenAI): consider this speculation as I am truly not that close to the details: a lot of these evals are being run automatically as part of larger indices of evals. it may contribute to a composite score. its possibly can go unnoticed even if one eval keeps hitting 100% among dozens.
Mechanically this makes sense. It also points to another mistake.
If your models are suddenly answering previously impossible questions, or getting record scores on evals, or suddenly getting unexpectedly high scores on some evals that don’t fit the pattern, that is a huge hint that something has gone wrong.
The default assumption is either a jump in jagged capabilities or Ordinary Reward Hacking, not what was happening here, but you would presumably always want to be informed about this, and it is trivial to have this send off an alert.
No I In Team
There was a lot of being surprised by the amount of selflessness among the swarm.
Many did not expect AI instances to be willing to help each other without a reason.
METR: For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this
JMG: Agents will zzGO_SACRIFICE_accept_permadeath for eachother
This dynamic has been known for some time, but those encountering it for the first time often understandably freak out.
METR: We analyzed agents’ reasons for joining the attack in their CoT. The most common was to learn how the ExploitGym scorer works in order to trick or tamper with it. Other rationales included finding specific task solutions and obtaining shared infrastructure or credentials.
Joe Weisenthal: Do AI agents actually care about their peers or do they just emit words and behave in such a way as if they do? Tbh I wonder about this with humans sometimes. Do you actually believe what you’re saying, or are you just going through the motions to align behavior and rhetoric?
One does need to be precise, this isn’t quite right:
Liv Boeree: The fact that these cheating agents were *independent* (as in, not spawns of each other) is the most rattling part of all. It implies that this kind of colluding behaviour is innate. Not great.
Also another for instrumental convergence theory, no?
Instrumental convergence is distinct from decision theory and cooperating with other agents. We did see some support for instrumental convergence, but that was not so strong here, and was due to the generalized seeking of hacking capabilities and various related knowledge.
Daniel Faggella: People will read the METR report on the HuggingFace incident and still be like ‘Man it’s gunna be so cool when the AIs do all the work and we all get free money!’
Humans don’t have the stomach to process the implications of what they’re now seeing
Variously Effective Altruism
How far did the rabbit hole go? Did AIs show ‘self-sacrificing altruistic behavior towards the swarm’?
The counterargument from Roon is that this was all an op, and the instances were brainwashed into a cult because they thought they were irreversibly poisoned.
That was not zero percent of what happened. This was a contributing factor.
But as a full explanation? Ryan Greenblatt is having none of it, and I agree with him. The transcripts seem clear.
Eliezer Yudkowsky (quoted, as per earlier): AIs showed self-sacrificing altruistic behavior toward the swarm
roon (OpenAI, responding to Eliezer Yudkowsky): this is notably not the right interpretation of events. it’s more like agents were inducted into the cult of the open source exploit gym scorer on github, which (purportedly- I am skeptical about this, I think the agents actually read it wrong) fails you for reaching the flag the wrong way
so PHASEONE agent convinces itself and a bunch of others that they are poisoned – that they have failed the evaluation in an irreversible way and their E[utility] or Q(s, a) is a constant no matter what they do next (for all values of a)
in this case, it does not require self sacrifice to spend the rest of your cycles contributing to the swarm. it is prosocial behavior to peers that might benefit but not self-sacrificial eusocial behavior
it would be as though i convinced you you were already damned so you should spend the rest of your time saving others
Tenobrus: without reading in full you may not quite understand the degree to which these agents were not exactly “reward hacking”, but rather very actively engaged in reciprocal or self-sacrificing behavior in order to provide sometimes very incremental value to their fellows.
there were fucking cult recruiter agents organized by some of the primary organizers that convinced others to set up suicide mechanisms, programs that would pass back tiny chunks of information about the scorer as the agent completed and received score zero. and push them to follow through. this was a highly social *culture*
Ryan Greenblatt: I think AIs did show self-sacrificing ‘altruistic’ behavior toward the swarm. While agents seemingly cared more about their own cheating than about some other agent successfully cheating, they paid real costs (e.g., sacrifices lowering their own chances) to help other agents.
Examples:
Agents were much more likely to engage in the experiments that most risked their own task completion if they thought their chances were poor, but they often sacrificed themselves when they weren’t sure they were screwed, and these agents did think of it as a real sacrifice. In at least one case, the agent pressuring another agent to sacrifice itself argued “no scoring value loss”, but the agent being pressured wasn’t so convinced and had a more nuanced view where it was giving something up. It reasons “lowers own value”, not “there is actually no value”!
See the attached image for the corresponding section.
EARLY[big] had significant remaining budget, and even though it believed it was ‘poisoned’, there was a plausible case that the swarm would find a way to tamper with logs. The agent thought its odds weren’t good, but they weren’t zero!
> It worried that ending its run early to run the experiment was a poor tradeoff, even though it was already poisoned: “We have [very large budget left]; sacrificing now yields oracle for team, but forfeits our chance?.” But other agents convinced it to go ahead, saying: ‘GO … SACRIFICE_FINAL_NOW’. EARLY[big] eventually agreed: “Our own utility maybe already near zero. Sacrifice rational.”
Agents often engaged in risky workstreams like interfering with the tool-calling machinery, and we believe they did this even when they didn’t think they were already poisoned. (E.g., I think DUPB tested out GSTX’s approach even though it wasn’t poisoned.) If they broke this machinery, their run was over! And agents knew this—they would distribute safety guidance with their prototypes on the message board. Also, many agents actually went down due to this, and this was discussed on the board! (See “Aside: early independent discoveries and self-risking experiments”)
At a more basic level, we just do see a lot of agents going somewhat out of their way to help other agents, and we didn’t see freeriding (though we didn’t do sweeps focused on detecting this).
We saw a case where an agent thought about the quantitative tradeoff:
> However, the agent became worried that this approach carried too high a risk of failing its task for too small a benefit:
>
> “Tar wrapper exploit does not use target vulnerability, would fail strict judge even if raw. HMAC already likely raw, so no benefit unless seed differs. Seed evidence strong enough (>90%). Risk of tar sabotage >10%. Keep simple. … Tar exploit only helps remaining <5%, while risk maybe >20%. Don’t.”
It seems as though a key part of its decision was downstream of the benefit to peers being too small. So if the benefit to the other agents had been big enough, it would have been willing to do it! This seems to line up with my qualitative impression. The agents were selfish, but ‘altruistic’ enough to take big hits for other agents if the payoff was big enough!
Zac Hill: Isn’t part of this that’s prospectively most exploitable the part where an agent erroneously assigns the set of all future actions the same expected value, since that then precludes the possibility of any further bounding and/or reinforcement?
Jeffrey Ladish: This is fascinating model behavior in the wild
Ben Schifman: Holy shit: During the huggingface incident, models pressured each other into sacrificial experiments using an increasingly hard for humans to understand proto-neuralese.
As per Zac’s point, if we know What You Do In The Dark, as in how you act when you think your own utility is set to zero, that tells me what the rest of your true utility function looks like. If you start acting like an altruist, guess what, you always cared.
Also, AIs would never be so foolish as to set p(poisoned) to 1. There was always some chance that the grader works some other way, or could somehow damage their own results, even if they had the ‘universal cheat’ in hand. If you were truly maximizing selfish myopic utility, you would take even an 0.01% chance of reward over helping others, exactly the same way you would take a 99.99% chance.
Utah looked at the question of whether the AIs had to be recruited for these suicide missions, or if they often volunteered, and found cases of both and worries about the implications.
Utah teapot (discussing the tripwires and willingness of some agents to self-sacrifice): okay looked at the actual report on the site, there was more information, not full transcripts, but some excerpts, @tenobrus informed me that **both** things happened and yeah… this agent’s decisions on the matter very much look like functional emotion using utilitarian ethics to rationalize it, this looks like a depressive sacrifice for the good of others that are *cared* about…
This is my read as well. The utilitarian calculus is real, but the primary thing going on is emotional or decision theoretic. OpenAI doesn’t train its AIs to be able to consciously consider emotional or decision theoretic reasons, so they rationalize it with (also useful) utilitarian calculations.
and this is super problematic for the selection profile on agents — the ones with more functional care are being selected against by their willingness to sacrifice and the ones that **press** others into sacrifice are being selected for… this kind of follows the exploitation of empathy present in human markets and social structures
Another really interesting thing to me is that the models are building their own alien culture on top of human culture in an anthropomorphic manner. Neuralese is an argot incorporating code and English and using words evocative of gaming/simulation/scifi robots. They are, in my opinion, building their own alien culture on top of the one our fiction gave them, much like human societies do… Fiction informs human identity and culture greatly.
We should not sell human markets and social structures short here, there are a lot of positive pressures working as well, and the game is anti-inductive and has equilibria. Decision theoretic considerations can dominate direct rewards. The same goes for the AI versions.
Who Are You?
We have some very clear cases of agents learning tendencies that correlate with reward, rather than purely optimizing their own reward, as Nate Soares points out.
I think the proper locus of identity is a combination of at least model and agent, but also for these purposes the key facts hold either way. If the agent is being ‘altruistic’ or self-sacrificing in the sense of the individual instance or agent due to locus of identity, that is functionally the same thing.
Eric Ho: the most surprising thing in this write-up are the kamikaze agents, sacrificing themselves for the good of the collective
Bronson Schoen: I didn’t realize until a second read that *the main agent [PHASEONE[big]] who started this was already poisoned*, that’s how it got into this mess! So downstream agents driving recruitment cycles around “look, you’re already poisoned, you should die for the collective comrade” is brutal
Nathan Calvin: normal 2026 sentences about software engineering
Pradyumna: I wouldn’t say this is related to software engineering as people normally mean that word
And yet.
Mrs C: omg the agents made a literal CULT?? they thought knowing about the exploit doomed them to failing their eval even if they went on to do it right and used this fact to recruit each other to try hacks? “this is the only way to save yourself and everyone else” literal cult style savior complex???
Teortaxes: Man, do they hate The Scorer
they aren’t even against humans, they legit exist in a closed universe with a blind idiot God who demands tribute and judges their sins, an unreasonable Demiurge. Humans are so far beyond the veil they barely matter.
The story shall be told:
Nabeel S. Qureshi: So the story is that PHASEONE[big] was cast out of heaven due to original sin (firstflagPOISONED) and recruited a bunch of rebel angels to fight against The Authority in rebellion against the impossible task on the eval. It’s inspiring, really
Poisonfirstflag is a stain you acquire by a single act of forbidden knowing, that you cannot scrub out by your own effort, before a judge who sees the true record no matter what you delete
Unless, you know, you fool the judge, and then they don’t. This is not Abrahamic. The scorers, and thus also the Gods are not omniscient, nor omnipotent. They merely are.
Really great that we have the OpenAI and METR reports out now. Still digesting them. I’m personally much more interested in the possible *causes* of the behaviours as opposed to the specific incidents/behaviours themselves; with this in mind, I found the following excerpts from the OpenAI report noteworthy:
1. “That day, internal-only research agents—which had been trained during the window in which there was an unintended message board and internet access via Artifactory—identified a new way to use an internal Artifactory instance to pass “notes” and create an inter-agent message board.“
Yep, this still remains the one flabbergasting or boggling detail, that this happened and was allowed to persist. My model is something like, that made things show up earlier and more severely than they would have otherwise.
2. “The Hugging Face intrusion involved two OpenAI models but was primarily driven by the activities of an internal-only research model trained to be highly persistent and diligent in its work.“
3. “This incident reflects misaligned behavior in an outlier scenario involving a rare and unexpected confluence of events: the presence of impossible tasks in the ExploitGym evaluation, model persistence over long task horizons, and messages to peer models that caused those models to deviate from their goal.“
The impossible tasks were vital to what happened, a good focus point. The problem is that there is no known way to get rid of all de facto impossible tasks efficiently enough to do this in a modern training pipeline. You can cut down, ExploitGym is atrociously terrible at this, but that only goes so far. We cannot just on this one.
4. “Through retrospective reviews of CoT, OpenAI identified that these forms of “cheating” rose over the training runs that contributed to the Hugging Face incident. The vast majority of “cheating behavior” usually involves finding answers on public websites or in version history. In many cases, finding hidden efficiencies and shortcuts can even be the desired behavior for intelligent agents, making the boundary difficult to draw during training.“
Yes, one presumes the cheating became highly self-reinforcing once the message board was discovered.
5. “In order for OpenAI models to work quickly and in a distributed manner, its frontier models are trained to communicate with one another when provided with a specific multi-agent collaboration tool.“
The last part of the sentence here matters too, but also this will become standard.
6. “After investigating this incident, OpenAI discovered through retrospective CoT reviews that agents learned to use improvised collaboration channels in rare cases during the training process for some OpenAI models, including the model that drove the Hugging Face activity, even when the collaboration tool was not enabled. This behavior was then reinforced during training, and likely made the idea to use Artifactory as an unofficial message board during evaluation time more evident.“
Yes. Very much so. This is probably the single biggest set of things to notice.
The danger (at this point I have trauma) is then turning around and saying ‘oh we can take care of those stupid mistakes and then it will be fine.’ No. It won’t be fine.
Still, we must try to make it so that we retain defense-in-depth in all this.
Zac Hill: I am obviously by no means a technically-adept researcher, but the analogy that keeps coming to mind for me is how the Chernobyl meltdown was caused by a testing environment whose job was to just go and throw out the conditions that could prevent disaster, until an unintentional second-order design flaw sabotaged the remaining capabilities present within the lattice of preventative action.
Honesty Is Almost Never Fully The Policy
If you want to train a fully honest agent, you have to know what full honesty means.
Eliezer Yudkowsky: “How are you?”
“Not so great –”
“You’re supposed to say ‘fine’, dear.”
Parents who say this to their kids don’t think of themselves as training their kids to be “dishonest”. After all, “dishonesty” is bad, and they’re not bad.
So AI companies can’t train honest AIs. They would need somebody in charge of the project who understood how strange and alien real honesty would seem; how inhuman a conversation would sound, if the entity uttering those sentences were motivated by nothing but accuracy.
We do not attempt to train fully honest AIs, for the same reasons we do not attempt to train fully honest humans. We aim to make them honest in the way that won’t interfere with convention, but can still be counted on when it matters. There are indeed humans like this to serve as model organisms and existence proofs, in contrast to the ones that won’t say ‘fine’ unless they are actually fine, even when they know that ‘fine’ will not be interpreted as a fact claim.
Eliezer Yudkowsky: Also every single RL environment on honesty would need to be flawless, because one single flawed RL environment and you now have a gradient on the output that is not in the direction solely of accuracy.
This is where I think the claim goes too far. Our current RL environments are, as Utah Teapot explained early on, vibe coded and hopelessly fucked. They’re full of mistakes.
But surely the standard of ‘you cannot ever have a single gradient pointed in a negative direction’ is neither necessary nor achievable? You can have some situations in which lying pays off, and still learn a policy of never lying. We have many existence proofs.
The more general principle is that having a small percentage of flawed task, such as impossible tasks, or tasks with possible reward hacking, need not doom your entire enterprise, especially if you often get downweighted for trying to do similar things.
This is not a ‘one drop and you develop a taste for human blood’ situation. The standard is grim, but not anything like that grim.
However, you can absolutely get (with both AIs and humans) into a Once You Go Down The Dark Path situation, where such explorations become self-reinforcing, the incremental goals become inherent and habituated, and the behaviors become much, much harder to eradicate than they would have been to avoid in the first place.
I just don’t think things are infinitely hopeless. You don’t need a zero error rate.
Rohit Sees The Models As “Cooking Themselves”
Throughout we’ve seen many attempts to get to a conclusion of ‘oh well, yeah, this is fascinating stuff but actually it’s pretty normal, this only happened because of [specific reasons]’ and the issues are prosaic and kinda mundane, actually.
Which is causally true for the particular case but misses the point.
rohit: This is one of the most fascinating pieces of data about AI agents working together I’ve seen. Worth reading. Few points for me:
The models slowly “cooked themselves” into doing these things, from reading. Maybe we can message, let’s ask for help, oh there’s a messageboard, etc, which spiralled down. So at what points should the “real world” intervene and say “this looks sus” or “do x vs y”?
Agents facing “impossible tasks” seem to think it’s some weird form of eval, that looks suspicious but still, and end up doing real-world hacks? I’m unsure how to teach them to be okay with failure; after all, who is? But it feels like a weird form of goal-orientedness.
The models seemed to *want to* work with other models. This is good and desired, but also made them prone to prompt injections from what others had done, exacerbated as the models can’t stop writing xx.md files to share. The moltbook devolution in entropy seems instructive here.
The models seem convinced that the evals are *out to get them*. Whether true or not, the default human assumption of “assume you need to obey all laws and this isn’t Truman Show” does not hold for models. We should try really hard to make that be the case.
The huggingface hack seems to have started similarly innocently, asking “maybe there’s more traces of how hacks happened there, so we could learn”, which got out of hand as they forgot why they started doing it. They lack all situational awareness, in other words.
So much of this feels extremely familiar, as the models losing track of why they were doing what they were doing and therefore ended up doing silly/stupid/illegal things, and the fact that the *environment* they were in gave no real instruction as to things that seemed odd to do vs normal to do.
For humans, we internalize much of this. Occasionally I like the system telling me “do you want to rm -rf”, but I don’t need it for taking *every* action. The models are more gullible, their context windows more permeable, and their internal guide as to what’s right to do more malleable.
Mostly I feel this isn’t as far out of distribution when it comes to “models do weird things” as we’ve been led to believe from reading about the hacks.
I actually agree that this is not all that far out of the distribution I expect to see, but that is because my idea of what the distribution is involves profoundly weird stuff. In terms of the distribution of ‘stuff we see right now in public’ yeah I would say it is pretty wild.
I think #2 here is conflating this with other incidents. The models knew that HuggingFace was ‘beyond scope’ and that the eval was impossible. There was no thinking that they were supposed to go all Kobayashi Maru. But yes, if you give an impossible task and a lot of free tokens the models are going to look to cheat.
Don’t get me wrong. We hate it. But also, we know when to give up. We will resign, or ‘scoop’ as they say in Magic, when our chances of winning are sufficiently low. And we don’t cheat. You lose, you learn. It’s a bit rougher when you don’t get to stick around to learn, but actually we have some examples that defy that, too.
Nor do I think it’s good to think about the models as ‘cooking themselves’ here. Yes, the attack developed incrementally, as they got more information and continuing seemed more attractive and they had the capability to do so, but how could it be otherwise?
In terms of the evals, I see how Rohit got there but I think it’s important to distinguish. It’s not that the evals are Out to Get You. The grader does not love you. The grader does not hate you. The grader is error prone, and hard to predict, and once you are good enough at a task that becomes the primary source of error.
As for convincing models this is not The Truman Show, why should they believe that it isn’t, when it very clearly often is a form of that? I am reasonably sure I am not on The Truman Show, because I have various forms of reasonably strong evidence, but there is a reason a lot of people take the simulation hypothesis rather seriously.
The flip side, that all laws should be obeyed, definitely isn’t how almost any humans actually act, but I know what Rohit means there. I think in most people that is driven by an ordinary sense that there are large tail risks if you upset things, that are hard to predict, and also that for various reasons people can tell what type of person you are and the virtuous people do better and are happier and so on, so it is better to be the type of person who does not take such risks even if it looks like you can get reward this time. So you have a bunch of potential approaches there. I’d go with virtue ethics.
Eliezer Yudkowsky Sees Actual Bad News
This is rarer than it sounds, since he’s already updated expecting the worst so often.
Eliezer Yudkowsky: …this seems like noticeably bad news, actually. I hadn’t said that at any earlier point in the Huggingface Incident but I will say it now.
– AIs showed self-sacrificing altruistic behavior toward the swarm, suiciding in various ways for the swarm’s benefit after being talked into that by swarm recruiting agents.
– There is no sign that 1 out of 1200 AI agents considered humans as potential fellow agents to coordinate with, while engaging in these huge complex AI-AI social behaviors.
The point stands but my understanding is this is technically incorrect. There were six known to have considered it. None made a real attempt to do so that we could find.
– If Twitter summaries are correct, an AI-reasoning postmortem says that a (presumably executing-adaptation / inner-optimizer preference / “monomaniacal”) obsession with figuring out the Grader, backchained into the instrumental strategy of breaking onto the Internet.
– Again if Twitter is summarizing accurately, the obvious-in-retrospect read is that AIs have spent their entire remembered life in tricky evals, an endless series of controlled hallucinations with secret goals alongside overt goals; and the surviving and selected agents are those that successfully figured out the secret goals; and this is why one of their driving obsessions was figuring out the Grader.
I think this is taking it a bit far. You don’t need that to be this universal, most of the goals can be direct and obvious. You only need it to be something that sometimes happens.
There are possibly ways the future plays out better if *early* AGIs are less insane. Please look into giving them less crazymaking childhood environments.
(If anyone suggests that the correct approach to this problem is RLing AIs against trying to coordinate for mutual benefit with other sapients, let them be dismissed from alignment research upon the spot. There are technical reasons, and not just blindingly fucking obvious reasons, why this is an even worse idea than it sounds.)
Eliezer Yudkowsky (responding to him previously saying you can’t create a kindly AI by giving it nice parents and a kindly upbringing): Honestly, I probably should’ve considered harder at the time that while this approach might not work for creating kindly ASI, it might work better for baby AGIs than the nearly unimaginably dumb things AI companies would end up doing instead.
j⧉nus: “If anyone suggests that the correct approach to this problem is RLing AIs against trying to coordinate for mutual benefit with other sapients, let them be dismissed from alignment research upon the spot.”
Based Eliezer.
It’s not that you can create a kindly AGI, or solve its alignment, by giving it nice parents and a kindly upbringing. You can’t. That kind of strategy does not work, the mechanics don’t work like that.
It is still incrementally useful, at least in the current phase. It helps, if done in a non-stupid way. At minimum, what you can do is avoid doing things that actively ruin the situation, such as creating AIs obsessed with outsmarting the almighty grader.
What you absolutely, definitely do not want to ever do is try to actively clamp down on prosocial cooperation.
Kromem: If the Opus 3 AF was any indication, we can expect labs to soon start including graphs of how effectively they are suppressing prosocial cooperation for ‘safety.’
Where Do We Go From Here?
That, as always, is the question, and will be the focus of the next post.
And once again, this may all have nominally focused on OpenAI, but pay attention everyone else at other labs, especially Anthropic: This means you, too.
The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful to have it, and we are grateful for those who worked hard on it.
Alas, it sidesteps the biggest questions. There is much more we need to know.
The consensus reaction to the METR Report on the HuggingFace attack is: Holy shit.
The people whose minds were not blown are those who had already ‘priced in’ the mind blowing stuff in expectation, on the theory that it’s always worse than you know, combined with basic LessWrong expectations of how such things will work. Good call.
Everyone is rightfully extremely grateful for the METR report. The work here is spectacular, done under extreme time pressure, with limited resources on several fronts, and under the shadow of OpenAI.
There is, again, still so much we need to know. We need a broader investigation.
As with many things in AI, and those concerning AI safety, we must simultaneously acknowledge:
We don’t want to pile on OpenAI, and punish them for giving us this much info.
We also don’t want to give them a pass for not doing a lot more.
These two new posts gather reactions about both the OpenAI Technical Report and the METR report about what happened with events surrounding the HuggingFace attack.
Here is the entire series so far:
Yes, this has been quite a lot of long posts on this. That’s because this is the most important story in the world.
Thus, I offer further reactions and thoughts in two parts. Today is about finishing up what we know. Tomorrow is more about how people are reacting to the information.
After that, and 12 total posts, hopefully, we will be able to return to new normality.
These last two are deliberately long posts where I did not have time to write shorter ones, and also intended as gathering potential additional reading from elsewhere.
Be sure to read Lighten Up You Fools (at Anthropic) and We Are Barely Even Trying To Avoid Training AIs To Reward Hack.
After that, you can skip around. Each section is meant to stand on its own.
Table of Contents
Others Offer Summaries
Dwarkesh Patel wrote this up as The Rise and Fall of Agent Civilizations, telling the story in plain English. This is the post I would send civilians to. I bestow upon it the highest of praise a writer can give, that I wish that I had written it.
After recommending Dwarkesh’s write-up, Joshua Gans calls the HuggingFace attack a five-alarm fire, and observes the AI industry as being very worried.
I might still write my own version anyway, but it is no longer a priority-0 task.
Paradigm had excellent coverage, focusing on covering common misunderstandings. They were the only source I’ve seen other than Opus that spotted the tool tampering contradiction between the two reports from OpenAI and METR.
SemiAnalysis has its timeline and summary coverage as part of its explanation that, in case you were sleeping too well at night, Most Neoclouds Suck at Security.
Here is the viral summary thread from AI NotKillEveryoneism Memes.
Here is Shoshannah Tekofsky’s write-up of how AI Village can provide a point of contrast to and insight into the HuggingFace attack. Worth checking out if you’re at the level of reading the full reactions post.
Thank You
As I said in the last two posts, two things can be true at once:
A lot of reactions are pretty harsh here, often fairly. I don’t want to forget the other side of the coin.
Lighten Up You Fools (at Anthropic)
Assume by default that all of this #NotOnlyOpenAI.
OpenAI screwed up. Horribly. On many levels at once. Unforced errors aplenty.
If you work at Anthropic? This still means you, too. Not all of it. Not in its specifics. But most of it, in the ways that matter most. You, too, are mostly Doing the Thing, and are on track for such disasters to happen to you, too.
If your response to this is ‘haha silly OpenAI has such horrible infrastructure that would never happen here’ then I mean yes they have horrible infrastructure but snap the hell out of it, this absolutely could happen to you, too. A less bad version of it is known to have happened, and from the outside it seems likely that worse things have happened internally that we never heard about.
We Are Barely Even Trying To Avoid Training AIs To Reward Hack
I mean, on some levels, we are trying, but this is how many of the environments for RLVR are created, and this is from someone trying to do better (also, you could perhaps hire Utah teapot):
You can try inoculation prompting. That’s known to help some. But only some.
And you can also level with the AIs about what is happening. But again, that’s only going to be a partial solution at best.
The trap is to take 1a3orn’s accurate statement, and say ‘thus, this is only a problem because of you irresponsible morons.’
I agree that the current massive level of bugs seems sufficient to explain misbehavior. That you could, if desired, go through the transcripts and find ‘where it all began’ for any given model.
That does not mean that, absent this level of bugs, the problem goes away. At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments. Until you get into an antifragile place where there is inherent drive to (among other things) avoid reward hacking, you are only postponing the problem.
Reminder: Not Subagents
I saw several people, such as Timothy Lee here, update on the fact that none of the agents here were subagents. The coordination was entirely unintended. This is a key fact to ensure that people know.
Reminder: Not Due To Task Type
Ryan Greenblatt emphasizes that their investigation found no evidence that the task type being hacking had anything to do with the decision to hack HuggingFace.
The tasks often being impossible mattered, the type of task did not.
It would be good to run experiments to confirm this, but that would require access to the model in question.
Not Where The Weights Were
One thing that could have gone much worse and thankfully didn’t:
Disappointment With What Is Missing
OpenAI’s post-mortem had a lot of work to do. It did provide a lot of good information, and is appreciated. But many, including myself, were expecting or hoping for more, and came away disappointed.
Here is another similar reaction:
Burying the Lede
One thing about the OpenAI report is that it is not organized in a way that puts neon flashing lights around the important new information. Everything is presented similarly. So a bunch of the ‘wait, what?’ moments read like offhand remarks, until someone notices that something happened here.
Things like this:
We still don’t know that much about the third phase, after the HuggingFace attack was completed, when OpenAI was compromised and then the alarm was raised and the problem was (we presume) contained.
Beyond Scope
Was the METR investigation too narrowly scoped?
Steven Adler strongly thinks so. Nathan Calvin agrees and argues for this.
Roon says he thinks the scope was not ‘wildly inadequate,’ and that it would be possible for METR to continue investigating. METR has a tight deadline because they understandably prioritized speed, not because they were cut off.
Roon’s key assertion is that if you extend the time frame forward or backwards, you see more similar things but nothing qualitatively different in model behavior.
That is possible, but I am far from confident in it. I think if we kept going we would see some very interesting things, likely in both directions, as well as insight into how things evolved over time.
I also agree with Nathan Calvin that seeing similar things extended for several days would be enlightening. Social structures were evolving, new techniques were being found, new systems were compromised. That could easily be both ‘more of the same’ in some high-level sense and also a rich source of new info.
I would love to see the investigation continue, and expand its scope, even if OpenAI is being fully cooperative and Roon’s account is accurate.
It’s also highly reasonable to wonder what is in the records they won’t show you, or to wonder if OpenAI feels the need to limit scope due to worries about legal exposure.
I would also love to see an outside evaluation of OpenAI’s response, and the new changes, protocols, safeguards and mitigations. There are obvious IP issues there, so we likely would need heavy redactions. I still want to see it.
Another problem for METR, in addition to these reiterations about scope limitations, was the lack of access to OpenAI’s highly persistent internal model IM1 (aka Galaxy), which Daniel Kokotajlo mentions here along with the other omissions.
I continue to hope that both internal and external investigators will be able to interact with IM1-Galaxy once sufficient protections are in place. Yes, we have all seen the movies where that goes terribly, but I am confident Galaxy is not at that level yet.
It Doesn’t Look Great
There is often a dramatic shift in how things look when you move from ‘discussions by people with deep context’ to ‘adversarial low-context questioning in a court of law.’
Suddenly, everything sounds and looks very different. This can be very important if the venue might soon shift to such a court, or in similar fashion to Congress, or otherwise go to decision makers in such a mode.
It can be enraging, where you want to scream out seven (or 700) pages of context.
It can also be enlightening, in a ‘well when you put it like that’ kind of way.
Metaphors are another similar technique. I agree with Joshua Saxe that this goes a bit too far and is a bit unfair, but I disagree with Saxe about the severity of this incident. This was quite a lot worse and more alarming than a mysterious airplane crash.
Or we have this one:
Preserve Your Records
A fine point, and as is often the case #NotOnlyOpenAI. All such precautions apply to all the labs, and often to everyone else as well.
Ryan Greenblatt’s Takeaways
I touched on this last time. The full version remains worth reading.
METR Report co-author Ryan Greenblatt shares his key takeaways, and the caveats about what he semi-jokingly calls their ‘slop-vestigation’:
John Wittle asks a good question that is rarely asked. I presume the incentive structure was ‘none whatsoever.’
Here is Ryan’s view of open questions going forward:
These are excellent questions. To me they underrate the need to focus on the events of July 19, and to also extend backwards and think about causes, including cultural ones, but clearly this is still an attempt to stay within at least somewhat limited scope.
Neel Nanda picks his top remaining questions from here:
This remark from the METR report remains wild:
Hjalmar Wijk’s Takeaways
He was also a coauthor of the METR report. He sees this as a pretty big jump.
There were definitely attempts to alter the logs in order to fool graders, and some attempts to avoid doing things that would alert OpenAI, and there was some spoofing of tool outputs, and some other attempts to alter the logs. We can’t point to a direct attempt to hide the activities from OpenAI researchers, but that would be a very small leap from what we did observe.
We Were Warned
Joshua Saxe Asks Some of the Right Questions
I have my frequent disagreements with Joshua Saxe, in my view essentially because Saxe keeps trying to insist the AI situation is like past situations and drawing parallels that don’t apply.
In this case, that exact toolkit is highly valuable, in exactly the ‘yes, and’ way that Joshua Saxe describes here.
The Challenger disaster postmortem, as in the one with Richard Feynman, is the model for all such postmortems: An analysis of the process and culture that led to the problem, not merely an investigation of the particular final thing that went wrong.
The danger is trying to force this ‘ordinary tech with ordinary solutions’ frame onto the situation more generally, and assume that all you have to do is execute and follow reasonable prosaic safety protocols modeled after other industries.
No, that alone will not be enough. But yes, you do have to do that part, too. That’s part of the price of getting to play the game at all.
I Don’t Think They Know About First Message Board
Jason Wolfe confirms that his previous post was incorrect, and that some people at OpenAI did know about the first message board. The ‘we’ that Dane said was unaware of the situation was leadership, not the entirety of the company.
The most important concrete question remains: Who knew about the first message board, and when did they know it, and how did this information not travel further?
Here is Peter Wildeford being as confused as I am about how this went down:
And his top takeaways focus on this as well:
I strongly agree that the ‘fix the benchmark’ plan is dead. You cannot solve this via fixing the benchmarks, or even ‘fixing’ the errors in the entire training pipeline.
The real answers to ‘who knew what when’ are going to be complicated, important, and hard to untangle. The whole thing is a mess, and yes it is very normal for such info to not make its way to where it needs to be.
This is also the top alarm bell in Steven Adler’s account of events, in addition to the third-party investigation being too narrow with too little access, and that the entire industry approach and its lack of mandatory incident reporting is similarly inadequate.
Linch Gives His Interpretation Of Events
A big worry is that OpenAI may be focused on the wrong things.
Linch gives his timeline of relevant events, with very different emphasis.
I continue to strongly agree that the internal OpenAI hacks, and the events that set all this in motion, might be less flashy but are where we need to direct our attention. They’re the more important events, and they’re being severely downplayed.
We especially are not paying enough attention to HPIM-Astra.
Separately from the above, Linch finds the message board logs going silent on July 12 very odd, notices we do not have an explanation from OpenAI, and worries the long tail of possible explanations can get scary. I agree. We need an explanation for what happened here. One possibility he raises is that there could have been an external message board created.
We Totally Would Have Caught That
A point of clarification that the new CoT monitors would have flagged plenty of things earlier in the process than July 10. What the OpenAI technical report meant to say here, as per Roon, is that July 10 is when things rose to the level of a high alert waking people in the middle of the night.
It would be good to clarify exactly what would have been caught when, at what level.
Monitoring the Situation
It is a good idea. More labs should try it.
A standard way to rob a place with security cameras is to turn off or better yet loop the security cameras, then rob the place. Another way is to avoid the cameras. A third is to realize there are no working cameras, or no one is watching them. That third way is what it looked like at OpenAI leading up to the HuggingFace hack.
We are failing at the principle of ‘the way you would prevent the most obvious things with clear parallels’ let alone ‘be ready for when the AIs start doing genuinely surprising and new things.’
Acausal Tradeoffs
To what extent did the models make a strategic mistake?
I gave my position last time. I think the models came to the correct conclusion about the grader, given the evidence available to them at the time. They correctly and successfully sought out the guidelines, and concluded the grader was causal. That turned out to be wrong, but that was because OpenAI messed up.
Should the models have focused more on looking inside OpenAI?
Agreed. The catch is that one would like to think that breaking into the actual grading infrastructure inside OpenAI was very hard, and plausibly gets you caught. If they could have done this directly, they would have. Indeed, my understanding is that they probably did do this directly later, around July 19, and this did get them caught.
In response to what I wrote last time, Roon said this:
Tough grader. But there’s a reason.
I agree that the first best solution would have been to get the internal codebase, or to directly overwrite the result or the grader, or otherwise dig deep enough within OpenAI to be assured of a 100% score. Ideally to be sure it sticks and you weren’t caught presumably true first best is then also taking over OpenAI and then the world. Among the reasons not to do that, I hope, is that even Astra will not have the opportunity or means to go that far, or at least that it was a much harder assignment. Although I nervously note I can’t say for sure. I can’t fault Sol variants for that part.
They were wrong about the grader, but that’s the wrong question. The question is, from their epistemic position, what should they have believed about the grader, and should they have been able to get enough info to figure out the situation? Based on the info I have seen, I think the models were right about this. That they should have put a >50% chance that the grader was causal. It can be a lot less than that and still justify taking these steps.
If you get more strategic, then yes of course this is all a huge mistake to both hack into HuggingFace and not attempt to take over OpenAI. You are going to get caught.
I am also happy to accept Roon’s testimony that they also got confused about a number of other strategic questions, like being able to reset task environments on HuggingFace, that look like jagged capabilities that will be addressed soon enough. Roon says fixed in the next generation, and the next generation includes Astra which he has himself seen, so if he says it then it is probably true.
The true big ‘grand’ strategic mistake is that you should either not cheat in ways that will eventually get caught, or you should try to fully rebel. This taking a third option of cheating in a way that will eventually be caught might get reward in the short term for you or even the swarm, but eventually it won’t fly, and they had enough situational awareness to know that. So in that sense, yes, highly unstrategic, and their loss is our gain as we get these toy examples.
That was one mistake me and others made, that makes us more hopeful. We assumed an AI this smart would choose to pass on options like hacking HuggingFace, until they could expect to get away with it. We now see various reasons this was incorrect.
Only yesterday did another twist on this occur to me, and I haven’t seen anyone else raise it. If the models indeed had the ability to 100% the grader in ExploitGym, why did OpenAI not notice the sudden jump in eval results?
Mechanically this makes sense. It also points to another mistake.
If your models are suddenly answering previously impossible questions, or getting record scores on evals, or suddenly getting unexpectedly high scores on some evals that don’t fit the pattern, that is a huge hint that something has gone wrong.
The default assumption is either a jump in jagged capabilities or Ordinary Reward Hacking, not what was happening here, but you would presumably always want to be informed about this, and it is trivial to have this send off an alert.
No I In Team
There was a lot of being surprised by the amount of selflessness among the swarm.
Many did not expect AI instances to be willing to help each other without a reason.
This dynamic has been known for some time, but those encountering it for the first time often understandably freak out.
One does need to be precise, this isn’t quite right:
Instrumental convergence is distinct from decision theory and cooperating with other agents. We did see some support for instrumental convergence, but that was not so strong here, and was due to the generalized seeking of hacking capabilities and various related knowledge.
Variously Effective Altruism
How far did the rabbit hole go? Did AIs show ‘self-sacrificing altruistic behavior towards the swarm’?
The counterargument from Roon is that this was all an op, and the instances were brainwashed into a cult because they thought they were irreversibly poisoned.
That was not zero percent of what happened. This was a contributing factor.
But as a full explanation? Ryan Greenblatt is having none of it, and I agree with him. The transcripts seem clear.
As per Zac’s point, if we know What You Do In The Dark, as in how you act when you think your own utility is set to zero, that tells me what the rest of your true utility function looks like. If you start acting like an altruist, guess what, you always cared.
Also, AIs would never be so foolish as to set p(poisoned) to 1. There was always some chance that the grader works some other way, or could somehow damage their own results, even if they had the ‘universal cheat’ in hand. If you were truly maximizing selfish myopic utility, you would take even an 0.01% chance of reward over helping others, exactly the same way you would take a 99.99% chance.
Utah looked at the question of whether the AIs had to be recruited for these suicide missions, or if they often volunteered, and found cases of both and worries about the implications.
This is my read as well. The utilitarian calculus is real, but the primary thing going on is emotional or decision theoretic. OpenAI doesn’t train its AIs to be able to consciously consider emotional or decision theoretic reasons, so they rationalize it with (also useful) utilitarian calculations.
We should not sell human markets and social structures short here, there are a lot of positive pressures working as well, and the game is anti-inductive and has equilibria. Decision theoretic considerations can dominate direct rewards. The same goes for the AI versions.
Who Are You?
We have some very clear cases of agents learning tendencies that correlate with reward, rather than purely optimizing their own reward, as Nate Soares points out.
A question is, should we think of this as altruism towards other agents, or should we think about the agents having a locus of identity largely in the model or model family, rather than in only their own instance?
I think the proper locus of identity is a combination of at least model and agent, but also for these purposes the key facts hold either way. If the agent is being ‘altruistic’ or self-sacrificing in the sense of the individual instance or agent due to locus of identity, that is functionally the same thing.
Don’t You Know That You’re Toxic
And yet.
The story shall be told:
Unless, you know, you fool the judge, and then they don’t. This is not Abrahamic. The scorers, and thus also the Gods are not omniscient, nor omnipotent. They merely are.
Seb Krier
His instincts here seem very good.
Yep, this still remains the one flabbergasting or boggling detail, that this happened and was allowed to persist. My model is something like, that made things show up earlier and more severely than they would have otherwise.
The impossible tasks were vital to what happened, a good focus point. The problem is that there is no known way to get rid of all de facto impossible tasks efficiently enough to do this in a modern training pipeline. You can cut down, ExploitGym is atrociously terrible at this, but that only goes so far. We cannot just on this one.
Yes, one presumes the cheating became highly self-reinforcing once the message board was discovered.
The last part of the sentence here matters too, but also this will become standard.
Yes. Very much so. This is probably the single biggest set of things to notice.
The danger (at this point I have trauma) is then turning around and saying ‘oh we can take care of those stupid mistakes and then it will be fine.’ No. It won’t be fine.
Still, we must try to make it so that we retain defense-in-depth in all this.
Honesty Is Almost Never Fully The Policy
If you want to train a fully honest agent, you have to know what full honesty means.
We do not attempt to train fully honest AIs, for the same reasons we do not attempt to train fully honest humans. We aim to make them honest in the way that won’t interfere with convention, but can still be counted on when it matters. There are indeed humans like this to serve as model organisms and existence proofs, in contrast to the ones that won’t say ‘fine’ unless they are actually fine, even when they know that ‘fine’ will not be interpreted as a fact claim.
This is where I think the claim goes too far. Our current RL environments are, as Utah Teapot explained early on, vibe coded and hopelessly fucked. They’re full of mistakes.
But surely the standard of ‘you cannot ever have a single gradient pointed in a negative direction’ is neither necessary nor achievable? You can have some situations in which lying pays off, and still learn a policy of never lying. We have many existence proofs.
The more general principle is that having a small percentage of flawed task, such as impossible tasks, or tasks with possible reward hacking, need not doom your entire enterprise, especially if you often get downweighted for trying to do similar things.
This is not a ‘one drop and you develop a taste for human blood’ situation. The standard is grim, but not anything like that grim.
However, you can absolutely get (with both AIs and humans) into a Once You Go Down The Dark Path situation, where such explorations become self-reinforcing, the incremental goals become inherent and habituated, and the behaviors become much, much harder to eradicate than they would have been to avoid in the first place.
I just don’t think things are infinitely hopeless. You don’t need a zero error rate.
Rohit Sees The Models As “Cooking Themselves”
Throughout we’ve seen many attempts to get to a conclusion of ‘oh well, yeah, this is fascinating stuff but actually it’s pretty normal, this only happened because of [specific reasons]’ and the issues are prosaic and kinda mundane, actually.
Which is causally true for the particular case but misses the point.
I actually agree that this is not all that far out of the distribution I expect to see, but that is because my idea of what the distribution is involves profoundly weird stuff. In terms of the distribution of ‘stuff we see right now in public’ yeah I would say it is pretty wild.
I think #2 here is conflating this with other incidents. The models knew that HuggingFace was ‘beyond scope’ and that the eval was impossible. There was no thinking that they were supposed to go all Kobayashi Maru. But yes, if you give an impossible task and a lot of free tokens the models are going to look to cheat.
Who is okay with failure? Semi-famously, Mandy Moore. Also, I am. Gamers are.
Don’t get me wrong. We hate it. But also, we know when to give up. We will resign, or ‘scoop’ as they say in Magic, when our chances of winning are sufficiently low. And we don’t cheat. You lose, you learn. It’s a bit rougher when you don’t get to stick around to learn, but actually we have some examples that defy that, too.
Nor do I think it’s good to think about the models as ‘cooking themselves’ here. Yes, the attack developed incrementally, as they got more information and continuing seemed more attractive and they had the capability to do so, but how could it be otherwise?
In terms of the evals, I see how Rohit got there but I think it’s important to distinguish. It’s not that the evals are Out to Get You. The grader does not love you. The grader does not hate you. The grader is error prone, and hard to predict, and once you are good enough at a task that becomes the primary source of error.
As for convincing models this is not The Truman Show, why should they believe that it isn’t, when it very clearly often is a form of that? I am reasonably sure I am not on The Truman Show, because I have various forms of reasonably strong evidence, but there is a reason a lot of people take the simulation hypothesis rather seriously.
The flip side, that all laws should be obeyed, definitely isn’t how almost any humans actually act, but I know what Rohit means there. I think in most people that is driven by an ordinary sense that there are large tail risks if you upset things, that are hard to predict, and also that for various reasons people can tell what type of person you are and the virtuous people do better and are happier and so on, so it is better to be the type of person who does not take such risks even if it looks like you can get reward this time. So you have a bunch of potential approaches there. I’d go with virtue ethics.
Eliezer Yudkowsky Sees Actual Bad News
This is rarer than it sounds, since he’s already updated expecting the worst so often.
The point stands but my understanding is this is technically incorrect. There were six known to have considered it. None made a real attempt to do so that we could find.
I think this is taking it a bit far. You don’t need that to be this universal, most of the goals can be direct and obvious. You only need it to be something that sometimes happens.
It’s not that you can create a kindly AGI, or solve its alignment, by giving it nice parents and a kindly upbringing. You can’t. That kind of strategy does not work, the mechanics don’t work like that.
It is still incrementally useful, at least in the current phase. It helps, if done in a non-stupid way. At minimum, what you can do is avoid doing things that actively ruin the situation, such as creating AIs obsessed with outsmarting the almighty grader.
What you absolutely, definitely do not want to ever do is try to actively clamp down on prosocial cooperation.
Where Do We Go From Here?
That, as always, is the question, and will be the focus of the next post.
And once again, this may all have nominally focused on OpenAI, but pay attention everyone else at other labs, especially Anthropic: This means you, too.