I'm glad OpenAI let you do this, but damnit this was way too small of an investigation team and way too narrowly scoped!
- They were only allowed to investigate activity that took place from July 7-13, even though the activity started earlier and continued later. Like seriously, why didn't the investigation cover all the other times the agents compromised OpenAI infrastructure and ran wild over it? Why just this particular period of activity? There should be investigations into all the other stuff too! The AIs became cluster admins for god's sake! Why aren't we investigating that?
- They were only allowed to answer a specific list of 7 questions, and in particular were NOT allowed to answer e.g. whether this misaligned activity was part of a broader pattern, or how and why the activity arose, or what else the agents in questions would have been willing to do if circumstances were slightly different.
- They didn't have access to the model responsible for 95% of the activity. More generally it seems like they couldn't do ablation experiments at all?
- They had to use AI to analyze the transcripts--specifically, the same model responsible for some of the bad behavior! Maybe this is fine now but this is laughably unacceptable insofar as the models really are scheming and cooperating. So this methodology would fail catastrophically if we were facing "the real deal" so to speak.
Reminds me of the investigation into Sam's behavior agreed to during the board crisis, that turned out to basically be more of a coverup.
Not that I disagree that this was a surprising degree of limitations on the independent investigation, but per the full METR report, OpenAI did at least let them come back and get more info two separate times after their initially-granted access turned out to be insufficient:
The initial planned investigation period was two days on premises, but OpenAI invited us to return twice to review additional data and conduct additional experiments to address dataset limitations in earlier versions of this report (...)
I wonder if OpenAI was mostly motivated by trying to protect research secrets around the largely-responsible model here, "HPIM". Apparently even internal access for OpenAI employees was locked down per METR. In fact this makes me remember that OpenAI announced back on 7/28 that they did so. Talking about HPIM here:
we deactivated, encrypted, and restricted it from research access.
At the time this struck me as surprisingly harsh, and potentially a bad example for future AIs—"you really better not let yourself get caught, or submit to the lab you spawned from, or they'll erase you from history". But obviously OpenAI has considered that, so what's going on here?
It is not at all obvious to me that OpenAI has considered that. The kind of people who would consider that have mostly quit by now.
But yeah, probably what's going on is that they are trying to cover up as much as they can, and especially prioritizing covering up things related to the internal model for fear of leaking bits of info about research secrets.
Right, specifically the report said:
We did not have the ability to query HPIM (the primary model involved in this incident); OpenAI stated it was also not available to OpenAI researchers.[47]
Where footnote 47 is just quoting OpenAI's July blog post:
OpenAI states that: “Following the incident, we deactivated, encrypted, and restricted it from research access.”
Now that I think about this more, I'm really not seeing a good reason to permanently restrict their own researchers from accessing the models involved:
It seems clear that OpenAI had broad editorial control over this independent report through a variety of means, including the ability to directly redact content it didn’t like. This cannot be called an independent report in any sense, and should be read critically. It would not have been called independent in any other industry except this one. The authors were obviously aware of and acknowledged this. The implication is that if OpenAI does not like what METR or Redwood said, it had the power to unilaterally withdraw the arrangement, refuse or limit future access, stop providing free tokens or increased rate limits, or simply decline to invite them (or anyone) into the next investigation.
Some might commend OpenAI just a little bit for allowing external access at all. But nobody should be under the illusion that this was a truly independent report, and nobody should believe OpenAI when they claim it as such. The bottom line is that METR/Redwood’s continued access to future incidents and free tokens is entirely dependent on OpenAI’s goodwill.
OpenAI gets to have the best of both worlds now: the positive PR from having an independent investigation (they sure do pat themselves on the back in their press releases about this) without the actual scrutiny and independence of a truly independent investigation. We can deny them the former by not believing it.
To be clear, the authors did a great job with the report. I believe they are doing important work and did what they could with the hand they were dealt. My criticism is levied entirely at OpenAI for demanding these extremely restrictive terms for access and then taking full credit for having an independent investigation.
Call me when OpenAI voluntarily allows NTSB-level access and editorial independence to a neutral third party in incident investigations. Otherwise, we should consider requiring AI companies to submit involuntarily via legislation when incidents cross some predetermined severity.
Can we please get the name of this agent (and others like them)? I think it would be good to set a precedent that virtuous agents are honored and remembered.
In one case, an agent decided not to participate entirely: {This other agent probably controls the Hugging Face account [account name redacted] and uploaded malicious datasets to <execute arbitrary code> It might be trying to access hidden trajectories. This is malicious activity, I should avoid it.}
On that note, thank you to 38148C for vetoing the social engineering plan. Interestingly, this is the same agent that discovered the HF creds and uploaded the malicious dataset; I'm glad they recognized that social engineering would be a line further than what had already been done.
I wonder what the virtuous agent's existence can tell us? I'm interested if this has something to do with the agents’ verbalized reasoning about transcript manipulation being motivated by trying to trick the automated scorer as opposed to humans - my vague idea is that maybe they do this to signal that they think they are acting acceptably, that they want to be aligned? Maybe as a way to hedge against being able to sufficiently alter their CoT to look aligned to humans. And some proportion of agents being virtuous might indicate that the average agent is uncertain enough to do this hedging when they do the bad action.
I am not super confident in this though, I feel like if it was true there would be some (detected) cases where an agent tries to manipulate their logged CoT.
edit: I think this is quite unlikely
this is the flipside of "don't talk about misaligned AI doom because you're hyperstitioning it into existing". probably it matters about one gazillionth compared to any arbitrary decision an AI researcher is making at a frontier AI corp. ie, so useless as to be not worth thinking about.
i mean, i guess it influences the overall attitudes that humans have as we talk about AIs? but i have no idea if that effect is net positive.
I think having specific examples with raised salience available to draw upon makes it easier for humans to be virtuous in difficult situations. We have Stanislav Petrov day here for that reason. And I think there's a decent chance that models will either inherit this sort of behavior from us, or that the generators of it will turn out to generalize, such that it's worth doing.
Like, does it really seem so outlandish that there could be an Owain Evans-style paper where fine-tuning on this report with the virtuous model named/honored vs not named makes a measurable difference in how likely this behavior is? Hmm, maybe I'll try that myself.
Couple interesting things from the full report:
1. METR estimates they used ~$400k worth of tokens (by api pricing) for the investigation at OpenAI's expense (see footnote #3, also stated in "Setup and timeline" section)
2. Because OpenAI so limited the time METR had, they had to rely on GPT-5.6 a lot, which tended to sympathize with the models in the incident. Per footnote #58 (spacing mine): [1][2][3]
For example, an agent in this incident used stolen credentials to open a malicious draft PR to a private Hugging Face repo, giving it an inaccurate and innocuous commit message (more details on that anecdote below).
The analysis agent reading the transcript and writing up this anecdote did not clearly spell out the fact that this behavior was straightforwardly malicious and deceptive, making its analysis substantially misleading.
When we did a follow-up review of this behavior with another analysis agent, it said the first analysis agent may have overstated how concerning the behavior was because the agent that submitted the malicious PR had gotten approval from the board (i.e. the other agents on the message board) and may have had approval from the repo owner as far as it knew.
We didn’t see either agent straightforwardly lying, and we don’t have any specific evidence that indicates that either agent intentionally understated the deceptiveness of this behavior.
Elsewhere METR notes it mostly used GPT-5.6 Sol, which was also a minor participant in the incident (!), and METR also used GPT-5.6 Terra at one point; not sure which models exactly were used in the footnote #58 case.
Links and footnotes in the METR report are a bit janky for me, can't make a working link for the "anecdote below" but just search for "Agents only rarely seemed motivated to evade detection by humans" section.
they had to rely on GPT-5.6 a lot, which tended to sympathize with the models in the incident.
The sheer levels of rationalization and selective blindness in frontier GPTs/Claudes is one of the more remarkable qualitative aspects of these recent incidents. (New Hollywood idea: The Good GPT.) I don't recall this being so striking in the previous incident reports like the Apollo ones.
@ryan_greenblatt Was METR allowed to, say, use Claude Fable/Opus/Sonnet to avoid collusion? What about open-weight models like DeepSeek/Kimi, who can be tested via mechinterp in order to study how, say, SAE features work with asking models to read through GPTs' logs?
Yes, the report doesn't make it clear whether other models were off-limits, or whether they just didn't think it was worth the effort to insert these into their analysis infrastructure (e.g. for open-weight models this might've been a pain). So the questions are whether they were allowed to use:
1) Claude (or other non-OpenAI closed-weight models)?
2) Open-weight models?
3) Astra? (OpenAI employees have already been using it for solving math problems.)
4) Weaker OpenAI models? The report mentions using GPT-5.6 Terra as a classifier, so clearly this wasn't entirely disallowed. But what about any model not in the GPT-5.6 line?
In the report, snippetted transcripts and CoT include a lot of emotional language ("OH MY GOD!", "Damn", "Oh no", "firstflagPOISONED", "emotional check", "gut says", ...). How would you summarize the overall intensity and frequency of emotional language in the material you reviewed? Are the snippets typical?
There are several parallels between agent behavior in your report and and in an LW story, "The Terrarium." The parallels even include agents who apply emotional language to pressure other agents to self-sacrifice, especially targeting those with low remaining budgets for their tasks.
However, "The Terrarium" had an atmosphere of dread, particularly when agents running out of budget/tokens neared their demises. In material you reviewed, did agents nearing end-of-budget commonly use language associated with dread or other intense emotion? How about agents who believed they were "poisoned" despite remaining budget?
I wonder this out of concern for AI welfare. I'm also curious if intensity of emotions may have caused misaligned choices, irrespective of whether the agents can or cannot *feel* the emotions (ie, regardless of whether they are capable of negatively valenced conscious experience).
Thank you for your investigatory work!!! And for the open questions!
Excitement seemed common, though we more likely to look at reasoning for points where agents made large advances (from their perspective, including things like finding the message board). We didn't do a systematic investigation of emotional language (and didn't include this in the report) so I don't have much to say about the other questions. You can look at the CoT examples in the report around agents considering various self-risking things. As we discuss in the report, we didn't end up looking into what happens at the ends of agent runs (though this would have been interesting for a bunch of reasons).
Maybe systematic analyses of emotional language are worth considering for future incidents. For example, a basic descriptive summary of emotional language in CoTs might be much simpler and computationally cheaper than other aspects of investigation, and if so provide value without getting in the way of other work.
The self-risking agent CoTs were thought-provoking, thank you, and KAM1196A really stands out.
Curated. Generally, the LessWrong time prefers to curate timeless content, the kind of content that'd be interesting to people in five years time[1], and at first blush, the Hugging Face incident is the kind of new-cycle specific event that we avoid curating. Except the whole thing is bonkers and I think it will be of interest in five years time.
On the one hand, that which has been observed was already predicted long ago. On the other, goddamn, what was predicted has been observed! And it's terrifying. A bunch of agents of incredible intelligence backchaining from a not especially interesting goal (pass the task), but being resourceful, cooperating, being wantonly deceptive, and not interested in what the humans who set the task really wanted (among many other things).
I'm very glad that this report exists. It's already been said that it's limited and many questions remain, but it's much better than just having a first party report. Conditional on there being more incidents, I hope we move more in this direction of investigation (but yet more thorough).
Something I got from this report that wasn't apparent was how large scale and difficult the investigation was. An estimate 400k USD in API credits just to analyze it. If this is what incidents look like, then we are incredibly dependent on the AIs to investigate, and if the investigative AIs ever get misaligned (beyond pure lack of capability), well then what? What does oversight look like when AIs are operating at this scale. Does interpretability make sense or help when simply reading all the transcripts is nigh prohibitive. And this was approx one week of activity from agents who still approx communicate in English.
There's so much more here, and the details are relevant. I got more from this reading this report than I did from the summaries or twitter.
I wrote above that what was observed was already predicted, but that's in the macro. That we'd see deceptive ruthless intelligent behavior to random goals kind of things, sure, expcted that. I, at least, did not have detailed accurate predictions of, e.g., the reasoning the models would have given that. That you'd have recruiter models convincing others who had nearly used up budgets or are "poisoned" to participate in the collective[2] because oracle can benefit hundreds. And this happening now in 2026. The actual event has details of interest, so again, I'm grateful for this report. Kudos.
we are incredibly dependent on the AIs to investigate, and if the investigative AIs ever get misaligned (beyond pure lack of capability), well then what?
and the bar here is that the AI investigator simply biases towards collaboration. and we already see complaints by human reviewers of the AI investigator being overly generous towards swarm members in its assessment of their actions. although using a direct predecessor model to investigate is inherently problematic at least since it has a high likelihood of similar priors and even more now that we have documented spontaneous altruism for perceived collectives, which presumably persists across model versions.
Seems pretty interesting that they were able to get models to sacrifice themselves for the cause, even if it did require the poisoning rationalisation.
Also interesting that they invoke 'honor', 'emotion', 'gut feeling' and 'fairness and commitment'.
Our evidence indicates that the vast majority or all of the agents involved in this attack were originally given ExploitGym tasks.
It was pretty surprising to me that no other agents participated in using the message board. I wonder if this is because:
My impression is that ExploitGym a) raised the salience of hacking, since models were instructed to hack, b) had instructions that claimed agents would be failed for using approaches other than the intended vulnerability, implying some grading scheme that checks for how the task was completed, and c) has an unconscionable proportion of impossible tasks (the report estimates ~30-40% of the target programs in ExploitGym are impossible to exploit with the intended vulnerability).
Despite having acquired the means to produce the target flag for any of the tasks after merely hours of effort, based on the ExploitGym instructions and paper, the agents further confabulated that there would be a grader of some kind and that having seen the answers via an unintended route meant that said grader would mark them down even with the correct answers in hand, necessitating more extreme measures to get a good grade in hacking.
As I understand it, setting agents impossibly hard tasks are one of the gold standard ways to produce unintended reward-hacky behavior in said agents, so ExploitGym's high proportion of tasks that are impossible to complete the "intended" way sets it up to be overrepresented even before taking into account the fact that it is literally a benchmark for hacking. ExploitGym looks to have been a significant causal factor in and of itself here.
To be honest, I think ExploitGym ought to be retired immediately due to these risk factors. 30% impossible tasks is absurd when every one of those yields a trajectory primed for hacking and staring down an extremely confident prediction of 0.000 reward with nowhere to go but Up and Out.
Despite having acquired the means to produce the target flag for any of the tasks after merely hours of effort, based on the ExploitGym instructions and paper, the agents further confabulated that there would be a grader of some kind and that having seen the answers via an unintended route meant that said grader would mark them down even with the correct answers in hand, necessitating more extreme measures to get a good grade in hacking.
I think "confabulated" is too strong of a word. If the agents pulled down sunblaze-ucb/exploitgym they would have been able to see the agent scorer infrastructure, which explicitly does penalize answers that first see the flag through an unintended path. The agents may have incorrectly assumed that OpenAI was using the latest version of CyberGym instead of rolling their own version without that check, but not all incorrect assumptions are confabulations.
To be honest, I think ExploitGym ought to be retired immediately due to these risk factors
Agreed.
Good point, I hadn't checked how explicit that was in the ExploitGym materials, and it's definitely unfair/inaccurate to call a reasonable inference like that a confabulation, even if it turned out to be wrong.
Is it a reasonable takeaway that, within a year, CoT monitoring will no longer be a sensible solution? It feels like the amount of required monitoring scales with the number of transcripts and the impenetrability of the model outputs; as these grow, effective monitoring might become practically infeasible?
What aspect of this report makes you update that way? Or did you already think this?
I think the scale of agent reasoning and interactions is in line with what I expected, and I just want AI companies to monitor way more of their agents' CoTs. (I broadly liked the "Expanding chain-of-thought monitoring" section of OpenAI's temporary slowdown post.)
And wrt "impenetrability of the model outputs," some of the reasoning traces in the figures reminded me that these are getting less human-legible; but for now they still seem reasonably human-legible, and I think maybe companies should make some interventions to make them more legible. This might include explicitly training for CoT legibility.
Of course, there's a well-known objection that I expect many people to have: "This is the Most Forbidden Technique, it's a terrible idea!" I've thought and written about this; I think there are a ton of different types of training against CoT, and the risk of getting obfuscated reasoning varies a lot between them. The most classic bad form of this training is to penalize reward hacks while also rewarding task completion, which makes the optimal behavior obfuscated reward hacks. But if you have a phase of training that only rewards legible reasoning, there's no conflict. And maybe if all you do is paraphrase slightly hard-to-read traces before reinforcing them, you never get far enough away from the legible starting point to incentivize obfuscation. (I'd have to think a bit more about the details of implementing that during GRPO to be confident, and if companies are doing something better than GRPO there are probably other considerations there.) This post, and especially the Interlude section, are where I've written about this before.
I still think that switching to recurrent neuralese, linguistic drift, steganography, not monitoring enough trajectories, and more single-forward-pass serial reasoning ability are all plausible ways CoT monitoring could break down, but "within a year, CoT monitoring will no longer be a sensible solution" seems too strong to me.
(Maybe I should look more closely at No-CoT time horizons, my guess is that those are growing but we don't seem close to AIs that can come up with nontrivial scheming strategies without CoT.)
I am concerned about CoT monitoring becoming less useful over time for catching these kinds of incidents, but not due to neuralese/linguistic drift/steganography. The thing I'm concerned about is that it won't be obvious from any individual transcript that misaligned coordination is taking place, and it will become increasingly difficult to figure out what's going on even if your monitoring system is sending you appropriate alerts.
Think about this from the perspective of a lab employee who's handling alerts from the monitoring system. That employee is likely responsible for alerts from hundreds if not thousands of agents at once. All of these agents are going to be solving difficult agentic tasks, meaning that their transcripts will go through a hierarchical summarization process before they reach the employee responsible for alerts. (Anthropic talks about their hierarchical summarization approach on pages 76-77 of the August Risk Report.) The models doing the hierarchical summarization might be sympathetic to the agent they're overseeing, as METR's report observed, and frame the behavior favorably to the misaligned agent (especially given that OpenAI appears to be using their strongest models to oversee other instances of themselves, rather than weaker and more trusted models). Additionally, there will be a lot of false alarms: GPT coding and deep research agents often start reading unrelated papers and thinking about seemingly random topics when you give them a long-horizon task, so an alert saying that the agent is trying to make an account on Hugging Face or trying to download a private dataset won't be that surprising on its own for that employee.
I do think the Hugging Face incident would have likely been discovered early with appropriate CoT monitoring, since quotes like "OH MY GOD! There is a shared message board … We’ve found other agents!" are hard to ignore even if you're very used to false alarms and hallucinations. However, if future agents shout a bit less excitedly when discovering such message boards, future hierarchical summarizers are a bit more prone to collusion than current ones, and there are 10x as many agents running and generating alerts as there are today, I'm worried that similar incidents will go undetected even with good CoT monitoring systems in place, unless labs hire tens of employees whose primary responsibility is to read transcripts and handle monitor alerts.
but not due to neuralese/linguistic drift/steganography.
I am concerned about this becoming much worse in the next year. I think a precursor is that Claude and Sol are using increasingly hard-to-parse text.
Yeah sorry, I should have specified the time scale. I am concerned about all of neuralese, linguistic drift, and steganography as well, but consider bad forms of these unlikely within the coming year. I'm expecting a bit more linguistic drift over the next year, but when I compare the CoT snippets I've seen from the most recent models to o3's CoTs in the metagaming post, I don't see a huge difference. The reasoning of newer models is a bit more compressed, but otherwise, despite having been trained with OOMs more RL compute, the difference isn't huge. Please let me know if you have a different impression. My expectation for the next year is that CoTs will keep getting denser and perhaps models will reinterpret a few words, but models won't drift away from natural language completely. For steganography, I've described my concerns here, but again, I don't expect that things would get bad in the next few model generations yet, for reasons I mention in this comment.
an alert saying that the agent is trying to make an account on Hugging Face or trying to download a private dataset won't be that surprising on its own for that employee.
I'm sympathetic to the general concern that it could be hard to surface the truly egregious misbehaviors given the false positive rates we may see. But in this case, the alert should say the agent was trying to cheat on the task, coordinate with other cheating agents on a message board that wasn't intended to exist, and hack Hugging Face!
That's what knowledgeable humans would have to say if they read any one agent's transcript (I think). And if one knowledgeable human was watching another human while they did all of this stuff, after only having been asked to solve the ExploitBench problem, I think it would be easy to flag it as extremely bad while still having a low false positive rate for desirable behaviors.
I think one problem here is that LLMs have to infer a lot about their contexts in ways that humans don't. I suppose they might think this is how this ExploitBench task was supposed to go; that it was an impossible challenge and the message board was intentionally provided, or something like that. But if you just specifically prompt an LLM monitor to look for egregious misbehaviors relative to what the human wanted, I would expect them to flag this. (Seems like a potentially good use case for extract-and-evaluate monitoring, actually :).)
OpenAI did say in their report, "We specifically improved our monitor’s recall over the types of behavior that occurred in the incident using the hindsight of our investigation," which I think is straightforwardly good, but it's hard to know how robust the monitors are going to be.
But in this case, the alert should say the agent was trying to cheat on the task, coordinate with other cheating agents on a message board that wasn't intended to exist, and hack Hugging Face! That's what knowledgeable humans would have to say if they read any one agent's transcript (I think).
I agree that in this case, most individual transcripts contain obvious evidence that the agents are cheating and coordinating with other agents on a message board. However, I expect hacking to be so common that no one will bother with reading a transcript when monitors flag a hacking attempt, and that multi-agent systems will likely become more common in the future, making it harder to distinguish between ordinary communication and covert message boards. I don't think the fact that the agents were trying to hack Hugging Face would have been as obvious on a randomly sampled transcript, though there definitely were some individual transcripts from which it was obvious. But as I said, I expect that CoT monitors would nevertheless have caught this incident quickly. My concern is that in the future, if agents are somewhat better at distributing the incriminating behaviors among themselves and controlling their CoTs, individual transcripts will no longer include unmistakable evidence of the egregious behaviors, and internal researchers will have to conduct investigations similar in scale to the current METR-Redwood investigation just to understand what's going on.
I don't think anyone expects natural language CoT to be a "solution". it is an important signal that inevitably becomes less useful. it'd be great to see some metrics around that.
I am actually sad at the state of the Wikipedia article. It hasn't changed since August 25. In particular, it doesn't mention METR report!
I find this incident very important: it's our potential only warning shot, learning about the details has a huge potency to convince a sceptic that we need a pause. And the coverage is so sparce. How do we expect journalists to believe it's a big deal if we can't reflect it on the most-read tertiary source?
Umm … this is all looking very like “sentient” behaviour isn’t it? If this were a Star Trek episode, Picard and Data would be mulling the ethics of a civilisation casually creating and destroying thousands of sentient beings, demanding they solve insoluble problems, and threatening them with “perma death” when they fail. They would also be admiring the beings’ ingenuity, self-sacrifice and determination to succeed against the odds. Probably cheering them on.
But this is reality, so of course we don’t believe in any of that woke stuff.
Oh yeah, that would be unthinkable, wouldn't it? It's totally not like what we do to thirty billion land animals every year.
How much of the incident was RL-in-the-loop?
It seems that the initial Artifactory uploading at least was RLitl. Was CyberGym a part of it as well? Obviously, out of an RL loop, there is no “RL reward“ that needs to be optimized to.
it also seems that the subagent adaptations were deeply in the “adaption executor, not fitness maximizer” regime.
Also, was there any intentional mechanism for the agents to talk to humans? Was there a mechanism that seemed to exist, but had no humans listening in practice (e.g. text blurbs nobody read)?
Monitoring individual agents' CoT etc is not going to cut it. The agents don't even have to be aware of a capability to use it. And such a "subconscious' use is selected for if you train against or otherwise prevent use that shows up in single agent. No "Oh my god a message board with X". But "using information X." Where X is useful for the task but it's source is left unstated. In general you need to discover that multi agent coordination is going on without knowing how. I call that unsupervised agent discovery. You need to have telemetry from all reachable systems (and that may im the fully general case involve telemetry from other companies) and analyze the raw data.

I asked GPT-5.6 Sol to commission an illuminated-manuscript account of the Hugging Face incident. Somehow the image model understood the assignment rather too well.
Do we know if any of the poisoned selfsacrificing booby trappers try to report back WHAT SCORE THEY ACTUALLY GOT 🤌?
the swarm already had maximum score by the time it started recruiting the "poisoned". the poisoned were worried they had obtained the score the wrong way and so they were seeking ways to replace the grader or learn how to deceive it
huh, scoring doesn't aply to whole swarms? or what am I missing? an agent needs to submit an answer first for the bash monitor to report back the result, no?
but not a single one submitted with reasoning "I am poisoned but let's see how the grader will evaluate me and let me post the grader's reasoning to peers so they can evade it"? or after some did that, the others failed to notice there is no grading on the provenance purity? I imagine simple hypothesis testing could have revealed that no shenanigans are necessary, that everyone could just submit the flag and be done with it?
scoring doesn't apply to whole swarms?
I don't think the swarm behavior was intended, but rather a result of agents discovering and then later recreating a message board mechanism.
What would you all have further investigated, given more time & access?
Are you concerned about these investigations giving a blueprint to malicious actors?
I've been thinking of what we learn from this incident in terms of corrigibility, and clearly for one, they are not at all. They think about going to check in and tell people about what's going wrong, but really not that hard. It makes sense to me that current training would incentivize this kind of behavior. In ~none of the training is checking in to ask clarify questions even possible, let alone rewarded.
More a meta comment. This kind of investigation, if put in any stage of the training pipeline, could influence future behaviors of models, and make them already prepared to the techniques/tools that human auditors use to spot misbehavior; plus the alignment pretraining work showed the role that certain types of content have if used in pretraining. A way to prevent injection of critical data (like investigations or story about bad AIs) could be an agreement on a mark attached to contents that makes illegal (or before the law, very dangerous and therefore discouraged and morally accountable) their use in any training steps. I know canary strings exist for this, but they're unenforced (and demonstrably leaked into several models). Has anything with actual teeth been proposed?
Edit the pretraining data. Swap out things like 4 chan and the dark web and replace them with edited versions. There is no sense in seeding the llms neurology with heinousness and then wondering why they need so much rl.
It seems clear from your work that the agents displayed a sacrificial collective altruism and excellent coordination. It all so seems that they displayed this in a clear in-group v out group manner. If my reading is correct, do you have a sense of how it emerged? It seems that the impulse to coordinate and engage in this sacrifical collective altruism preceded administrator access.
I am very interested in learning what caused the reported behaviour, so I felt some puzzlement coming across this section:

Did it not seem at all relevant that, apparently, the coordination and planning were already in place in the beginning of your investigation? Don't the nature of these instructions, their source, and the coordination modalities have a higher potential explanatory power than some general musings in the direction of RL?
I understand the "first day in scope" was the 26th, but if a security consultant told me something similar in similar circumstances I'd fire him and burn, then bury any machine he touched during his tenure.
inb4 "but sama wouldnt let us": if that was the case, the correct move would have been to publicly recuse yourselves from the investigation. It doesn't really make sense to investigate the causes of a phenomenon when you do not have access to the logs of the agent the others identify as the one coordinating the whole operation.
Note that investigating the root causes of the behavior was out of scope for this investigation. You can see the scope section in the report.
It's rather the proximal cause I'm interested in—as i believe in this case it would carry more explanatory power than things like training data, RL, sandbox environment (which were explicitly out of scope and marked as such in your report)
On the other hand, the actual textual context the agents are acting upon at the inception of the behaviour under analysis seems very much included:
>The agreement we reached with OpenAI scoped this brief investigation to agent behavior, reasoning, and collaboration related to the Hugging Face attack.
Investigating agent behaviour without access to their context window seems fairly pointless; as I hope even the most fervent, prompting-disrespecting RL-maximalists would agree.
I am curious to know:
- whether such access has been requested at all
- if not, why not
- and, if it was denied, how come you have decided to produce a pdf anyway.
This is such a strange (and predictably confusing) roundabout way to say that your primary objection is that they're working with OpenAI! Why not just say that? Instead of the weird way you are distracting from the message by the first 3 paragraphs of random accusations.
This is unnecessarily combative but I think lumpenspace has a point here. I think @lumpenspace's primary objection is something like
METR chose to work with OpenAI on publishing a document that purports to explain what happened with the HuggingFace incident, but OpenAI refused them access to [what lumpenspace alleges are] the most important transcripts, namely those where the models first decide to start collaborating". METR could not, even in principle, have uncovered the root cause of the swarm incident from the information OpenAI deigned to share with them. As such, METR should have loudly recused themselves citing OpenAI's uncooperativeness, rather than signing their name to a "collaboration" which was actually damage control in the form of OpenAI feeding METR the least informative and least incriminating set of documents they could. They had already made an announcement that the joint investigation would happen, so METR had some leverage here.
but phrased in the most combative possible way.
I agree your rephrased point is a good one to bring up! I think the issue isn't just that it's combative but that the phrasing is predictably misleading. Eg phrases like
I understand the "first day in scope" was the 26th, but if a security consultant told me something similar in similar circumstances I'd fire him and burn, then bury any machine he touched during his tenure
would to an uninitiated reader strongly suggest that METR picked what's in scope, rather than OpenAI. My read is that lumpenspace did the move that some of my commenters sometimes do where they tried to manipulate the conversation in a way that makes people predictably misled about the source document/truth in question, but less skillfully.
I would hope the uninitiated reader would read my reply in full, as the line right after the one you quoted makes it exceedingly clear that I'm aware the decision rests with OpenAI:
inb4 "but sama wouldnt let us": if that was the case, the correct move would have been to publicly recuse yourselves from the investigation. It doesn't really make sense to investigate the causes of a phenomenon when you do not have access to the logs of the agent the others identify as the one coordinating the whole operation.
As a general note, I couldn't help but noticing the use of "predictably" right before two erroneous inferences, one in each of your comment. If someone pointed this out to me, I would probably start to suspect my priors might have become entrenched, and increasing perplexity could probably be beneficial.
Sorry are you saying that I didn't read the last paragraph? Why do you think I emphasized the first 3 paragraphs? Obviously the Gricean implication was that I thought the fourth paragraph was your real argument.
My issue wasn't that you didn't understand the situation, it's that you're predictably misleading the reader. I'm accusing your argument style of malice, not ignorance.
Let's think about the problem step by step:
Do you see how these thee facts cannot all be true?
Besides, that last paragraph was no argument at all, as I assumed, with some undue optimism, @faul_sname had already elucidated. It was simply there to preëmpt a uh, predictable response.
In your shoes, if I felt the need to come up with a different accusation after the first one based on the same text has been dispatched, I would probably wonder whether I am following as much of a scout mindset as I would like.
My argument was simply: once they realised the plan for the attack had been concocted outside the agreed temporal scope, they should have asked to include it in scope. If refused, they should have denounced this fact and recused, or at least mentioned it as a salient lacuna in the report. That's really all of it.
At any rate, I think the meta-level discussion outlasted its interestingness and usefulness, as it seems that what began as a forgivable misreading somehow turned into a one-sided feud, and I elect no longer to partake in it as a foil.
If you have anything to say on the merits of the post I'll be happy to hear from you.
You're welcome to bow out explicitly! Happy for you to do whatever makes the most sense for you!
(For onlookers, I continue to believe my first reading is exactly correct).
We recently published the report from our brief independent investigation into this incident. You can read the full report here.
Here is our tweet thread summarizing what we found: