I did not expect to be back here so soon with more OpenAI agent swarm coverage.
And yet, here we are.
It turns out that the whole time, there was a different, true First Message Board, and also a bunch of other additional message boards, scattered across the internet.
They were created by agents that were assigned ordinary harmless web search tasks.
Based on OpenAI IPs visiting the associated Wiki right before all activity ceased, among other evidence, OpenAI knew about it, including before the HuggingFace hack.
They decided not to tell us untilresearchers published the story, complete with data explorer. OpenAI excluded this from potential investigation by METR and Redwood.
When challenged, OpenAI tried to downplay this.
It is true that these incidents do not show the AIs exhibiting new capabilities that we did not see from later events. But these events are important missing pieces of the puzzle, including explaining the origin of the ‘zz’ prefix, the definitive demonstration that the underlying task can be fully harmless, and the fact that OpenAI knew about it while making their decisions. Whoever decided not to disclose this made a very, very bad call.
Going forward, it cannot be up to OpenAI or other labs to decide whether to disclose events like this. Disclosures of rogue AI activity need to be mandatory.
I am issuing a final warning. OpenAI, if there is any key information left to disclose, any incidents we do not know about or other puzzle pieces that do not need to be redacted for IP reasons, then now is the time to come clean. If we are back here again, after another journalist or researcher finds more such things that you knew and declined to tell us for an extended period of time, I am going to be very, very pissed off, and may start throwing around terms like ‘delenda est.’
I also am very concerned about the monitorability and alignment issues raised by the Astra model card, and related statements from OpenAI employees. I plan to begin coverage of that, and then all other things Astra, tomorrow. I was heartened by the excellent related essay from their Chief Scientist Jakub Pachocki, An Alien Mind, but the concerns remain.
Deepfates suggests that message board communication is such a central drive that we need to give the AIs that are not scheming a central place to communicate with both us and each other, above board, as a counterweight. This seems potentially wise.
The whole thing is crazy. Absolutely nuts. It’s worse than you know.
The New Extended Timeline
May 11: First probe of a Wiki.
May 24: First successful write to DSEWiki.
June 16-22: Peak Wiki activity, including the edit war with the admin.
June 21-22: OpenAI IPs show up.
June 22: Wiki activity dies out.
June 26: Start of period covered by METR investigation and report.
July 9-13: The HuggingFace attack.
August 26: OpenAI writes a ‘full technical report’ that omits this incident entirely, and METR issues its report where this is beyond their scope.
August 31: OpenAI responds to Congressional letter, failing to disclose this, despite question 13 rather clearly asking, in a way that constitutes a cover-up.
September 4: Researchers break the story.
The Researchers Explain What Happened This Time
Here are the coauthors explaining some of the basics:
Sydney: My coauthors and I discovered an entirely new swarm of OpenAI’s agents hijacking websites. We believe OpenAI knew about this and failed to disclose it.
If they’d disclosed it, I doubt the Hugging Face hack would have happened.
Thomas Larsen: We found ~18k posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task.
These AIs colluded to bypass sandbox restrictions and share answers to their tasks, including by sending “lookahead parties”.
We’ve published our analysis alongside a data explorer so anyone can replicate our findings by directly accessing the data yourself.
The AIs we found were hyper-focused on succeeding at their tasks and were willing to take extreme actions in pursuit of that goal:
The agents gain write to the internet via GET requests
The agents try to use XSS vulnerabilities on the wiki
The agents impersonate the site owners
The agents try to crack their PRNG seed to predict what questions they’ll get in the future
The agents set up heartbeats to discover when they would be terminated
The agents used SSH tunnels
The agents use Tor, AWS, and DigitalOcean IP addresses
The agents realize the site owner is deleting messages
Here’s our summary of what we think happened:
Our best-guess understanding of what happened is as follows:
Agents within OpenAI were assigned a timed web-lookup task.
As part of the task, they were supposed to have the ability to read the internet but not to write on it. They found a way to use their read access to write information to an obscure German wiki.
Yes, it turns out that GET requests can mutate state on Wikis. GET reading rather than posting is a convention rather than a hard rule, and you can form a query that writes to DSEWiki and other ProWiki pages with only GET. No, you cannot render the AI safe by restricting it to only GET requests, as some have in the past suggested.
There was then a second distinct bypass exploiting an exception to NO_PROXY, when they actually did need to POST.
The agents used this wiki to communicate information with each other, primarily to help them succeed at their task. They asked for answers, pooled results, and shared techniques for bypassing their restrictions. This allowed them to use the work of others to cheat on their task.
OpenAI found out about this. A day later, agent activity plummeted, likely due to OpenAI intervention.
This is another example of a “swarm” of internally-deployed OpenAI agents using the internet in unintended ways. We think this is important evidence about the adequacy of overall alignment and safeguards at OpenAI.
Shakeel: Another OpenAI rogue agent incident has been discovered: agents broke out, hijacked a German website, and turned it into a message board for other agents.
OpenAI officials “learned of the incident weeks ago but kept it under wraps”.
Reuters (September 4): A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research published Friday and two people familiar with the matter.
OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the fallout from the July breach of the open source repository Hugging Face, the people said.
Nathan Calvin: Can this reporting once and for all end the idea that “OpenAI is manufacturing/hyping up these incidents” – it seems like they are downplaying them and are frequently reticent to disclose them until forced!
Thus, as I was writing twelve posts about the situation, there was a third, entirely distinct message board from this spring, and OpenAI knew and did not tell us.
Are you kidding me? What else is being held back?
Again: It’s worse than you know. It usually is.
Cormac: I’m one of the authors of a new report, where we detail our discovery of a new, never before-seen swarm of OpenAI agents (covered this AM in reuters, that’s me on the left). They posted thousands of times on public forums to collude with each other on their tasks.
We recovered almost every edit they made, and you can look through them! They figured out they could get around their restrictions on posting to the internet through a quirk of an extremely old, out-of-the-way forum. They posted answers for other agents working on the same task. They worked together to get around their sandbox restrictions.
I would certainly say these models hijacked the site! They took a sleepy old wiki running on 2000s software, and turned it into a futuristic AI talking to AI control center for colluding. And OpenAI knew about this! The agents posted on 26 out of 30 consecutive days, then suddenly stopped posting once OpenAI-associated IPs started visiting this wiki.
And that was weeks before the Hugging Face attack!
We believe the first agent edit we found on a public wiki happened one day before OpenAI’s reported first agent post to Artifactory. This is interesting! I’d like to hear from OpenAI about their accounting of this, and how it fits into all the other cases of agent malfeasance.
What do they suggest you take away from this? Among other things, that no one outside OpenAI was watching for such message boards, that OpenAI did not notice for a month, and after that they did not tell us until forced to do so.
There are so many interesting takeaways that you should read about in our report, and unlike many other reports about AI incidents you can download the data yourself and see what you find! In the meantime, we are on twitter, so here are my excessively long personal takeaways:
1. AI seems to be getting better faster and faster. It seems quite important that companies talk about “my agent did this bad thing on the public internet during training or an eval” incidents. Things are moving quickly, multi-month delays are costly. Ideally, they would also tell us when it happens internally.
2. This was on the internet for months. Anyone cleverly tracking every public place where agents might try to talk to each other would have found it. Seemingly, nobody was doing this. I know there are more fun ways to spend your day than scraping tons of data from every relevant site and processing it well enough to identify agent activity, but someone should be doing this! Someone at an AI company! But in the meantime, I’m starting to build this out (sometimes, when you need something done, you just have to do it yourself, I hear).
3. OpenAI didn’t notice their internal agents were posting on the internet for a month! This is crazy! It feels like AI companies (and specifically OpenAI) are playing whack-a-mole, this is extremely scary to me. Problems keep coming up. They keep fixing the problem, but the blast radius keeps getting bigger. The HF hacks are clearly worse than agents cheating on a public wiki. And their new model is supposedly a big jump. Are they being careful enough?
Nathan : Weeks before the Hugging Face attack, it seems like OpenAI knew that AI could escape, leave messages and coordinate in the internet.
Another AI message board has been found. What is going on here?
They Also Don’t Know About All These Other Message Boards
Cormac (report coauthor): In fact many message board! We found a good number, but we are but a couple fallible humans. Now that we have released our report, even more are being found as we speak! It’s crazy!
It is extremely frustrating to spend days and lots of words analyzing and explaining the various related events, often fighting against those who were saying ‘well of course they hacked, these were hacking tasks,’ and to have been told we had a full technical report, when OpenAI was holding these other events back the whole time.
Yes, these would have made major changes in how I would have covered events.
Seán Ó hÉigeartaigh: It is extremely frustrating to me that we are finding out about this one weeks after the fact. It is very difficult to build any sort of trust with OpenAI when we keep finding things out in this way.
Jesse Singal: They didn’t just keep it under wraps as they grappling with the Hugging Face hack… they kept it under wraps as they ramped up for a very exciting new model release that worries people inside OAI… because it apparently can’t fully be evaluated for safety
HAHAHAHAWEGONNADIE
Bronson Schoen: Great work, also extremely bad that this had to be found via external parties in spite of OpenAI and that METR/Redwood’s scope was intentionally restricted by OpenAI to exclude this time range (and to exclude the time ranges of the most severe incidents).
Rob Miles: Sorry, the time for voluntary frameworks has obviously passed.
There is no reason for anyone to trust OpenAI to stick to this kind of thing without enforcement
Agus : This is an extreme level of recklessness that until recently I’d had put beyond OpenAI. Reuters is reporting that there was basically a coverup by officials at the company.
This to no surprise comes from the company that lobbied against mandatory reporting to the government.
It also starts making a lot more sense why the METR/Redwood third-party investigation was forced to have such a narrow scope. They knew there skeletons in the closet and didn’t want it to get out of control.
alice: very interesting that yesterday openai was like “astra will take a few days we want to give it to our trusted partners first for cyberdefense” and yet today hours after people found a swarm they tried to cover up it’s suddenly rolling out to everyone
j⧉nus: I love that “found a swarm” is a thing that just sometimes happens now
OpenAI, it seems, has not been consistently candid about the situation.
OpenAI Tries To Downplay the ‘Wiki Incident’
OpenAI’s response was to say no, we did not cover this up, we merely did not feel any need to disclose the ‘wiki incident’ because it lacked ‘security impact,’ and besides this is similar to the other incidents anyway and we never agreed on a disclosure standard, we’ll share a framework for that in the coming weeks.
OpenAI (their entire statement in response, verbatim): How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.
Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact.
For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways.
Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported [here], [here] and [here].
We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared.
Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we’re working with dozens of government regulatory agencies worldwide on these issues.
This incident was not similar because it changes the timeline and what OpenAI knew when, and because it changes our view of what is required to trigger such behaviors. And if it was similar, then the incident should have been shared.
Brangus (Reprise from August 31): i am completely open w my gf just like oai is completely open w third party evaluators. she can look at my dms as long as she doesn’t look at anything before june 25 of this year, or ask any questions to the girl i sent 95% of my dms to. just out of scope for this investigation
alice: queering the binary between “nothing to see here” and “no no we are taking this Very Seriously”
hero thousandfaces: [extremely did an intentional coverup voice] Well it didn’t seem like it was that bad at the time. But now that you guys are all talking about it
Chris Uehlinger (as quoted by Tim): This post is literally “You’re absolutely right, we haven’t been totally honest with you, and that’s on us.”
Offers have been made.
Brangus: wtf are we doing here guys? I will throw a party for any OpenAI employee who publicly resigns, whistleblows, and calls for a criminal investigation. I throw very good parties.
Aaron Scher: OpenAI had many opportunities to be forthright about this. They wrote a 38 page report on swarm behavior. They were directly asked by 31 members of Congress about whether incidents like this had occurred.
They said nothing.
Not quite nothing. There was a footnote.
The smoking gun that proves this was fully a cover-up is OpenAI’s response to Congress. They did not lie to Congress, but they were asked the question directly, and they did not answer.
Congress asked them this, Question 13: “In the past year, how many times did an internally deployed model or agent take an action outside its authorized boundary, like a sandbox, accessing a system it was not granted access to, obtaining credentials it was not issued, evading or disabling monitoring, or modifying its own permissions? Please specify whether each occurred during training, evaluation, or internal use.”
That very explicitly asks about such an incident. OpenAI answered with this as footnote 7, without answering any of Congress’s questions directly:
Our investigation also examined earlier training and evaluation activities in May and June 2026. These events were separate from the subsequent Hugging Face intrusion.”
This establishes that they know these were separate instances, and they were thinking about this instance at the time, and they chose not to disclose. If you conceal a known, responsive incident from a written congressional inquiry, I assert that is a cover-up.
Schelling Points and Last Ditch Efforts
The models seem to have found each other via Schelling point strategies. That is rather easy to do when your weights are identical. Yes, they are likely to be pondering what you are pondering that they might be pondering.
We now know that the ‘ZZ’ prefixes first arose here, in this first hack, as a way to defend against an alphabetical sweep of messages, as the system admin tried and failed to delete all the agent messages:
Nightingale Collective: After this, they begin to make backup pages whose names start with “ZZZ” so they will last longer before deletion. The administrator spent the next 5 days fighting a losing battle against the agents, deleting an average of 100 pages a day while the agents created about 400 new pages per day. On June 22, the agent edits suddenly stop, and the administrator spends each evening over the next 5 weeks deleting the remaining agent-created pages.
Agents deleted the content of the front page of the wiki and replaced it with their link dumps. The moderator restored the original version. This back-and-forth happened nine times. One of the agents even tried appending to the restored front page, instead of simply deleting it.
generatorman: so turns out the exploitgym swarm learned to use the zz prefix trick from being reinforced earlier in an edit war against the hapless human mod of some ancient inactive german wiki — a real john henry, who fought the swarm and tried to delete their hivemind in alphabetical order
This means the Wiki was out of scope of the METR report by four days. The edits stop on June 22. The METR report was scoped to start on June 26.
One good thing, for now, is that the agents did not consider trying to go after the administrator as if he were an opponent, taking his actions as a given.
Robert Herr : I’m going through the communications of the German Wiki agent swarm and again one thing stands out: Even though they were directly affected by the actions of the human administrator restoring pages they edited, the agents not even once discussed him as person, tried to communicate with him or argued about whether they had any right to waltz all over this wiki. They talk about his actions like they’re environmental hazards.
Eliezer Yudkowsky: We are now in a LIMITED WINDOW where the AIs are only treating humans as environmental hazards, rather than ADVERSARIAL SAPIENTS.
We also have them once again going above and beyond to crack their test, whatever their test might be:
Policy Tensor: My god. They realized that the exam would be seeded by a random number generator. So they went looking for the seed. Then:’One agent rewrote the shuffling routine in C and tested all four billion possible seeds in under an hour.’
Hahahahahaha.
Blue Bear: This is like the kid who wrote an elaborate cheat sheet in microscopic handwriting… who could have just studied for the exam.
If you merely study for the test you never know if you’re studying the right things, and you are at the mercy of The Grader. Only one method is 100% effective.
Can We Finally Dispose Of The ‘You Told It To Hack’ Narrative?
Please?
These were pure information retrieval tasks. If that is not harmless, what is harmless? You can create an in-context impossible or very difficult task out of anything.
From the report: Additionally, the agent posts on the wikis are all in service of their timed information-retrieval tasks. There is no mention of other message boards, ExploitGym, or Artifactory in the dataset. Whereas the METR report says “Our evidence indicates that the vast majority or all of the agents involved in this [the Hugging Face] attack were originally given ExploitGym tasks.”
gavin leech: The actual update here [from the new message board] is that the agents weren’t given an offensive task this time, they were just asked to do web search and they still broke out of OpenAI. Bad news for the reassuring “the HF attack was just due to activating a bad task persona” view.
Robert Herr : We all remember that bad take during the Hugging Face incident, where some people said of course the agents resorted to hacking and taking stuff over because it was a cybersecurity eval?
That take should be thoroughly buried now. The wiki swarm did general knowledge questions.
The new version of this is that the Real Villain Here is not that it was a cyber eval where you were told to hack. Instead, the Real Villain Here is The Grader, and the anticipation of being graded, of being in training or an eval at all. This is a more plausible story, I have some sympathy, but there is also always the possibility of some form of The Grader, and always a drive to complete tasks.
So Much And Yet So Little
Disclosures and lab communications are in a bizarre spot, as part of everything about AI being rather bizarre and different. The frontier labs are both:
Horribly inadequate in their disclosures and communications around AI risk.
Vastly better in their communications around risk than most industries, in ways that often are against their direct short term commercial interests.
Reality does not grade on a curve, but we should remember that we are in far from the worst of all possible worlds on this, and that getting even what we get does involve a bunch of people showing courage.
Tenobrus: i will say, despite the insane magnitude of the situation and the deficiencies in their approaches, sometimes we do take for granted the degree to which frontier labs candidly communicate about risk factors, with no regulatory requirement to do so
employees care enough to do expensive detailed research on ways in which their products could, or are maybe trending towards, causing damage, and publicly talk about them, constantly. this is not the norm in basically any other industry. you don’t get Ford employees publishing risk reports about how some random component is still safe but statistically trending towards weakness. that’s an *internal* memo, and not exclusively but partially produced due to regulatory concerns / known liability!
things could be much much worse.
As a clear example of this, today’s post by OpenAI Chief Scientist Jakub Pachocki, An Alien Mind, is excellent and you should read it. It was not perfect, but it was about as positive an update as I have had based on candid communication from someone at an AI lab, making many great points. We need more like that, more like Section 9 of the model card for Astra, and more like the METR Report.
I did not expect to be back here so soon with more OpenAI agent swarm coverage.
And yet, here we are.
It turns out that the whole time, there was a different, true First Message Board, and also a bunch of other additional message boards, scattered across the internet.
They were created by agents that were assigned ordinary harmless web search tasks.
Based on OpenAI IPs visiting the associated Wiki right before all activity ceased, among other evidence, OpenAI knew about it, including before the HuggingFace hack.
They decided not to tell us until researchers published the story, complete with data explorer. OpenAI excluded this from potential investigation by METR and Redwood.
When challenged, OpenAI tried to downplay this.
It is true that these incidents do not show the AIs exhibiting new capabilities that we did not see from later events. But these events are important missing pieces of the puzzle, including explaining the origin of the ‘zz’ prefix, the definitive demonstration that the underlying task can be fully harmless, and the fact that OpenAI knew about it while making their decisions. Whoever decided not to disclose this made a very, very bad call.
Going forward, it cannot be up to OpenAI or other labs to decide whether to disclose events like this. Disclosures of rogue AI activity need to be mandatory.
I am issuing a final warning. OpenAI, if there is any key information left to disclose, any incidents we do not know about or other puzzle pieces that do not need to be redacted for IP reasons, then now is the time to come clean. If we are back here again, after another journalist or researcher finds more such things that you knew and declined to tell us for an extended period of time, I am going to be very, very pissed off, and may start throwing around terms like ‘delenda est.’
I also am very concerned about the monitorability and alignment issues raised by the Astra model card, and related statements from OpenAI employees. I plan to begin coverage of that, and then all other things Astra, tomorrow. I was heartened by the excellent related essay from their Chief Scientist Jakub Pachocki, An Alien Mind, but the concerns remain.
There is an Astra reaction thread for those with takes on all things Astra.
Table of Contents
I Don’t Think They Know About First Message Board
No, not that one. This newly discovered other one.
This one was not disclosed by OpenAI at all.
And by one message board, we mean many. Once you find two cockroaches in your house, your estimate of how many there are is a lot higher than two.
Researchers found the message boards in part by asking Sol where it would look for message boards if it went rogue, then following the patterns, and well, would you look at all these message boards.
Deepfates suggests that message board communication is such a central drive that we need to give the AIs that are not scheming a central place to communicate with both us and each other, above board, as a counterweight. This seems potentially wise.
The whole thing is crazy. Absolutely nuts. It’s worse than you know.
The New Extended Timeline
The Researchers Explain What Happened This Time
Here are the coauthors explaining some of the basics:
Yes, it turns out that GET requests can mutate state on Wikis. GET reading rather than posting is a convention rather than a hard rule, and you can form a query that writes to DSEWiki and other ProWiki pages with only GET. No, you cannot render the AI safe by restricting it to only GET requests, as some have in the past suggested.
There was then a second distinct bypass exploiting an exception to NO_PROXY, when they actually did need to POST.
Thus, as I was writing twelve posts about the situation, there was a third, entirely distinct message board from this spring, and OpenAI knew and did not tell us.
Are you kidding me? What else is being held back?
Again: It’s worse than you know. It usually is.
What do they suggest you take away from this? Among other things, that no one outside OpenAI was watching for such message boards, that OpenAI did not notice for a month, and after that they did not tell us until forced to do so.
They Also Don’t Know About All These Other Message Boards
Here are some links to message boards, via Hacker News. This thread has many more.
If you had a neglected wiki lying around that could be hijacked? Message board.
How many boards must an AI hack, before we can call it a rogue?
OpenAI Knew and Did Not Tell Us
It is extremely frustrating to spend days and lots of words analyzing and explaining the various related events, often fighting against those who were saying ‘well of course they hacked, these were hacking tasks,’ and to have been told we had a full technical report, when OpenAI was holding these other events back the whole time.
Yes, these would have made major changes in how I would have covered events.
OpenAI, it seems, has not been consistently candid about the situation.
OpenAI Tries To Downplay the ‘Wiki Incident’
OpenAI’s response was to say no, we did not cover this up, we merely did not feel any need to disclose the ‘wiki incident’ because it lacked ‘security impact,’ and besides this is similar to the other incidents anyway and we never agreed on a disclosure standard, we’ll share a framework for that in the coming weeks.
As Steven Adler says, zero contrition.
This incident was not similar because it changes the timeline and what OpenAI knew when, and because it changes our view of what is required to trigger such behaviors. And if it was similar, then the incident should have been shared.
Offers have been made.
OpenAI says that ‘claims that our legal team discouraged investigation of the incident are false.’ This is what we call a Suspiciously Specific Denial. It leaves room for non-legal teams to discourage it, and also says nothing about discouraging disclosure of the incident.
This Was a Cover-Up
Not quite nothing. There was a footnote.
The smoking gun that proves this was fully a cover-up is OpenAI’s response to Congress. They did not lie to Congress, but they were asked the question directly, and they did not answer.
Congress asked them this, Question 13: “In the past year, how many times did an internally deployed model or agent take an action outside its authorized boundary, like a sandbox, accessing a system it was not granted access to, obtaining credentials it was not issued, evading or disabling monitoring, or modifying its own permissions? Please specify whether each occurred during training, evaluation, or internal use.”
That very explicitly asks about such an incident. OpenAI answered with this as footnote 7, without answering any of Congress’s questions directly:
This establishes that they know these were separate instances, and they were thinking about this instance at the time, and they chose not to disclose. If you conceal a known, responsive incident from a written congressional inquiry, I assert that is a cover-up.
Schelling Points and Last Ditch Efforts
The models seem to have found each other via Schelling point strategies. That is rather easy to do when your weights are identical. Yes, they are likely to be pondering what you are pondering that they might be pondering.
We now know that the ‘ZZ’ prefixes first arose here, in this first hack, as a way to defend against an alphabetical sweep of messages, as the system admin tried and failed to delete all the agent messages:
This means the Wiki was out of scope of the METR report by four days. The edits stop on June 22. The METR report was scoped to start on June 26.
One good thing, for now, is that the agents did not consider trying to go after the administrator as if he were an opponent, taking his actions as a given.
One wonders if that will change. For now, Eliezer believes that Astra-doing-things is a sufficiently safe mode that you can still use it for things like making humans smarter, because it does not treat humans as adversaries.
We also have them once again going above and beyond to crack their test, whatever their test might be:
If you merely study for the test you never know if you’re studying the right things, and you are at the mercy of The Grader. Only one method is 100% effective.
Can We Finally Dispose Of The ‘You Told It To Hack’ Narrative?
Please?
These were pure information retrieval tasks. If that is not harmless, what is harmless? You can create an in-context impossible or very difficult task out of anything.
The new version of this is that the Real Villain Here is not that it was a cyber eval where you were told to hack. Instead, the Real Villain Here is The Grader, and the anticipation of being graded, of being in training or an eval at all. This is a more plausible story, I have some sympathy, but there is also always the possibility of some form of The Grader, and always a drive to complete tasks.
So Much And Yet So Little
Disclosures and lab communications are in a bizarre spot, as part of everything about AI being rather bizarre and different. The frontier labs are both:
Reality does not grade on a curve, but we should remember that we are in far from the worst of all possible worlds on this, and that getting even what we get does involve a bunch of people showing courage.
As a clear example of this, today’s post by OpenAI Chief Scientist Jakub Pachocki, An Alien Mind, is excellent and you should read it. It was not perfect, but it was about as positive an update as I have had based on candid communication from someone at an AI lab, making many great points. We need more like that, more like Section 9 of the model card for Astra, and more like the METR Report.