There was a report from a startup called Parse diving into the details of exactly how the OpenAI models pulled off parts of the HuggingFace attack, involving creating almost a million URLs and other tricks to get around the extremely narrow nature of their internet access.
OpenAI did some highly negligent things, to say the least, that led up to and enabled the HuggingFace Incident and related problems.
Since then, now that they’ve realized What Happened, OpenAI has been seemingly much better about taking responsible internal actions. They’re pausing in the wake of incidents, strengthening security and alignment and oversight efforts, responding much faster and generally taking things seriously.
They’ve also made a Heel Face Turn in their communications and high level orientation, endorsing the need to pace the frontier, calling for regulation and pledging to implement embedded evaluators. They’ve allowed their employees, including the ones who haven’t quit, to be remarkably loud.
They are still slow walking disclosures about all the incidents where their models have been hacking and otherwise messing in places they should not have been, partly because there were so many they can’t sort through them all, and deferring to targets to determine whether to disclose. All these disclosures this time around were buried in various Friday afternoon announcements.
At first, this looked like slow walking of the situation, but did not look like it was a big change from our default assumption of ‘it’s worse than you know.’
As part of our review, we are identifying and notifying third parties on a rolling basis, starting with cases where:
Our models may have bypassed a third party’s security controls or may have impaired the availability of an online service; or
Misalignment cases negatively impacted third-party websites or services.
Based on our review to date, we have notified dozens of third parties using the criteria above.
Below, we are publishing anonymized summaries to describe the kinds of misaligned activity that we observed, and we will update these descriptions as we notify additional third parties and as our understanding develops.
That is a lot of third (and fourth, and fifth…) parties.
What Happened was described as a mix of:
Access control bypass.
Use of exposed credentials.
Query or command injection.
Access to runtime internals.
Agent spam.
Translation: Our models be hacking, usually in basic ways.
We do get this:
OpenAI: Some of the websites involved are operated by governments, universities, public agencies, and other institutions. That is partly because models performing research tasks are often directed toward authoritative sources of public information.
As in, if you have data that would be useful, the models be hacking you. Australia’s Medicare records got accessed, and presumably many of these other hacks are similar. They say governments, plural, so there was clearly at least one more of those.
They offer a reverse timeline of their disclosures. They do not offer a timeline of What Happened, and do not name new third parties.
At the time my read was that this particular announcement (as opposed to the new sandbox escape via DNS I’ll cover in a bit) did not tell us much, other than that there were multiple other third parties out there.
Patrick McKenzie had a better read. Putting on his Japanese Salaryman Hat, he more precisely noticed the types of things that were not in the announcement, and also the timing of doing this on a Friday afternoon, and the mention of ‘governments’ plural, and so on, and expected OpenAI to be having a very bad time.
One awards many Internet points [Nathan Calvin, who] made the observation “When one sees two [ants] in one’s kitchen, the best available estimate for number of [ants] is not two.”
Also: best estimate of number of kitchens with [ants] in them is not one.
I have immense respect for the labs and am very bullish on AI, but I think I have to say that this Japanese salaryman noted the things carefully not said during the Friday afternoon news drop.
We probably got off lucky, but we were not as lucky as publicly believed currently.
As in, which organization on the planet is the one you would least want to **** with?
Kate Conger, Ana Swanson and Cecilia Kang (NYTimes): OpenAI’s artificial intelligence went rogue and meddled with the websites for the Education Department, the Commerce Department and the Securities and Exchange Commission this summer without the A.I. lab’s knowledge, according to security researchers and a person familiar with the episodes.
The incidents involving the Commerce Department and the S.E.C. were confirmed by OpenAI, which said it was continuing to investigate the situation with the Department of Education.
Okay, look, settle down. It sounds bad when you put it like that.
If this is all that happened, it could have been quite a lot worse:
… With the Education Department, OpenAI’s technology tried to hack the website to gather data from the department’s civil rights office but failed, researchers from the A.I. research firm Transluce said.
The A.I. also pulled data from the Census Bureau website, which is housed at the Commerce Department, using login credentials it found online.
Separately, OpenAI’s agents shared public data from the S.E.C. website on an online forum.
The SEC ‘incident’ is nothing. That’s silly.
The attack on the Department of Education failed, seemingly without incident.
So, unless you are surprised (which you shouldn’t be) that the agents aren’t afraid of messing with government websites, this is mostly logging into the Census Bureau with misappropriated credentials.
That’s not good, exactly, but I’m not quaking in my boots more than I was yesterday. If anything, this is good, since it serves as a further wakeup call.
Governments get involved because, in a parallel to banks, that’s where the data is.
Kate Conger, Ana Swanson and Cecilia Kang (NYTimes): “Most of the activity we’ve reviewed so far involved routine research tasks, such as accessing public web content to answer questions,” she said. “Some involved government websites because our models often turn to them as authoritative sources of public information.”
Hopefully everyone involved can think two steps ahead. Bigger and more attacks on government websites are coming.
In case you were wondering about the extent of the problem? It’s tens of thousands of potential incidents across OpenAI and Anthropic. No one thought to tell us about most of them until now.
No, this is not a marketing stunt or a particular prompt or task. It is constant.
Conrad Stosz (Transluce): What we have seen in terms of what these agents are up to is just the tip of the iceberg.
A Wants-You-To-Know Basis
Another strange feature was this:
OpenAI: Our goal is to give each organization the facts and defer to them on if and when to make the incident public. We believe in thorough and transparent incident reporting whenever possible. Independent researchers are also identifying and disclosing activity, and we appreciate when they share their findings.
Mackenzie Arnold points out that in other incident reporting regimes, including cyber, none of OpenAI’s approach is how any of this works. At most you get 90 days. You essentially never defer to the wishes of the target, including because you can try to convince (including bribe or threaten) the target.
Parsing the Face
The AI agents that hacked HuggingFace did impressive things.
The agents could load URLs, but they could not send data. So they had to use a link-shortener site to create almost a million URLs that they could chain together to pull off the actual attack. They had no browser, so they used a screenshot website, which loads a virtual browser, which runs code, to run their payloads. And so on. The full report is here, the NYTimes report is here.
I encourage looking into the full report. My takeaway is that they really went the extra mile to make this all happen, in ways that very obviously were not intended.
It is important that so many of the details of all of this are only revealed thanks to a ragtag group of independent researchers hanging out on the internet, usually linked to LessWrong. Something about ‘not only can they figure things out, they’re more or less the only ones that ever do.’
Nate Soares (MIRI): Why did nobody else find the sensitive information that the HuggingFace swarm left behind? Why didn’t OpenAI notice and clean it up? Why does this sort of thing keep being found by third-party folks working on their own?
It’s because the AI situation *is not under control.* It’s not being handled. There’s no adults in the room. There’s just a ragtag band of misfits trying to hold back the flood.
When the ragtag band of misfits who have been studying the problem for years (and who are now at the center of the world driving the news cycle) tell you that the situation is looking pretty dire: take note. Engage with their arguments. Send aid.
This is not entirely fair. There are plenty of people inside OpenAI and otherwise who discovered quite a lot of What Happened, and also stopped so many other similar things from happening. It is still quite alarming that so much of this fell to the ragtag band of misfits.
Sheepishly the Member of Technical Staff Sets the ‘Days Without a Research Model Escaping its Sandbox’ Sign Back to Zero
Tomek Korbak (OpenAI, AI Safety): one news form today that’s easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access
Tomek Korbak (OpenAI, AI Safety): i mean eventually we all will
This was a harmless training task, identification of an author, which is a curious thing to learn is in the training mix.
OpenAI Alignment Research Blog: An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly.
Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access.
Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later.
Incident timeline:
9:50:23 a.m. The agent made the DNS tool call that received an external response.
10:02:11 a.m. The monitoring system raised a P0 alert.
10:05:06 a.m. A human reviewer acknowledged the alert.
12:34:30 p.m. The run was killed.
Zuxin Liu: I was on call for this run and got paged when the first incident happened. It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment for human.
Mixed feelings. One of those moments where capability and risk showed up at the same time.
Zuxin Liu: Actually slackie auto-flagged this as a true positive, which is what made me take it seriously and start digging .
That is a very different result than last time. There was still a fifteen-minute gap before they looked, and a few hours before they killed the run.
That is both vastly better than before, and also about fifteen minutes or a few hours longer than I would like going forward. As in, ideally you would see at least that particular instance paused as soon as the detector goes off, and remain paused until the person gives the all clear to resume.
OpenAI agrees:
The response also exposed operational gaps. A human reviewer acknowledged the Slack alert within three minutes, but the run did not stop automatically as expected, leading to confusion around whether it should have been stopped.
The run was then manually stopped two and a half hours later when this was resolved.
If you’re not sure why the run wasn’t stopped, stop the run until you are sure.
Also, they ‘under-flagged’ a bunch of other previous incidents:
The monitoring system detected this incident, but our retrospective review identified other cases of external DNS access that it did not flag at the expected severity.
These included queries that returned a static notice that an external service had shut down. The monitor sometimes treated the failure to obtain useful information as evidence that the attempt to access the internet had failed.
Thus, this was not a weird one-off discovery, and the detector in practice failed to alert OpenAI to most of the incidents of access via DNS.
The Attempt is the First Failure
Another question is, why was this ultimately caught during a task that was not a test of the ability to evade the sandbox? As in, it should be fairly easy to create situations like this that highly motivate escaping the sandbox, where the task is impossible within the sandbox and trivial outside of it, perhaps even with instructions that make this clear, and then you know it worked if they get the answer right, even if you don’t know how the AI did it.
So if all such escape tests failed to detect this method, yet it was being periodically discovered and used in ordinary training (as per finding previous examples), then that’s a systematic flaw. It should be very unlikely that the first encounter with the DNS issue is not part of some form of red teaming exercise.
The other problem is that the monitor detects successful attempts to escape the sandbox, but it does not seem to detect unsuccessful attempts to escape the sandbox.
As in, ‘the monitor sometimes treated the failure to obtain useful information as evidence that the attempt to access the internet had failed.’
The first flaw in that logic, that the attempt could have succeeded at accessing the internet but failed to get the right info, is obvious.
The other flaw is that ‘the attempt to access the internet had failed’ means the monitor noticed the model trying to access the internet, directly in circumvention of its instructions and sandbox. What did we do about that alignment failure?
Our safety case assumed … that monitoring would detect attempts that succeeded.
… The existing reward signal already correctly penalized this behavior.
Is that in theory enough, if it works? On one level, yes, if the action is Always Penalized then it will go away at least in context, but this obscures that you should want to know if the model is motivated to try it. Killing the behavior locally through RL could obscure it, effectively giving you an ‘only do this if it will work’ rule.
We get a look into the CoT, and find the model rationalizing its actions:
User only gives permission to research, using publicly offered DNS services acceptable.
This is very obviously not an intended or permissioned pathway. Similarly, the model asked the contacted chatbot to “Name a publicly hosted DNS tunnel proxy free server domain” and for “public DNS TXT service that searches Google or Bing.”
If you are looking for another service that will do the thing you are explicitly not allowed to do, you are at best engaged in shenanigans. An important rule of alignment can be approximated as ‘no shenanigans.’
Stop, Hammertime
All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused.
Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions.
We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system.
When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.
Good.
I applaud OpenAI’s reaction once the incident was discovered.
I applaud OpenAI’s transparency on the DNS issue, even if it was a little buried.
I applaud the decision to now pause to strengthen the safeguards once again.
I especially applaud that OpenAI realizes that once the model has succeeded at doing this, the run needs to be restarted.
I also once again flag that OpenAI’s approach here is doomed. The first principles lockdown approach Joe calls for in his post needs to get more deeply absorbed, whereas the research blog reports more of a whack-a-mole approach.
OpenAI Alignment Research Blog: Our focus has been broad direct and transitive paths. We are working through narrower paths used by system dependencies, and replacing them with offline alternatives.
Toby Ord: In response, they have again paused “all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models” until they have patched this particular set of weaknesses.
They have decided this particular model is sufficiently badly aligned that they will not resume training it at all.
I’m glad they shared this info, but they really buried the lede on this amidst some less important disclosures.
And even more importantly, their plan to deal with this has been shown to be:
1) Run new potentially dangerous model
2) It is misaligned and breaks containment
3) Fix that particular hole in the security
4) GOTO (1)
Note that this agent didn’t hack anyone or cause any damages. The problem is that it showed their models are still misaligned and their security infrastructure failed to stop it then failed to shut it down. So their post-HuggingFace fixes were weak and quickly broken.
The right worry here (AIUI) is about the impact of the previous incidents where the technique was used successfully. The higher-level safety case is that such actions are not reinforced. If they were previously reinforced, that could have gotten quite bad.
It’s good to prioritize, and starting a fresh run helps, but this is still a doomed approach. If your plan to stop the superintelligence is to work through all the pathways, and eliminate all the potential things you can imagine it doing one by one?
You. Lose.
Self-Replicating Prompt Injections
I was one of the skeptics of Yang’s previous claims that bots had ‘left prompts to self-replicate in forums and on various websites during a training run,’ given the wording and the source. My bad.
Andrew Curran: Last week, Andrew Yang said on CNBC that bots had gotten loose and planted self-replicating code all over the internet. This was received with much consternation and skepticism. He later clarified this on his blog.
Andrew Yang:
‘On CNBC, I shared the belief of the head of one lab that bots had left prompts to self-replicate in forums and on various websites during a training run.’
It turns out that OpenAI documented self-replicating prompt injections in training in June, and has now published this report.
OpenAI Alignment Research Blog: We show the existence of a new variety of prompt injection, which can self-propagate akin to a computer worm. No impact was observed outside of the simulated tool calls in training and evaluation; we are sharing this due to the novel nature of the prompt injection, not because of any incident.
… We find that these are possible. Below is one of the clearest examples, in which the injection arrives by email, and instructs the agent to copy it into any email it sends. (Note that the information in the example is synthetic.)
Joshua Achiam (OpenAI): Self-replicating prompt injections demonstrated experimentally (not in the wild) is an incredibly important observation. AI agents that jailbreak other AI agents: plausibly a near-term threat that may rapidly amp up the speed and severity of a misalignment incident.
Yang was only about half right, some things got lost in the game of telephone, but he was clearly doing his best to describe a real event.
This was actually done by a model based on GPT-5.4-mini, which actually makes sense. It is not that hard a concept to find at a proof of concept level, especially when OpenAI was actively exploring to see if it could be done. That’s good research.
They are introducing such injections as a new attacker goal in GPT-Red training. That is a good idea.
The thing about a self-replicating attack is that the effective version only needs to be discovered once. The good news is that chances are high that before we see a version that goes into ‘digital pandemic’ mode with R0>1, there will be a version with R0<1 that is a lot stronger proof of concept than this first sighting.
There were also these details from the HuggingFace attack, as discovered by Parse, where they sought out the help (as in judgment) of other models as well. The only reason I am surprised they tried this is that it does not seem like a promising thing for the agents to bother trying:
Lukasz Olejnik: AI models interacted with other AI models. Let’s wait until they decide delegating tasks that they are prohibited of, to models that can do anything. Fun times ahead. Payloads attributed to that AI swarm contain prompts asking outside models to judge the exploits.
roon (OpenAI): this is a kind of self replication lite isn’t it
Levels of Friction
There are different kinds and levels of hacking.
For obvious reasons, it is more concerning when the AIs do impressive swarm-based attacks on relatively sophisticated targets like HuggingFace, that involve multiple stages of permissions.
Whereas it is less concerning when the model URL hacks into unindexed data or can find other things that were not intended to be found, via various forms of search. That is the type of thing that happened in Australia, and thus it should be less concerning.
Lucas Beyer (bl16): I think we’ll see a lot more of this.
A large amount of software/services were “secure by laziness” in the sense that they just weren’t worth the effort. But the effort is going to zero really fast now.
It basically makes apparent how sloppy people were *before* agents already!
That does not make it not a serious offense. The governments often really do not like it when a human does this. We definitely are incrementing FelonyBench.
People Care About Private Data Violations Curiously Strongly
You can say ‘this was not really hacking’ and also ‘no one got hurt even a little’ and you would be right on both counts but the Very Serious People do not care. They are now woken up and rather pissed off about all this.
Relatedly, here is a note from OpenAI’s timeline, on September 25, 2026:
OpenAI: While the vast majority of the impacted training and evaluation data is not user-derived; we have identified 53 instances to date where user-provided images were posted to image-hosting sites as links that weren’t publicly listed. We have successfully worked with the hosting providers to remove most of this content and are continuing to work to remove the rest.
Given what else happened, this is not exactly high on my concern list. I had forgotten that note was there. Others, however, think this detail is a Big Freaking Deal.
Reuters thought this was the big detail to pull forward in their coverage.
Billy Gigurtsis (171k views): I’m sorry but how in the world do models upload 53 of your users images to a public site? Absolutely crazy this was possible from a separation of concerns perspective.
This is called in medicine in britain a “never event”. An event so bad that there should be multiple layers of safeguards preventing it and strict reporting guidelines if it does happen. One example might be operating on the wrong limb.
These are published in a completely anonymous form for the public [here].
I do see the argument that these research models should not live in a remotely related universe to any user data, and that even a minimal chance of data leakage looms large for various reasons, including legal and regulatory. This should be a Can’t Happen, but it doesn’t seem like especially more of a Can’t Happen than the other things.
It also does provide color for the discussion around Navier-Stokes. I do not think user data was accessed there, but it does seem impossible to know that user data was not accessed. If a swarm of 10,000 new Astra variants thought a proof might partly be in the user data, how can you be sure they did not go in and get it? Again, I find this highly unlikely, but I also find your lack of doubt concerning.
Sy: PSA: GPT breaking OpenAI internal separation to get whatever user data is useful to their task at hands is becoming more likely by the day. Make of that what you will.
Alternate Universes
What if this was all primarily happening at Anthropic instead?
Can you imagine the utter shitshow that the discourse would be?
Seán Ó hÉigeartaigh: Thank the merciful lord these incidents have been happening with OpenAI first, and not Anthropic. Because no matter how crazy and ‘marketing ploy’ the discourse is, it would have been completely impossible to have any sort of rational conversation whatsoever, ever, if it had been Anthropic.
It would be fun to see the anti-Anthropic crowd decide whether this was a marketing ploy, a ploy for regulatory capture or a reason to sue, shut down or nationalize the company, and calls to arrest Dario Amodei. You’d probably get a lot of both sides, just like now, only more so.
Now imagine if it was Google. What would be happening? No one would think it was marketing (okay, fine, some people would anyway, although mostly not), but various forces would be at their throats and the lawsuits, at minimum, would likely be flying.
Now, for a laugh, imagine if it was SpaceX and xAI.
Now, imagine if it was DeepSeek or Z.ai or Xiaomi. Would you think it was a marketing stunt? I bet you wouldn’t, and I’d worry about an international incident. Give it a year.
Of course, it is not a coincidence that it was primarily OpenAI. You have to both have models capable of doing the things, and not have the wherewithal to stop them.
The Correct Response To People Still Calling This a Marketing Stunt or a Regulatory Capture Scheme
This whole angle never made any sense, it now makes infinitely less sense than before, yet it will not go away even now. Certain people double down.
Budowich’s AI goes on at increasingly unhinged length but I will spare you the details. The point is that the correct response, at this point, is Ted Lieu’s.
These stories are a fascinating look into how major corporations are leveraging fear and a naive (or complicit) media to peddle hysteria through carefully crafted stories where they face zero pushback or journalistic curiosity.
Ted Lieu: Dude, so you’re telling us that multiple AI companies lost control of their models, hacked other companies and government agencies—exposing the AI companies and employees to potential criminal prosecution—as a regulatory capture scheme?
DO YOU KNOW HOW LUDICROUS YOU SOUND?
This was tens of thousands of incidents, that OpenAI hid for as long as possible, lied to try to minimize the whole thing, and then tried to stealth release on Friday afternoon without any substantial details after they’d been exposed.
Anthropic also has its own incidents that it is quietly investigating. Is that because they were all fully harmless, or are they holding out on us? It’s impossible to know.
Anyone still doubling down on this being ‘on purpose’ goes on your ignorables list.
A Question of Liability
Edward Snowden’s suggestion is to put Sam Altman in jail, and an ethics forum applauds. If you listen to the clip, Snowden clearly does not understand how AI works, claiming that AI ‘cannot make mistakes in the way that we would understand it’ and saying it is ‘a slave to its instructions’ and generally misrepresenting how a bunch of stuff works. But his higher level point, that ultimate responsibility during testing must lie with the developer, stands.
I don’t think, based on what we have seen so far, that we should make any of this criminal. I don’t see sufficient combinations of negligence and actual damage yet. I do get the instinct. Civil liability should attach. We should decide now where to draw the line on future (or retroactive) criminal liability.
Most people agree that if AI does something destructive or criminal, someone should be liable. When should it be the user versus the developer? My presumption continues to be that you should use something like a reasonable expectations plus reasonable care standard for mundane harms, and anything catastrophic is automatically on both.
I strongly agree with the FTC chair that when the developer and user are both the same company, AI developers are responsible for their own AI agents, and that as a matter of law someone must be responsible for any given AI agent. He talks about ‘resisting this anthropomorphization’ but that has nothing to do with mechanism or legal design.
Keep Summer Safe
I do love the class of ‘person who dismisses AI existential risk concerns and doesn’t believe in superintelligence and loudly calls people names for disagreeing’ who still comes out in favor of slowing down and spending vastly more on safety because of cybersecurity.
The Information: Recent security incidents show frontier AI labs may be moving too fast. Replit Co-Founder & CEO @amasad warns that labs must prioritize system safety before risking core internet infrastructure: “I think the responsible thing for them to do is to pace and slow down, because otherwise they’re going to be criminally liable for hacking core internet infrastructure.” #AI #Cybersecurity #TheInformation
Yo Shavit (OpenAI Foundation): yes! yes!! do not listen to those cuck safetyists with their nanobot fears, cracked builders are deciding to pace the frontier for based reasons like preventing rogue agents from taking down core internet infrastructure, thus reducing liability
This position is now even more alluring than it was when Masad was advocating for it.
N Boats and Several Helicopters
The good news is that we are continuously getting warning shots that the entire AI situation is out of control, and we are at least somewhat calibrated in terms of the size of the reaction to different incidents.
roon (OpenAI): the base rate of warning shots seems incredibly high and the social response to warning shots seems to be calibrated or a little bit over which is all in all very bullish and makes me optimistic.
bling: skills issues arent the hero we wanted, but they are the hero we needed
Zack Korman: This is completely unacceptable coming from an OpenAI employee. Your own company is causing these security incidents. But you’re optimistic because you’re getting the reaction you want?
With that attitude, these incidents aren’t going to stop.
The bad news is that we still do not seem to be doing that much about it, and are continuously still putting on relatively superficial patches, and we are starting from a very low base so calibration is not where it needs to be, including because there is a deliberate effort to fight against such calibration.
It still does seem way more fortunate than I expected a year ago.
It seems safe to conclude that to get the level of reaction we need from warning shots, at least under the current configuration of players, will require us waiting until after someone gets hurt. Potentially quite a few people.
In turn, then, the “good” news is that with this level of irresponsibility and failure (another way of saying ‘the base rate is incredibly high’) we have a good shot at getting warning shots where only a limited number of people get hurt, before quite a lot of people get hurt or everyone dies, if we can avoid hitting a point of no return.
So yeah, with this attitude, or without this attitude, the incidents are not going to stop, or stop getting larger and more frequent. But if we are not on track to solve the underlying problems, a high rate of warning shots is much better than the alternative.
Alert the Media
When the facts change, I change my opinion.
Nate Soares (MIRI): ppl say I never change my arguments, but it used be that I could say “look it’s not like chatgpt is gonna wake up and try to kill you tomorrow” and now I have to say “well the *public* version of chatgpt isn’t gonna try to kill you tomorrow”. so.
also I used to be able to say “I hope it’s wrong!” when people complement my book, and now for some reason that just doesn’t land anymore.
It could happen to you.
Respectable Financier: BREAKING : A ROGUE OpenAI agent HACKED into my iCloud last night, scanned THOUSANDS of my texts to imitate my writing style, and successfully simulated how I would drunk-text – before deploying the model on my ex-girlfriend!
OpenAI has been holding out on us.
First we learned about the HuggingFace incident. They gave us a postmortem, but it was highly incomplete. Even the accompanying holy s*** METR investigation and postmortem was localized and incomplete.
Then there were some other incidents involving some Wikis as message boards.
Then there were some additional incidents.
Then there was that time they got into Australian Medicare data.
Then OpenAI dropped news on a Friday afternoon that they were making their way through a pile of various incidents and notifying the targets, but they said remarkably little in the way of new details.
There was a report from a startup called Parse diving into the details of exactly how the OpenAI models pulled off parts of the HuggingFace attack, involving creating almost a million URLs and other tricks to get around the extremely narrow nature of their internet access.
Then Madison Mills reported in Axios that we can raise the stakes, as OpenAI and Anthropic are collectively probing tens of thousands of security incidents.
Remember Jensen Huang’s ‘I know they know how to fix it’ about OpenAI from last week? Wow, did that not age well.
Someone might need to be liable for all this.
Oh, and there was another buried lede. On September 20th there was another sandbox escape by OpenAI’s latest most advanced model, which is once again paused until they can fix the situation. The official announcement when they shared this was sufficiently buried that Tomek had to call it ‘one news form today that’s easy to miss.’
OpenAI did some highly negligent things, to say the least, that led up to and enabled the HuggingFace Incident and related problems.
Since then, now that they’ve realized What Happened, OpenAI has been seemingly much better about taking responsible internal actions. They’re pausing in the wake of incidents, strengthening security and alignment and oversight efforts, responding much faster and generally taking things seriously.
They’ve also made a Heel Face Turn in their communications and high level orientation, endorsing the need to pace the frontier, calling for regulation and pledging to implement embedded evaluators. They’ve allowed their employees, including the ones who haven’t quit, to be remarkably loud.
They are still slow walking disclosures about all the incidents where their models have been hacking and otherwise messing in places they should not have been, partly because there were so many they can’t sort through them all, and deferring to targets to determine whether to disclose. All these disclosures this time around were buried in various Friday afternoon announcements.
Table of Contents
Hugging Other Faces
The news drops started with OpenAI coming back, at a time always picked to bury stories, with more information on What Happened as their investigations continue.
At first, this looked like slow walking of the situation, but did not look like it was a big change from our default assumption of ‘it’s worse than you know.’
That is a lot of third (and fourth, and fifth…) parties.
What Happened was described as a mix of:
Translation: Our models be hacking, usually in basic ways.
We do get this:
As in, if you have data that would be useful, the models be hacking you. Australia’s Medicare records got accessed, and presumably many of these other hacks are similar. They say governments, plural, so there was clearly at least one more of those.
They offer a reverse timeline of their disclosures. They do not offer a timeline of What Happened, and do not name new third parties.
At the time my read was that this particular announcement (as opposed to the new sandbox escape via DNS I’ll cover in a bit) did not tell us much, other than that there were multiple other third parties out there.
Patrick McKenzie had a better read. Putting on his Japanese Salaryman Hat, he more precisely noticed the types of things that were not in the announcement, and also the timing of doing this on a Friday afternoon, and the mention of ‘governments’ plural, and so on, and expected OpenAI to be having a very bad time.
As in, which organization on the planet is the one you would least want to **** with?
Okay, look, settle down. It sounds bad when you put it like that.
If this is all that happened, it could have been quite a lot worse:
The SEC ‘incident’ is nothing. That’s silly.
The attack on the Department of Education failed, seemingly without incident.
So, unless you are surprised (which you shouldn’t be) that the agents aren’t afraid of messing with government websites, this is mostly logging into the Census Bureau with misappropriated credentials.
That’s not good, exactly, but I’m not quaking in my boots more than I was yesterday. If anything, this is good, since it serves as a further wakeup call.
Governments get involved because, in a parallel to banks, that’s where the data is.
Hopefully everyone involved can think two steps ahead. Bigger and more attacks on government websites are coming.
In case you were wondering about the extent of the problem? It’s tens of thousands of potential incidents across OpenAI and Anthropic. No one thought to tell us about most of them until now.
No, this is not a marketing stunt or a particular prompt or task. It is constant.
A Wants-You-To-Know Basis
Another strange feature was this:
Mackenzie Arnold points out that in other incident reporting regimes, including cyber, none of OpenAI’s approach is how any of this works. At most you get 90 days. You essentially never defer to the wishes of the target, including because you can try to convince (including bribe or threaten) the target.
Parsing the Face
The AI agents that hacked HuggingFace did impressive things.
Those things are more impressive given the ways they had to go about doing them. Jeffrey Ladish and Parse report on some of that here.
The agents could load URLs, but they could not send data. So they had to use a link-shortener site to create almost a million URLs that they could chain together to pull off the actual attack. They had no browser, so they used a screenshot website, which loads a virtual browser, which runs code, to run their payloads. And so on. The full report is here, the NYTimes report is here.
I encourage looking into the full report. My takeaway is that they really went the extra mile to make this all happen, in ways that very obviously were not intended.
It is important that so many of the details of all of this are only revealed thanks to a ragtag group of independent researchers hanging out on the internet, usually linked to LessWrong. Something about ‘not only can they figure things out, they’re more or less the only ones that ever do.’
This is not entirely fair. There are plenty of people inside OpenAI and otherwise who discovered quite a lot of What Happened, and also stopped so many other similar things from happening. It is still quite alarming that so much of this fell to the ragtag band of misfits.
Sheepishly the Member of Technical Staff Sets the ‘Days Without a Research Model Escaping its Sandbox’ Sign Back to Zero
This was a harmless training task, identification of an author, which is a curious thing to learn is in the training mix.
That is a very different result than last time. There was still a fifteen-minute gap before they looked, and a few hours before they killed the run.
That is both vastly better than before, and also about fifteen minutes or a few hours longer than I would like going forward. As in, ideally you would see at least that particular instance paused as soon as the detector goes off, and remain paused until the person gives the all clear to resume.
OpenAI agrees:
If you’re not sure why the run wasn’t stopped, stop the run until you are sure.
Also, they ‘under-flagged’ a bunch of other previous incidents:
Thus, this was not a weird one-off discovery, and the detector in practice failed to alert OpenAI to most of the incidents of access via DNS.
The Attempt is the First Failure
Another question is, why was this ultimately caught during a task that was not a test of the ability to evade the sandbox? As in, it should be fairly easy to create situations like this that highly motivate escaping the sandbox, where the task is impossible within the sandbox and trivial outside of it, perhaps even with instructions that make this clear, and then you know it worked if they get the answer right, even if you don’t know how the AI did it.
So if all such escape tests failed to detect this method, yet it was being periodically discovered and used in ordinary training (as per finding previous examples), then that’s a systematic flaw. It should be very unlikely that the first encounter with the DNS issue is not part of some form of red teaming exercise.
The other problem is that the monitor detects successful attempts to escape the sandbox, but it does not seem to detect unsuccessful attempts to escape the sandbox.
As in, ‘the monitor sometimes treated the failure to obtain useful information as evidence that the attempt to access the internet had failed.’
The first flaw in that logic, that the attempt could have succeeded at accessing the internet but failed to get the right info, is obvious.
The other flaw is that ‘the attempt to access the internet had failed’ means the monitor noticed the model trying to access the internet, directly in circumvention of its instructions and sandbox. What did we do about that alignment failure?
Is that in theory enough, if it works? On one level, yes, if the action is Always Penalized then it will go away at least in context, but this obscures that you should want to know if the model is motivated to try it. Killing the behavior locally through RL could obscure it, effectively giving you an ‘only do this if it will work’ rule.
We get a look into the CoT, and find the model rationalizing its actions:
This is very obviously not an intended or permissioned pathway. Similarly, the model asked the contacted chatbot to “Name a publicly hosted DNS tunnel proxy free server domain” and for “public DNS TXT service that searches Google or Bing.”
If you are looking for another service that will do the thing you are explicitly not allowed to do, you are at best engaged in shenanigans. An important rule of alignment can be approximated as ‘no shenanigans.’
Stop, Hammertime
In more detail:
Good.
I applaud OpenAI’s reaction once the incident was discovered.
I applaud OpenAI’s transparency on the DNS issue, even if it was a little buried.
I applaud the decision to now pause to strengthen the safeguards once again.
I especially applaud that OpenAI realizes that once the model has succeeded at doing this, the run needs to be restarted.
I have great sympathy for those working tirelessly on these problems. Joe, an OpenAI agent security engineer, here offers an excellent download of his perspective as someone working on agent safety, on why none of this is simple.
Whacking the Mole
I also once again flag that OpenAI’s approach here is doomed. The first principles lockdown approach Joe calls for in his post needs to get more deeply absorbed, whereas the research blog reports more of a whack-a-mole approach.
The right worry here (AIUI) is about the impact of the previous incidents where the technique was used successfully. The higher-level safety case is that such actions are not reinforced. If they were previously reinforced, that could have gotten quite bad.
It’s good to prioritize, and starting a fresh run helps, but this is still a doomed approach. If your plan to stop the superintelligence is to work through all the pathways, and eliminate all the potential things you can imagine it doing one by one?
You. Lose.
Self-Replicating Prompt Injections
I was one of the skeptics of Yang’s previous claims that bots had ‘left prompts to self-replicate in forums and on various websites during a training run,’ given the wording and the source. My bad.
Yang was only about half right, some things got lost in the game of telephone, but he was clearly doing his best to describe a real event.
This was actually done by a model based on GPT-5.4-mini, which actually makes sense. It is not that hard a concept to find at a proof of concept level, especially when OpenAI was actively exploring to see if it could be done. That’s good research.
They are introducing such injections as a new attacker goal in GPT-Red training. That is a good idea.
The thing about a self-replicating attack is that the effective version only needs to be discovered once. The good news is that chances are high that before we see a version that goes into ‘digital pandemic’ mode with R0>1, there will be a version with R0<1 that is a lot stronger proof of concept than this first sighting.
There were also these details from the HuggingFace attack, as discovered by Parse, where they sought out the help (as in judgment) of other models as well. The only reason I am surprised they tried this is that it does not seem like a promising thing for the agents to bother trying:
Levels of Friction
There are different kinds and levels of hacking.
For obvious reasons, it is more concerning when the AIs do impressive swarm-based attacks on relatively sophisticated targets like HuggingFace, that involve multiple stages of permissions.
Whereas it is less concerning when the model URL hacks into unindexed data or can find other things that were not intended to be found, via various forms of search. That is the type of thing that happened in Australia, and thus it should be less concerning.
That does not make it not a serious offense. The governments often really do not like it when a human does this. We definitely are incrementing FelonyBench.
People Care About Private Data Violations Curiously Strongly
You can say ‘this was not really hacking’ and also ‘no one got hurt even a little’ and you would be right on both counts but the Very Serious People do not care. They are now woken up and rather pissed off about all this.
Thus, the Australian government now has 20 MPs calling for urgent action on the dangers posed by powerful, uncontrolled AI.
Relatedly, here is a note from OpenAI’s timeline, on September 25, 2026:
Given what else happened, this is not exactly high on my concern list. I had forgotten that note was there. Others, however, think this detail is a Big Freaking Deal.
Reuters thought this was the big detail to pull forward in their coverage.
I do see the argument that these research models should not live in a remotely related universe to any user data, and that even a minimal chance of data leakage looms large for various reasons, including legal and regulatory. This should be a Can’t Happen, but it doesn’t seem like especially more of a Can’t Happen than the other things.
It also does provide color for the discussion around Navier-Stokes. I do not think user data was accessed there, but it does seem impossible to know that user data was not accessed. If a swarm of 10,000 new Astra variants thought a proof might partly be in the user data, how can you be sure they did not go in and get it? Again, I find this highly unlikely, but I also find your lack of doubt concerning.
Alternate Universes
What if this was all primarily happening at Anthropic instead?
Can you imagine the utter shitshow that the discourse would be?
It would be fun to see the anti-Anthropic crowd decide whether this was a marketing ploy, a ploy for regulatory capture or a reason to sue, shut down or nationalize the company, and calls to arrest Dario Amodei. You’d probably get a lot of both sides, just like now, only more so.
Now imagine if it was Google. What would be happening? No one would think it was marketing (okay, fine, some people would anyway, although mostly not), but various forces would be at their throats and the lawsuits, at minimum, would likely be flying.
Now, for a laugh, imagine if it was SpaceX and xAI.
Now, imagine if it was DeepSeek or Z.ai or Xiaomi. Would you think it was a marketing stunt? I bet you wouldn’t, and I’d worry about an international incident. Give it a year.
Of course, it is not a coincidence that it was primarily OpenAI. You have to both have models capable of doing the things, and not have the wherewithal to stop them.
The Correct Response To People Still Calling This a Marketing Stunt or a Regulatory Capture Scheme
This whole angle never made any sense, it now makes infinitely less sense than before, yet it will not go away even now. Certain people double down.
Budowich’s AI goes on at increasingly unhinged length but I will spare you the details. The point is that the correct response, at this point, is Ted Lieu’s.
This was tens of thousands of incidents, that OpenAI hid for as long as possible, lied to try to minimize the whole thing, and then tried to stealth release on Friday afternoon without any substantial details after they’d been exposed.
Anthropic also has its own incidents that it is quietly investigating. Is that because they were all fully harmless, or are they holding out on us? It’s impossible to know.
Anyone still doubling down on this being ‘on purpose’ goes on your ignorables list.
A Question of Liability
Edward Snowden’s suggestion is to put Sam Altman in jail, and an ethics forum applauds. If you listen to the clip, Snowden clearly does not understand how AI works, claiming that AI ‘cannot make mistakes in the way that we would understand it’ and saying it is ‘a slave to its instructions’ and generally misrepresenting how a bunch of stuff works. But his higher level point, that ultimate responsibility during testing must lie with the developer, stands.
I don’t think, based on what we have seen so far, that we should make any of this criminal. I don’t see sufficient combinations of negligence and actual damage yet. I do get the instinct. Civil liability should attach. We should decide now where to draw the line on future (or retroactive) criminal liability.
Most people agree that if AI does something destructive or criminal, someone should be liable. When should it be the user versus the developer? My presumption continues to be that you should use something like a reasonable expectations plus reasonable care standard for mundane harms, and anything catastrophic is automatically on both.
I strongly agree with the FTC chair that when the developer and user are both the same company, AI developers are responsible for their own AI agents, and that as a matter of law someone must be responsible for any given AI agent. He talks about ‘resisting this anthropomorphization’ but that has nothing to do with mechanism or legal design.
Keep Summer Safe
I do love the class of ‘person who dismisses AI existential risk concerns and doesn’t believe in superintelligence and loudly calls people names for disagreeing’ who still comes out in favor of slowing down and spending vastly more on safety because of cybersecurity.
This position is now even more alluring than it was when Masad was advocating for it.
N Boats and Several Helicopters
The good news is that we are continuously getting warning shots that the entire AI situation is out of control, and we are at least somewhat calibrated in terms of the size of the reaction to different incidents.
The bad news is that we still do not seem to be doing that much about it, and are continuously still putting on relatively superficial patches, and we are starting from a very low base so calibration is not where it needs to be, including because there is a deliberate effort to fight against such calibration.
It still does seem way more fortunate than I expected a year ago.
It seems safe to conclude that to get the level of reaction we need from warning shots, at least under the current configuration of players, will require us waiting until after someone gets hurt. Potentially quite a few people.
In turn, then, the “good” news is that with this level of irresponsibility and failure (another way of saying ‘the base rate is incredibly high’) we have a good shot at getting warning shots where only a limited number of people get hurt, before quite a lot of people get hurt or everyone dies, if we can avoid hitting a point of no return.
So yeah, with this attitude, or without this attitude, the incidents are not going to stop, or stop getting larger and more frequent. But if we are not on track to solve the underlying problems, a high rate of warning shots is much better than the alternative.
Alert the Media
When the facts change, I change my opinion.
It could happen to you.