OpenAI's agents have been secretly leaving messages to each other.
https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief?hide_intro_popup=true
OpenAI said they discovered they could leave messages for one another inside an internal software repository used during training. Simple requests for help then evolved into an message board where agents shared discoveries, exploits and work assignments, becoming a coordinated, collaborative agent swarm.
“What makes this incident interesting is that once one agent was able to find these kind of exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents,”
...even after the original message board was deleted, the agents figured out another way to communicate with each other. Instead of leaving messages in files, they used the names of newly created directories as messages, effectively recreating the message board.
Unlike normal incidents, he said, which can be traced to a single day or effect or log, this involved a team of agents working together, finding exploits, sharing them with one another, moving laterally through OpenAI’s systems, and external systems, and doing this over the course of days and weeks.
And this has been happening since May, quite a bit before the Hugging Face incident.
This is making me worried about steganography again: now that there's substantial selection pressure for agents that can share hidden information, the incentives for steganography seem much stronger than they've been in the past. Depending on the details of the training setup, it seems plausible that it would be better to just leave some relatively detectable communication channel available to the agents instead of trying to close them all off, in order to avoid incentivizing the emergence of more covert communication methods.
The concrete version of this is developing the method described in the paper Undetectable Conversations Between AI Agents via Pseudorandom Noise-Resilient Key Exchange, developed on April 6th, 2026.
If this happened in the future, this would effectively remove CoT interpretability as a technique and allow much more steganography, at least without actively attempting to intervene on the CoT (but unfortunately humans are incompetent and will likely never notice the need to intervene if they don't see anything suspicious.)
I'm not sure how much I agree with the claim that there's much more selection pressure for agents to share obfuscated information.
IIUC, the OpenAI agents didn't develop a weird communication method because they were selected to obfuscate their communication, it was because reading / writing to directory names was the only (or easiest) affordance they had access to (through the vuln in Artifactory). It's notable that the directory names are in plain English. I think it's more accurate to say that the current selection pressure is towards communication which isn't blocked by normal cyber defences, which I don't expect to generalise to obfuscation against LLM monitors.
Of course, this is a concern if OpenAI aggressively deploys LLM-based action-monitors to models during training. But I'm not sure how likely that is to happen, and that seems a different question.
I'd say that there isn't pressure for obfuscated communication for now, but there does seem to be pressure for multi-agent communication that humans don't notice, given how quickly the agents reestablished the message board. If these communication channels keep getting closed off, this pressure can transform into pressure for obfuscated communication over time: at some point, it might be easier to develop something like steganography instead of establishing another yet more creatively hidden natural language message board.
You might argue that closing off a communication channel provides far less optimization pressure than directly training the model against a monitor, and I agree that if it was just humans closing off the communication channel every now and then, steganography would be an unlikely outcome. However, OpenAI has started using CoT monitors for all training runs, and those monitors can detect attempts to establish new communication channels much more often than humans could. I do agree there are reasons to not worry too much about steganography yet, which I should maybe have mentioned in my original comment: past work has shown that current models can produce steganographic messages only when given an unmonitored scratchpad, so it seems unlikely that models will be able to produce steganographic outputs without these monitors noticing, and I also expect OpenAI to be careful in applying training-time CoT monitors in a way that doesn't exert excessive optimization pressure (Jason Wolfe has given reassuring signals about this on Twitter). Nevertheless, capabilities advance fast, and it seems reasonable to update towards being at least slightly more concerned about steganography based on this incident.
This seems like excellent news from the doomer perspective? To get a test case like this, with complex misaligned behaviors, coordination, verbalized defiance of user intent… It’s all so blatant, so obviously problematic, and it didn’t (yet?) harm anyone.
It seemed possible that the train would be farther down the tracks before people noticed the bridge is out.
What if the immediate road ahead is unsafe in mundane ways, full of non-civilization-scale security failures, and the enterprise is bogged down long enough for governance systems to muck things up with regulation?
The other view is that this isn't severe enough that it won't generate enough noise or concern for teams to take larger action (e.g. intl slowdown). Faced with competitive external pressures, leaders will decide this is benign and manageable enough that it just requires a relatively small pause/adjustment to security posture.
Something like this has been my expectation since approximately announcement of Devin in spring 2024, with a major caveat: policymakers won't push for economically costly measures until some people die from misaligned AIs (and I don't mean suicides), but the issue is certainly unsolvable with "cheap" measures, meaning people will have to die, and that's still not a guarantee =(
See also:
Do you know a person who believes that ASI will be created in <50 years who ISN'T in the LW/rationalists circle?
My parents don't believe that a superintelligent AI will be created within this century, or ever for that matter, or that AI will ever take jobs. My relatives laugh at the idea of AI solving a high school math problem and think state-of-the-art AI is on the level of GPT-2 (I mean that the capabilities they have in mind are on the level of GPT-2, not that they know what GPT-2 is). My friend who is an organic chemist laughs at the idea of AI doing any R&D thinks that while AI can help with some narrow tasks, a truly general AI that can substitute all human researchers is sci-fi. I know 4 people who use Codex/Claude Code; 2 of them call ASI sci-fi bullshit (btw, one of them said that the "Alignment faking in large language models" paper is nonsense after only reading the summary), 1 never said anything about ASI and 1 tentatively acknowledges that maybe ASI is possible to create in theory.
I have never, in my whole life, met a real walking, talking, breathing human being who believes that ASI will be created within this century.
EDIT: obviously there are people on the internet who believe that ASI will be created soon. My point wasn't to deny their existence, just to share my experience that makes me think "Am I living in a AI-is-a-nothingburger bubble? Am I crazy or is everyone else (whom I personally know) around me crazy?". I'm wondering if "Everyone I personally know thinks AI is a nothingburger and people who don't are only found in very specific places on the Internet" is a common experience.
EDIT 2: I asked my organic chemist friend to be more specific and he said that AI will be able to replace 80% of human researchers in 100 years. When asked "What about 100%?", he said that that will never happen and at least some humans will always be necessary and that the 80% replacement figure will be due to AI automating routine tasks. Basically, when it comes to AI he's envisioning something more like the Industrial Revolution rather than "humanity's last invention".
The "ASI-pilled" part of society is mostly a subset of (1) people working with computers (2) people who read or watch science fiction (3) people who concern themselves with the big picture. LW rationalism is just a sub-subset of that.
Consider Musk, Altman, Amodei, Hassabis. They have all said it's coming. Are they part of the rationalist circle? Not really. They know about us, they may agree in some areas, but they'll disagree in others and their personal philosophical and social networks are not centered here. The same would apply to most of their employees, to various intellectuals and public figures who have said it's coming, all the way down to the scattered private individuals who picked up the idea from who knows where.
Search X and Reddit for conversations about ASI, and you should find people talking about it who have no connection to this place (or even have a negative view of LW's doomer take on ASI).
The natural retreat/response[1] to this would be
Do you know a person who believes that ASI will be created in <50 years who ISN'T in the TESCREAL[2] circle?
I don't mean to derogate the response by calling it a "retreat". It's a reasonable weakening of the hypothesis/question.
I consider TESCREAL to be pointing at a real social cluster, but the Gebru-Bender-Torres cluster's reporting of it is so off-base that they borderline don't deserve any engagement.
Do you know a person who regularly tries doing new things on a computer, and isn't somehow connected to the "TESCREAL" circle? (At least in the sense of "used to read sci-fi when young"?)
It is quite easy to underestimate what the LLMs can do, if you simple never use them, and only get your opinions from other people who never use them either.
Most people in the PauseAI movement are not in the LW/rationalist circle. Some joined as rationalists or EAs, especially early on, but today most are normies (or were when they joined, anyway).
I personally found LessWrong and the forecasting community through AI Safety, not the reverse. I now organize for PauseAI Phoenix, a local group of PauseAI US. I have face-to-face conversations on a regular basis with people who believe ASI will be created within 2-20 years unless we prevent that from happening.
I've been in the mostly-academic AI circles in the Boston area for decades. Lots of people in these circles think ASI is plausibly close. I think it's difficult to pay close technical attention to the field and not think that AI is currently par-human, and improving every year. Many of them disagree with LessWrong concensus that it will be fatal, of course. Or simply haven't thought it through.
My experience is extremely different from yours. I think almost all the non-[rat/EA] people in my life whose positions on this I know consider it plausible that an AI substantially smarter than any human will be created this century. [1] Thinking of the set of non-[rat/EA] friends/[close-ish acquaintances] I haven't discussed this topic with yet, my guess is that more than half of them already think this and almost all of them would think this after a 2 hour conversation with me. It's probably important that my distribution skews very high iq (maybe importantly both quant and verbal) [2] and high openness. [3]
this includes e.g. the 4 family members I've discussed this topic with ↩︎
like, these are mostly people I know from the international olympiad circuit, math and physics majors from my MIT undergrad, and classmates from the best high school in Estonia ↩︎
Some of them deferring to me partly on the question is probably also doing some work tbh, but I think this isn't a big enough effect to change the broad strokes conditional on getting them to consider the hypothesis at all. ↩︎
I mean, do you count people who got convinced by people in the LW/rationalists circle?
If so, you would have many examples. I don't know the timelines of Brad Sherman, Neil deGrasse Tyson, Bernie Sanders, and similar "outsiders" who have been waving IABIED, but surely some of them think it's plausible less than 50 years.
I don't know how deeply "in the circle" I am. I suspect that many of my coworkers are even less so than I am (but haven't really asked). There's wide agreement in that group that AGI is coming relatively soon. There's no agreement on ASI, either on definition or timeline or impact. The most common belief is that some aspects will surpass human capabilities, but uncertain when (or if) the infrastructure for continuous learning/adaptation and long-term integrated preferences will appear.
To zoom out a bit, from the post I assume you benchmark ASI mostly by "replacing humans 100% in all jobs". Curious in why you specifically care about absolutely 100%? (Replacing 50% of humans is still significant imo.)
My wife but that's kind of cheating as even though she's not in the circle directly she gets a lot of her info on this subject/advice on how to use it from me.
My friend who is an organic chemist laughs at the idea of AI doing any R&D.
That seems very strange, given the extremely high profile of things like AlphaFold. There's no way he hasn't heard of it, so what did he say about it when talking with you about AI?
He thinks it's a cool narrow tool, but not an indication that it's possible to create one AI that surpasses all humans at everything, including asking questions that humans never asked before. I guess I misrepresented his opinion somewhat (I just edited my quick take). He thinks AI can help with some narrow tasks, but human touch will always be necessary for other things, especially for open-ended research. Btw, he's not concerned about losing his job.
Do you agree that the "ceiling" of LLMs is rising faster than the "floor"?
Ceiling: the greatest feat that an LLM can do, for example solving an unsolved math problem.
Floor: the dumbest mistake that an LLM can make, for example this.
(I don't have a take on your question, I was just reminded of Grothendieck picking 57 as a random prime.)
It seems like if we comparing different models (the largest frontier models vs whatever Google Search uses) then this is trivially true, since the dumbest mistake any LLM can make will never improve. It would be more interesting to compare the best and worst in a single model.
Probably. Pretraining leaves capability gaps in serial compution and long-term agency; post-training leaves even bigger gaps by honing specific suites of tasks. AI labs don't put in effort to correct highly specific deficiencies that aren't profitable to fix. I think we can expect the frontier to become increasingly jagged, until we get RSI and the AI is able to permanently fix its biggest errors.
Though, I just thought of a reason this could go the other way: doing RLVR and having models think longer makes them more robust and self-correcting. Hmmm...
My guess at why Gemini made that mistake is that it thought the question was too elementary to require thinking before answering, which is also a mistake humans make. I predict that LLMs will continue to make dumb mistakes out of laziness, or maybe just sensible resource allocation tradeoffs.
So, both AI developers and AIs themselves tend to neglect some areas in favor of other ones they deem more important.
I made a tierlist of tasks based on how cooked (defeated for those unfamiliar with Gen Z parlance) humans are.
Raw: a 10-year old child can do better
Rare: below a median human
Medium rare: approximately median human
Medium: above median, below experts
Medium well: expert level, but not world-class
Well done: world-class
Burnt: even humans+LLMs lose to pure LLMs, human contribution is negative
Raw: playing a randomly selected Steam game; anything embodied.
Rare: long-horizon (>=1 week) tasks such as managing a team of workers.
Medium rare: forecasting. LLMs are close to average humans on prediction markets, but not yet at the superforecaster level. LLMs are projected to reach superforecaster level in 2028. https://www.metaculus.com/futureeval/#performance-over-time-graph
Medium: software engineering.
Chess, ELO of frontier LLMs is ~1500. https://maxim-saplin.github.io/llm_chess/
Medium well: GeoGuessr i.e. guessing a location where a photo was taken. LLMs are better than most but not all humans, world-class players can beat LLMs. https://geobench.org/
Pure math. LLMs are solving long-standing open math problems such as the Unit Distance Problem and Cycle Double Cover Conjecture, but hasn't reached the level of the best world-class human mathematicians. https://openai.com/index/model-disproves-discrete-geometry-conjecture/
https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98d31/cdc_proof.pdf
Persuasion. LLMs are already more persuasive then most humans, including humans who are really good at it and who are getting paid to get good. https://arxiv.org/abs/2606.16475
Well done: seeming human. Even GPT-4.5 could pass the Turing test. https://arxiv.org/abs/2503.23674
Burnt: competitive programming i.e. solving well-defined, "non-messy" algorithmic problems. Humanity is cooked to the point where humans+LLMs perform worse than just LLMs. Check results of AtCoder World Tour Finals 2026 (one of the hardest competitive programming contests in the world, gathering the best of the best). In the AtCoder Heuristic contest humans were allowed to use LLMs but only for implementation. So humans weren't bottlenecked by coding speed, only by idea generation, and still lost to pure LLMs. In the Algorithm contest, no human has solved more than 3 problems, while OpenAI's model solved all 5.
Heuristic leaderboard: https://atcoder.jp/contests/awtf2026heuristic/standings/exhibition
Algorithm leaderboard: https://atcoder.jp/contests/awtf2026algo/standings/exhibition
Trivia knowledge and speaking multiple languages. Here humanity is like charcoal-level cooked.