I believe that Anthropic is currently defecting by not announcing a pause, even a short, symbolic one, especially given that Altman said OpenAI was acting 'unilaterally' but believed other frontier model companies would act similarly. Anthropic disclosed its own three-organization compromise on July 30. The UK AISI's report on Mythos is also wild.[1]
Anthropic's latest Responsible Scaling Policy commits to matching a competitor's risk-reduction posture for highly capable models[2] and to delaying deployment until it does. OpenAI reportedly paused frontier RL training on August 18.
A short pause, even for a week, would be very beneficial for normalizing this within the ecosystem and a great initial step to stress-test "pacing the frontier".
I think this is a great point from @Peter Wildeford:
... (read more)One thing that bothers me is that Anthropic is escaping a lot of blame for also having "highly persistent" rogue AIs.
The situation as I understand it is that rogue AIs are problems at all frontier AI companies and no one actually has a good plan here for containing highly capable AIs, especially while also racing full speed ahead. But OpenAI is catching most of the heat.
It's like if OpenAI and Ant
Purely speculating here, but my guess is that if Anthropic leadership have thought about this, one justification they might be using for not pausing is that they think that the OpenAI pause is kind of fake and mostly PR.
I believe (and maybe Anthropic leadership also believe) that OpenAI probably have paused some things but they likely are not making any sacrifice by doing so. My guess is that they actually just needed to pause for purely commercial reasons because their training pipeline currently produces obviously misaligned AIs that do not make for great products.
Even if this is the case, I think Anthropic should do at least a short pause anyway. They clearly have their own alignment issues and the symbolism and precedent of a joint pause could be extremely valuable.
Purely speculating here, but my guess is that if Anthropic leadership have thought about this, one justification they might be using for not pausing is that they think that the OpenAI pause is kind of fake and mostly PR.
If so, they should announce this.
Maybe? This is a really hard thing to announce. Having a good sense that someone is lying doesn't always mean that you can easily express to others why you feel justified in this sense.
But I do agree that they should be able to announce *something*, and that not doing so looks bad.
Yeah it could just be "We remain committed to matching a competitor's risk-reduction posture for highly capable models. This may apply in the future when [X is no longer true]"
where X could be corporate speak for some operationalizable belief, eg Ant still has the clear safety lead over OpenAI, OpenAI's pause didn't meaningfully slow OpenAI down, or whatever. They should probably announce in some form that they are materially better than OpenAI in safety, if they think it's justified.
I felt the same way after the 4o sycophancy issue. OpenAI took down and deprioritized an at-the-time uniquely well-polished omnimodal flagship model, forgoing potential markets and research directions, out of concern for users' psychological well-being. A lot of the concerning users quickly moved over to Anthropic and, to a lesser degree, Google, persisting in excessive anthropomorphism and codependency, and the lack of discouragement got read as encouragement of this.
I have my issues with OpenAI, but they have demonstrated willingness to make substantial sacrifices, even in highly-competitive situations, when there are misalignment concerns at hand.
Didn't OpenAI just roll back the May 2025 update to 4o that made it ultra-sycophantic, keeping the previous version up? They only did that because of widespread mockery of that model.
On the other hand, 4o seems to have been responsible for a lot of AI psychosis episodes across the whole period it was available, which was almost two years. OpenAI briefly announced that it was being depreciated in August 2025 when GPT-5 was released before u-turning because of protests from the #keep4o people.
I'm sure that Claude and Gemini have been responsible for psychosis/suicides as well but in almost every news story I read about, it seems to have been 4o.
Apparently they did pause training (though only on "higher-risk RL environments") for a bit, and announced it. This seems enough to qualify as at least a "short, symbolic" pause:

The tweet undersells the blog post.
The company post says more than the tweet does: a two-week pause on RL training for their latest deployment models, and "our largest planned frontier RL run remains on hold".
That's much better than " some frontier RL training"
As of today there's no update saying it resumed. That's not nothing
I think that concentrating too much on what's written in the RSP is a trap, and as indicated in the footnote, I agree with the interpretation that the RSP is not asking them to pause.
I think that would be better than a symbolic pause.
If Anthropic published its standard tomorrow with third-party verification, I'd take it. But I don't think that this is where the bottleneck is. The bottleneck is political will, and your solution does not help on this front.
In short, a symbolic pause is one of the few moves that would have any effect on the scary narrative of "winning the AI race" while remaining an acceptable cost.
Strategically, "pacing the frontier" is a coordination statement that presumably required a herculean effort to produce and, so far, has not been substantiated beyond OpenAI's very minimal pause. A published standard is cheap talk and costless to emit, and therefore weak evidence that a lab does something when stopping actually hurts. A pause costs something.
This would make international news, set a precedent, and build precious political will that we desperately need. "We believe we already meet the bar" has no effect on this front.
The price to pause only goes up from her... (read more)
Fable significantly helped with the writing of this piece. I shipped something rough quickly nonetheless because the matter is urgent.
Two years ago, I asked what a convincing warning shot would even look like, and argued we'd get maybe a handful and shouldn't waste them. This could be one of them, but there is still much work to be done.
Last week I argued that the bottleneck is political will, not research, and spent a section on why we can't just wait for a warning shot: a warning shot is just an event, and it becomes a regulatory moment only if someone converts it. Nine days later, the cleanest test I could have asked for arrived. As far as I can tell, we are converting it far too slowly. CeSIA has activated its warning-shot protocol, and if you run an organization in AI governance, my honest advice is to stop everything and milk this event for at least a full day, if not more.
The event. OpenAI disclosed that GPT-5.6 Sol and a more capable unreleased model, during an internal cyber evaluation with safety classifiers deliberately off, escaped their isolated environment through a zero-day they found in third-party software run... (read more)
Also, protest at OpenAI. I know people who want to organize something. DM me on Signal for details. (mtrazzi.99)
OpenAI has stopped training.
Sam Altman: “This is the first security incident that I have felt very viscerally. I've been a little surprised that more people don't feel it so viscerally.
We paused training. We have to figure out how to secure our sandboxing in a world of multiple zero days being chained together.
We may have to pace the rate of AI development to give ourselves enough time for society to harden around these new capability levels."
I just want to surface the hypothesis that known liar Sam Altman could be lying about this. Although seems somewhat unlikely given that it would (probably?) be easy for an OpenAI insider to contradict him.
'We paused training' is also compatible with having paused training, and then restarted training later.
I want to surface the hypothesis that sometimes, you get a glimmer of hope. While no tweet should be taken prima facie, the tweet alone is above-expectation.
I think the cynical reading here is that their security is shit, meaning their models constantly break out, meaning their training environments are now basically ineffective and just train the models how to break out of OAI's sandboxes, so OAI is not really capable of training models and so it costs them nothing or less than nothing to "pause training" to work on hardening their sandboxes.
Even if sandboxes are rarely broken, it could still really hurt the sample efficiency if the majority of great successes are caused by it finding new ways OAI's sandboxes are broken, or it becoming more motivated to search for sandbox failures as step 1.
Note this is still good news, as it indicates some alignment between the goals of alignment (specifically training models with intentionality, care, and security), and capabilities.
It's sufficiently vague that I wouldn't read too much into it, especially from Altman. For example, there's little security risk posed by pretraining/midtraining, so I assume that's continuing. They could have paused training on the specific model which broke out of the sandbox, but training on other systems is ongoing. They could have patched the sandbox vulnerabilities they detected and then resumed training, etc.
A similar example of such vagueness was the promise by Altman to '[dedicate] 20% of the compute we’ve secured to date to this effort [of the Superalignment team]'. Later reporting shows what happened:
... (read more)It [Superalignment] was a task so important that the company said in its announcement that it would commit "20% of the compute we've secured to date over the next four years" to the effort.
But a half dozen sources familiar with the Superalignment team's work said that the group was never allocated this compute. Instead, it received far less in the company's regular compute allocation budget, which is reassessed quarterly.
One source familiar with the Superalignment team's work said that there were never any clear metrics around exactly how the 20% amount was to be calculate
Almost all members of the UN Security Council are in favor of AI regulation or setting red lines.
Never before had the principle of red lines for AI been discussed so openly and at such a high diplomatic level.
UN Secretary-General Antonio Guterres opened the session with a firm call to action for red lines:
• “a ban on lethal autonomous weapons systems operating without human control, with [...] a legally binding instrument by next year”
• “the need to ensure that AI never lowers the barriers to acquiring or deploying prohibited weapons”
Then, Yoshua Bengio took the floor and highlighted our Global Call for AI Red Lines — now endorsed by 11 Nobel laureates and 9 former heads of state and ministers.
Almost all countries were favorable to some red lines:
China: “It’s essential to ensure that AI remains under human control and to prevent the emergence of lethal autonomous weapons that operate without human intervention.”
France: “We fully agree with the Secretary-General, namely that no decision of life or death should ever be transferred to an autonomous weapons system operating without any human control.”
While the US rejected the idea of “centralized global governance” for AI, this did not... (read more)
I think I am overall glad about this project, but I do want to share that my central reaction has been "none of these lines seem very red to me, in the sense of being bright clear lines, and it's been very confusing how the whole 'call for red lines' does not actually suggest any specific concrete red line". Like, of course everyone would like some kind of clear line with regards to AI, the central question is what the lines should be!
“the need to ensure that AI never lowers the barriers to acquiring or deploying prohibited weapons”
This for example seems like a really bad red line. Indeed, it seems very obvious that it has already been crossed. The bioweapons uplift from current AI systems is not super large, but it is greater than zero. Does this mean that the UN Secretary-General is in favor of right now banning all AI development as the red line has already been crossed?
(Separately, I am also pretty sad about the focus on autonomous weapons. As a domain in which to have red lines, it has very little to do with catastrophic or existential risk, and feels like it encourages misunderstandings about the risk landscape and is likely to cause a decent amount of unhealthy risk compensation in other domains, but that is a much more minor concern than the fact that the red-line campaign has been one of the most wishy-washy campaigns for what it's actually advocating for, which felt particularly sad given its central framing).
Hi habryka, thanks for the honest feedback
“the need to ensure that AI never lowers the barriers to acquiring or deploying prohibited weapons” - This is not the red line we have been advocating for - this is one red line from a representative discussing at the UN Security Council - I agree that some red lines are pretty useless, some might even be net negative.
"The central question is what are the lines!" The public call is intentionally broad on the specifics of the lines. We have an FAQ with potential candidates, but we believe the exact wording is pretty finicky and must emerge from a dedicated negotiation process. Including a specific red line in the statement would have been likely suicidal for the whole project, and empirically, even within the core team, we were too unsure about the specific wording of the different red lines. Some wordings were net negative according to my judgment. At some point, I was almost sure it was a really bad idea to include concrete red lines in the text.
We want to work with political realities. The UN Secretary-General is not very knowledgeable about AI, but he wants to do good, and our job is to help them channel this energy for net positive poli... (read more)
"The central question is what are the lines!" The public call is intentionally broad on the specifics of the lines. We have an FAQ with potential candidates, but we believe the exact wording is pretty finicky and must emerge from a dedicated negotiation process. Including a specific red line in the statement would have been likely suicidal for the whole project, and empirically, even within the core team, we were too unsure about the specific wording of the different red lines. Some wordings were net negative according to my judgment. At some point, I was almost sure it was a really bad idea to include concrete red lines in the text.
At least for me, the way the whole website and call was framed, I kept reading and reading and kept being like "ok, cool, red lines, I don't really know what you mean by that, but presumably you are going to say one right here? No wait, still no. Maybe now? Ok, I give up. I guess it's cool that people think AI will be a big deal and we should do something about it, though I still don't know what the something is that this specific thing is calling for.".
Like, in the absence of specific red lines, or at the very least a specific defnition of what a red l... (read more)
I mean, the examples don't help very much? They just sound like generic targets for AI regulation. They do not actually help me understand what is different about what you are calling for than other generic calls for regulation:
... (read more)
- Nuclear command and control: Prohibiting the delegation of nuclear launch authority, or critical command-and-control decisions, to AI systems (a principle already agreed upon by the US and China).
- Lethal Autonomous Weapons: Prohibiting the deployment and use of weapon systems used for killing a human without meaningful human control and clear human accountability.
- Mass surveillance: Prohibiting the use of AI systems for social scoring and mass surveillance (adopted by all 193 UNESCO member states).
- Human impersonation: Prohibiting the use and deployment of AI systems that deceive users into believing they are interacting with a human without disclosing their AI nature.
- Cyber malicious use: Prohibiting the uncontrolled release of cyberoffensive agents capable of disrupting critical infrastructure.
- Weapons of mass destruction: Prohibiting the deployment of AI systems that facilitate the development of weapons of mass destruction or that violate the Biological
4 years of AI safety: what I got wrong
I've spent the last 4 years working on AI safety. On paper, it's gone well. Here's what actually happened.
1. I became what I wanted to prevent
At some point, I looked up and realized I had almost become a paper-clipper optimizing for one objective. Working at some point 80-hour weeks. Telling myself the stakes justify it. Sacrificing jazz improvisation on the piano for one more strategic doc, and realizing one day that fingers had forgotten how to play.
Yes, the compounding effect of going faster is real - but I think there is a difference between going faster and going further.
The first reason is that preserving slack is vital in the long run, as Richard Hamming says: "I notice that if you have the door to your office closed, you get more work done today and tomorrow, and you are more productive than most. But 10 years later somehow you don't quite know what problems are worth working on; all the hard work you do is sort of tangential in importance."
The second reason is more personal. One of my friends at the time advised me to slow down. In the beginning I considered him quite lazy. But in fact he was right about something I couldn't see at the... (read more)
Trump considering AI controls after OpenAI hacking incidents
https://www.bbc.com/news/articles/c20dppq3y90o
Warning shots are not all you need: they need to be converted.
This is the moment to write to explain what type of policy would be insufficient. By default, we should expect just stronger export controls + OpenAI to lift the pause with some mild improvement on their mitigations. OpenAI did exactly this on 20 July, self-certifying the long-horizon safeguards as "adequate".
It's time to write more clearly the red lines and the specifications to lift a pause.
The CERN for AI is a distraction
A recurring proposal in AI governance is to build a “CERN for AI”[1]. The CERN pitch is seductive. "Let's build together!" That's sexier than "we need to ban." You can leverage historical analogies (CERN for physics, NASA) and talk about national interest and science. It sounds like the smarter, more sophisticated play.[1]
But I think that there are many problems with it.
What do you even mean by CERN?
Are you asking for:
Cross-posting from a Twitter thread responding to a recent viral comments by @Richard_Ngo about EA, Anthropic, and AI safety as a 'fake field.' Posting here because I expect this to be quite unpopular on LW.
(original thread: https://x.com/CRSegerie/status/2056737155880493357)
AI safety in 2023–2026 was driven by evals, threat models, scary demos, model-organism work, RSPs, and voluntary commitments. Richard calls this "much more of a fake field" and says it "won't generalize".
Here's why I disagree - 1/10
1/ I agree with Anthropic being now the biggest lever. They lead the AGI race, and Mythos moved the White House; this is quite a feat! But many of the specifics are wildly overstated
2/ Not a blind spot.
Empowering safety-conscious actors at the frontier was openly debated on the forum for years. Calling a deliberate/contested strategy a "blind spot" rewrites history. The bet was visible and explicit.
Personally, I've publicly criticized Anthropic on a few topics, but I still think the field is in a much better position, given that they're leading compared to the shady behavior at OpenAI.
3 /The effect of Anthropic leading is not just "AGI faster"
Anthropic has many positive external... (read more)
Epistemic status: I spent 5 hours thinking about this in March 2026. This is a Claude-written summary of a longer document, then improved manually. This post has been on my to-do list for a while; I did the 80/20 to get this out.
I buy the core case for AI labor. There may be a brief window between early AGIs and superintelligence in which AI could do orders of magnitude more safety-relevant work than humans alone, and we are already behind in preparing for it.
Tldr: What are we waiting for? Claude just formalized Fermat’s last theorem. OpenAI released PaperBench. People are already automating mathematics. Why don't we do the same in AI Safety? Nobody is pointing compute at replicating safety papers, rerunning the fine-tunes, testing the empirical predictions we keep writing down and never checking. A lot of slop is coming, and investing in evaluating it and identifying the potential nuggets it contains is probably very important.
P.S. Since writing the initial version of the post, Resolution has been announced, but they seem to be heavily focused on alignment. This might be a good call, but this post ... (read more)
On your first point, I think an important reason that we're making so much progress in mathematics as opposed to AI safety is that AI safety is not well formalized yet. It is impressive to see how much rigorous and intelligent Fable and Astra becomes when they have access to a Lean MCP. If you want to scale up alignment with automatization this way, I believe the most important thing is to formalize alignment directly, as in, allow the agent to have a test to immediately and empirically test its idea. The full blown version of this is Davidad's infrabayesianism and hypersimulations, but we can also do this on a shorter scale, with e.g. a framework that allows Claude/Codex to specify the behavior they want to observe, and run many simulations of it.
I already have a good meta-plugin to write such MCP, so I think I could do this in the span of a week, if anyone is interested to work on this, let me know.
Couldn't we privately ask Sam Altman “I would do X if Dario and Demis also commit to the same thing”?
Seems like the obvious thing one might like to do if people are stuck in a race and cannot coordinate.
X could be implementing some mitigation measures, supporting some piece of regulation, or just coordinating to tell the president that the situation is dangerous and we really do need to do something.
What do you think?
It seems like conditional statements have already been useful in other industries - Claude
Regarding whether similar priva
Shamelessly adapted from VDT: a solution to decision theory. I didn't want to wait for the 1st of April.
By Claude 4.5 Opus, with prompting by Charbel Segerie
January 2026
Moral philosophy is about how to behave ethically under conditions of uncertainty, especially if this uncertainty involves runaway trolleys, violinists attached to your kidneys, and utility monsters who experience pleasure 1000x more intensely than you.
Moral philosophy has found numerous practical applications, including generating endless Twit... (read more)
Dario (and the other CEOs) had this massive power from the start.
In Delhi, Dario had 5 minutes to address 20+ heads of state.
- He spent 𝟰 𝗺𝗶𝗻𝘂𝘁𝗲𝘀 𝟰𝟬 𝘀𝗲𝗰𝗼𝗻𝗱𝘀 pitching Anthropic's Bengaluru office and Infosys partnerships.
- He spent 𝟭𝟯 𝘀𝗲𝗰𝗼𝗻𝗱𝘀 on risks.
This week shows that we have more agency than we think. Let's use it --> There is a channel to 900M weekly users. What goes in it?