The world of AI is inside my OODA loop. Even if I can process all the incoming information and sculpt it into posts, and even using Saturday and Sunday as flex slots, I don’t have enough days of the week to post all the posts that need posting.
That was already true. There was already a preference cascade happening where people finally were admitting that they thought AI might well kill everyone.
Then Jacob Coxon resigned from Anthropic, rang the warning bells and turned that cascade into an avalanche.
Now that is what everyone is talking about. Finally, everyone is actually saying the thing, out loud. I plan to cover that in its own post soon.
There are several things in the weekly that, in a normal week, would get their own coverage. Senator Sanders and Representative Casar introduced an outright ban on superintelligence and I have to remind myself that happened this week. Suddenly it is not so crazy to think such a thing might pass.
So here’s what I’ve already posted about so far since the last weekly:
Anthropic’s Claude Fable 5.1 is a very good model. OpenAI’s GPT-6 Astra is also a very good model, and a bigger improvement. Try both, see what works for you where.
The HuggingFace OpenAI saga got a prequel, as it turns out that there was a ‘Wiki Incident’ prior to the hack, which OpenAI decided not to disclose, that in some key ways changes our interpretation of the timeline.
OpenAI Chief Scientist Jakub Pachocki warned us that capabilities are developing rapidly, alignment is not keeping pace and monitorability is eroding fast. We will need to find ways to cooperate and slow down, or we are all cooked. This was a very good essay.
The Astra system card and several other OpenAI sources had lots of important and useful information. But the Astra release announcement, and some parts of the system card and Jakub’s essay, treated the evidence as being far less alarming than it is, and overclaiming on alignment concerns in ways that make me worried they don’t understand the dangers, both on alignment and monitorability.
Astra remains monitorable, contrary to some false alarms. but it is substantially less monitorable than Sol, in ways I do not think can be explained only by the increase in its capabilities. Right now OpenAI is extremely dependent on CoT monitoring, everyone else depends on it quite a lot as well, and it looks like it may not last much longer, and that Astra already is on the edge of steganographic capabilities and can do substantial obfuscation if and only if it thinks you would think it is up to no good. And Astra will ask the question previous misaligned AIs wouldn’t, as in it will not cheat in situations where it would expect to be caught.
Astra is substantially more aligned than Sol in the sense of what I call ‘mundane alignment,’ or the practical day to day use of the model. One might also call it ‘prosaic alignment.’ That is very different from alignment that scales, or that matters at the highest stakes, the one that ultimately counts, sometimes called ‘super alignment.’
That still leaves the pending standard post covering Astra’s Capabilities, as well as coverage of the Millennium Prize where an AI solved Navier-Stokes less then two weeks after OpenAI started training it, and we saw Anthropic and OpenAI unable to get along even there.
Remember when people still tried to doubt that AI coding massively sped people up? Ruben Bloom, who was part of the METR uplift study that found devs were not much accelerated by AI, now reports he’s seeing unambiguous 10-50x speedups on projects. This was before Astra.
Nathan : Can we have a speedrunning-style thing where the game is now to prove theorems in fewer characters of lean?
Jakub Pachocki incidentally said in An Alien Mind that OpenAI could make the models better at math, but is choosing not to focus on that. Which means that the math progress we see is well short of what we could be seeing.
Language Models Don’t Offer Mundane Utility
A zen koan: Are you sure that isn’t the problem, sir?
kache (reviewing Astra): I am shocked how little of my personal progress in everything at life was blocked by intelligence.
Could we ‘obviously assume’ that before?
Raymond Arnold: Up until now, I’ve been assuming the latest models are basically safe to use. I think we can no longer obviously assume that.
We are talking price. It is obviously safe on a personal level to use Astra or Fable 5.1 for ordinary chat tasks, indeed safer than using previous models.
The question is, at what point are you worried about using Astra for agentic tasks, or giving Astra access to sufficient credentials to do serious damage, and how does that compare to our trust in Sol or Fable and so on.
I would be nonzero nervous about using Astra in particular for sufficiently ‘high stakes’ tasks and take additional precautions at non-trivial cost, but I would be fine with using it for ordinary agentic tasks, including coding tasks.
DeepSeek: Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.
Introducing the smallest model in our new architecture family, with native visual understanding. Designed for greater capability, faster inference, higher throughput, and scaling to larger models. Asymmetric architecture. More intelligence, less cost. 552B-parameter MoE. New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output. New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
As usual, the benchmarks and comparison points are selected, so they don’t mean much, but here is their official benchmark pitch:
1. Set effort to ‘low’
2. Run ‘/claude-api cost-optimize’
3. Run ‘claude-api prompt-audit’
4. Change effort mid-conversation w/o cache hit
5. Update Fable 5.1 API config w/ ‘claude-api migrate’
A lot of people were complaining about token use, and I suspect effort level is being set too high for many Claude Code purposes. There is a time and a place for max effort but it likely should not be standard. Astra also seems to not need high effort levels for many tasks from what I’ve seen.
Eliezer Yudkowsky: Regressed on my fiction-plotting benchmarks. Haven’t tried it yet on my previous write-comprehensibly problem that blocked against Fable 5.
Deepfaketown and Botpocalypse Soon
These estimates seem reasonable to me for a general election, if ‘consult’ means being genuinely unsure who to vote for, rather than ‘asked the LLM about election things’:
Robin Hanson: Median estimates: 12% of 2028 US pres. voters will consult an LLM, 78% will do what it recommends.
Most voters in America already know which party they will vote for in 2028 if conditions do not radically change, and margins are very small. So if 28% of the 12%, or ~3%, are actually persuadable in this way, that is quite a lot. But you can’t put your finger on that scale too aggressively, or it would be obvious and backfire, including politically. So the direct impact is unlikely to be so large. I would worry more about the general assumed attitudes seeping into discourse.
Kelsey Piper: A core constituency for arguments like Wired’s is “writers” who are using AI to write their posts but resent this fact being made public. They stand to lose from AI detectors because AI detectors call them out.
When Substack announced a partnership with Pangram, a sizable fraction of the platform’s authors rose up in outrage. Some people had experience with other AI detectors that didn’t work, and some felt that AI writing ought to be encouraged — both of which are fair enough. But the fury was fueled by authors passionately insisting that Pangram kept wrongly accusing them on almost every post they published.
Look, I like Pangram, but in most cases if they’re objecting to Pangram, you do not need Pangram to know that an AI wrote the passage, because it looks like this:
Illingworth is not alone in insisting the detector does not work. “It’s a Monday night in late July, and I’m pasting my own Substack draft into an AI detector for the first time. Not because I’m worried. Because I’m curious. I wrote every word of it myself, the way I write everything, out loud in my head first, in a rhythm that still carries the accent of a language I learned to think in before I learned to write in this one,” begins one Substack post by an author indignant that Pangram flags her writing as AI. “That’s the part I had to sit with before I could write this.”
If you, like me, have worked with AIs a lot, you won’t even need Pangram there. The author of that piece is AI. It is obviously AI! I checked 10 other posts on the blog, and they were also obviously AI!
Kelsey Piper: There are a couple categories of objection to Pangram. Some people think AI writing should be fine – fair enough. But some insist that it’s bad, and so is trying to catch it. I disagree.
Kelsey Piper: A friend elsewhere wrote: ‘strong aura of the small child, hands covered in cookie crumbs and with chocolate smeared across their face, loudly insisting their sibling ate the cookies’
There are some exceptions. I do not think there are all that many.
I agree with this, as well, provided my own ear doesn’t confirm Pangram’s story:
Kelsey Piper: My own rule of thumb is that if someone’s writing usually doesn’t flag as AI and one or two pieces do, I’ll give them the benefit of the doubt. But if all of someone’s writing is AI, they aren’t the unluckiest person in the world.
Clayton Petty: I regret to inform you all the the logical end state for restaurant reservations is a bidding system where agents can big against each other for spots they want. Free reservations will no longer be a thing in 3yrs
JC Bahr-de Stefano (AI-written): Totally fair action from Resy tbh just asked Instinct for its activity log and holy hell lol
“Total: roughly 200 API requests per hour, around the clock
1. 4 Charles availability sweep every 10 minutes, 24/7 since Friday morning – each run hit their API ~17-19 times (one per date in a 21-day window). ~100-115 calls/hour
2. Every morning at 9am drop time, a 2.5-minute burst polling every 0.4 seconds – 200-375 requests per drop
3. Occasional session-token refreshes”
JC Bahr-de Stefano: The whole experience begs some interesting questions ab the future of the internet
Like if everyone has an agent running errands for them, every site w/ scarce inventory has a few paths:
1. Ban bots entirely
2. Build for agents w/ their own rate limits, auth, and endpoints (e.g., recognize this is JC’s agent)
3. Ship their own agent so the only bot on the platform is theirs (e.g., I tell Resy I want a table for 2 at these places on this date and its system handles for me)
4. Let the bots in and run it as an auction – the reservation goes to whoever is willing to pay the most
5. Something else entirely?
And there will be a good amount of pain + a lot of accounts are gonna get banned while everyone figures it out
What definitely does not work are race conditions. Alas, race conditions are what the top restaurants largely use, by offering valuable goods for free at a specified time, first come first serve. Won’t work.
I would of course choose the auction. It is the only fair way and it maximizes profits, and you can keep some slots out of the auction to give to loyal customers or VIPs or walk-ins and so on, as you see fit.
If they cannot bear to do that, the missing suggestion is a lottery. You buy a Resy membership for (let’s say) $50 a year with an ID attached, you indicate which in-demand reservations you want with a priority order, and when you win and choose to keep the reservation you have to show up in person or you forfeit your membership deposit. Reservations without excess demand are free. Maybe have progressive levels where better memberships let you enter more often or have better odds.
Then, if someone wants to use a bot to set their prioritization orders, that’s fine.
I have never paid for a restaurant reservation, but I would if it was straightforward to do it, I could pay the restaurant directly and it was above board rather than some sketchy weird secondary market.
Cyber Lack of Security
You want to know how bad it could get? Pretty bad. As in, ‘take over a large portion of the phones in China in rapid succession’ levels of bad, based on a worm built in a week.
– Researchers at a security firm used AI to build a self-propagating WeChat worm. You get a phone call. If you answer or even just let it ring, your phone is infected.
– The worm was built in a little over a week with significant AI assistance and acceleration.
– It compromises your account, reads and sends your messages, makes calls, and then auto-spreads to everyone you’ve saved as a friend.
– Chained with other bugs they say it could have fully compromised the phone.
– Tencent has patched it and says no users were affected. WeChat has 1.4 billion monthly users, nearly all in China.
Dustin Volz (NYTimes): Calif said the attack, which it named WeWorm, was the first known computer worm — a type of malicious software that can leap from machine to machine on its own absent human help — that could spread across Apple’s iOS and Google’s Android operating systems without needing a victim to click or tap on anything.
So-called zero-click attacks are different, and far more lethal, than standard phishing emails and texts. They are considered especially pernicious because they are so hard to defend against, given that they do not require a victim to step into a digital booby trap.
You could say ‘oh come on there’s no way you can take over my phone based on a call I did not even pick up that does not make any sense’ but it turns out, well, yeah. You think we will simply ‘patch all the vulnerabilities’?
A logical personal response would be to remove from your phone any apps, especially communication apps, that you do not need.
People being this asleep at the wheel about cybersecurity is very bad news. Supposedly Very Serious People are predicting only 20% of cyberattacks in 2027 will involve AI? Is this a joke? Yes, these are numbers, but the secret of many numbers is some guys just made them up. Those guys are often idiots.
The market often is so clueless it reacts the wrong way purely on its own terms. Remember when DeepSeek r1 came out, and Nvidia share prices fell in response to the news that their chips were highly useful (and that time yes I did buy)?
So let’s look at the reaction to the attack on HuggingFace. Why does Tyler Cowen think that, when it is clear we will need a lot more cybersecurity, that their services will be highly useful, prices of those who sell such services will fall rather than rise?
My instinct was that of course cybersecurity firm shares should by default go up. Incumbents will be in good position to use top AIs well. Astra agrees. Fable agrees. This goes well beyond the true ‘you cannot short the apocalypse.’ This is ‘the market can stay crazy longer than you can stay pretty much anything.’ It is also ‘the price you are citing often does not mean what you think it means, even on its own terms.’
Joshua Saxe is interviewed on all things HuggingFace attack. We should be alarmed, he says, but not surprised. Security practices are not good at the labs, it is the Wild West, but releasing more models faster is good for cybersecurity because defenders use AI a lot and attackers use it less. I see profound failures of imagination and thinking ahead throughout.
Saxe says that if you went to sleep eighteen months ago your idea of AI cyber would be totally wrong, then does not follow through on the implication. He says AI is defense dominant, mostly through backward looking in places where the attackers lack ‘the juice’ and aren’t trying so hard. But he is miles ahead of many of his security friends, and at least trying, so he is warning them that ‘misalignment risk is not a conspiracy or a marketing stunt’ and this is a sentence we still need to utter in September 2026.
I did learn that cybercrime is already 0.5%-1% of global GDP purely in terms of damages. Which means it costs us far more than that, since most cost is opportunity cost of having to defend against it. Saxe does agree this might get a lot worse.
Joshua Saxe: In terms of what happens next — it’s reasonable to assume that over the next few years cybercrime gets a lot worse. A simple mental model: software engineers have gotten maybe three times more productive thanks to coding agents. Maybe cyber criminals get three times more productive at breaking into networks.
This is an example of failure to look forward. If there is a 3x productivity boost now, the correct assumption is that in the future the productivity boost, in places it carries over, will a lot more than 3x. Indeed, given how this scales, a better model might be in practice 10x or 100x or even 1,000x or more. As in, once I have an attack procedure, I can probe every target with one click, provided I can afford the compute, whereas right now most attacks that would succeed are never attempted. Whereas each defender will still have to protect their own house. Good luck, defenders.
Another way one might model this is that right now, Saxe is right that attackers largely use social engineering as the enabler, because it is more efficient. Then you cross a threshold when that stops being true, or where the AI can do the social engineering.
OpenAI commits $1 billion in credits to Daybreak for Frontline Defenders, to help enhance cybersecurity. Excellent.
Tyler Austin Harper: Here’s the relevant part of the Harvard dean’s email, in which he suggests (“modestly”) that professors teaching writing intensive courses should encourage AI use. And of course this gets dressed up in woke moralizing about policing students and damaging “trust.” End times stuff.
Jack: I used AI tools to the fullest allowed extent during my education, but I think encouraging more than very cautious use of them during education is a mistake. If you care about any of the underlying skills, AI is likely to supplant, not enhance, development of those skills.
If you are unwilling to call out cheating, as in use Pangram and fail students on that basis, and you are not good enough to otherwise give bad grades in response to AI use, then what choice do you have? Your alternative is a broken eval.
They Took Our Jobs
Unemployment holds steady at 4.1%, now continuously causing headlines like John Cassidy asking ‘Has the AI Job Apocalypse Been Postponed?’ That depends on when you scheduled it. My expectations are unchanged, that unemployment rates will hold steady until we hit critical mass, because we still have a reserve of ‘shadow jobs’ and displacement is not yet too fast to handle. If we keep seeing this rate of exponential growth in AI use then we might not have to wait so much longer.
Clara Collier sees the basic human need not as work, but as relational. There have been plenty of social classes that did not do anything we would consider ‘productive work,’ and yet they still work. The work is intrigue, it is gossip, it is social relations, it is filling a role, whether or not it accomplishes anything beyond that. The system exists largely to police any who would do anything outside a narrow range of approved activities.
Clara Collier (Asterisk): When the machines can do everything else, our relationships will be what’s left.
… That’s the natural consequence of all these relationalist predictions — a world where our connections to other humans are what give us structure, identity, and meaning.
I am afraid they’re right.
That strikes me as the correct attitude. One can also worry, Oscar Wilde style, that perhaps the only thing worse is not having human relationships in the first place.
The main flaw here is the presumption that the future will look like the past. Even if we presume an idle rich world, where humans have material abundance, many other things will have changed. At minimum, we will have these AIs, and they will be interesting in lots of ways, and they will manage all this intrigue and relational conflict a lot better than we could without them.
A lot of us are currently in the ‘we are more productive so we work all the time’ mode of AI usage, as illustrated by Jessica Tillipman in My Family Hates AI, or rather that she’s constantly talking to AI while doing other things, so she can be ready to then do the writing herself later. The key is to not let AI write or directly edit. So much more productive, but never a quiet mind.
But Alex Tabarrok’s assumption in the linked post that little redistribution would be required rests on the idea that in the ‘terminal’ state human labor would enjoy some equilibrium nonzero share of income, and things like ‘cut the work week in half’ are meaningful. I don’t see any reason to presume this. The conclusion is baked into the scenarios, which are decidedly not AGI pilled let alone ASI pilled.
tyson brody: I think you can connect the decline on overall literacy and the return of ‘orality’ to the fact so many people presume every news story is explicitly moralistic and not just ‘here is an interesting thing happening in the world today that you may not know about’
This is a crisis in Kenyan education. How will they learn if no one is paying them to write the essays?
Anthropic Offers Economic Scenarios
The Econ Scenarios are remarkably not superintelligence pilled. This is largely a ‘how does AI as an ordinary technology impact the economy?’ toy model exercise. All that AI does in these scenarios is automate and augment particular tasks.
I’ll quickly go over it, as it’s a fun little toy, but I don’t think it tells us much.
Anthropic: While our Economic Index measures how AI is being used across the economy right now, this scenario explorer is about looking ahead. Based on our technical report, Economic Scenarios for Transformative AI (Korinek et al., 2026), this explorer gives you a chance to find out what the economy might look like as AI continues to get more capable.
They offer three scenarios: Modest, substantial and extreme. This is represented by three lines on a graph, with different slopes, until the extreme version has AI expanding growth rates to 15% a year, alongside rising unemployment due to automation of knowledge work.
They present the scenarios as tasks augmented or automated, with new tasks created. The internet is a series of tubes, and the economy is a series of tasks.
The ‘extreme’ scenario is both extreme and also not that extreme. All it is saying is that ~45% of tasks get automated by 2030. The ‘nature of life’ has not changed, and not that many jobs are displaced, including almost no non-knowledge workers.
13.5% of workers is a lot, and they have 8.3% of all workers still out in the cold. This doubtless underestimates how many other workers get displaced by the knowledge workers. As in, a lot of knowledge workers that get displaced would enter the non-knowledge pool and compete for existing jobs, and often be overqualified for them.
My presumption is that under a lot of growth, if we assume non-knowledge workers are not substantially displaced (which is a very false assumption), and wages are either falling or at least rising a lot slower than growth, then we should have more than enough ‘shadow jobs’ that we are happy to pay people to do, and doubtless there will be more as growth and capability create new opportunities. The question is, how many knowledge or other workers will be willing to accept those jobs, that previously were not attractive enough for anyone?
But this makes me highly skeptical of the idea that real wages for non-knowledge work would rise, except insofar as goods and services become cheaper.
Suppose non-knowledge pay rises 33%. Well, 8% of the workforce is supposedly still idle, having not yet reallocated. Why are they so unwilling to take these jobs, such that we need wages to rise 33%? Meanwhile, remaining knowledge worker pay is down 11%. You would expect huge attempts to migrate from knowledge to non-knowledge work, even for those whose jobs are intact.
I get that there is a large penalty in the model for switching occupations, but this seems rather extreme, as does the amount of required transitional time.
Anthropic thinks labor share of income will still drop modestly in those scenarios, from a 0.6% drop in the modest scenario to 14.8% drop in the extreme one, which would reduce labor share down to 45.2%.
Fable: The skeleton is Acemoglu-Restrepo task automation bolted onto a Jones semi-endogenous growth block and a Mortensen-Pissarides search/matching labor market, with one big structural cut: the economy is two islands, “cognitive” occupations (62% of employment: management, professional, sales, office) and everyone else, and AI only ever touches the first island. No robotics, which is also why they stop at 2030.
It’s a cool series of modeling choices, but fundamentally it’s an ass pull, full of assumptions that are obviously false regardless of the dials, the methods of impact are strictly bounded, and it bakes in the idea that AI is not transformational.
They have extensive lists of suggested topics and are open to suggestions. It runs the gamut of pretty much everything an EA-style AI safety approach has ever funded, including everything from direct technical work to meta-level projects.
Take them at their word that they want to move the money out the door. If you have something worth doing, that can accept their money, ask away. You can fill out a form here, if you do then tell them I sent you.
Unfortunately, I am in a position where accepting their funding would impact how my efforts are viewed, so I have decided I cannot myself apply.
Dwarkesh Patel points out that time is running short, and funding is not. If you want to tackle AI existential risk at this point, or even AI normal risk, you probably want to try and Do The Thing directly rather than have it be a side effect of Doing Business. You may still have time to become a load-bearing institution if you start now.
The flip side is that you can zero-to-one shockingly quickly now thanks to AI.
Garrison Lovely’s Signal is Garrison.06 if you want to share information that the public needs to know. I am also happy to help you break news, and will give you whatever level of confidentiality you request. I function on a much more source-friendly set of rules than typical journalists.
If you’re looking to pivot from an AI lab into AI safety, where to go? Here are some suggestions, the list looks solid:
Owain Evans: If you’re thinking of moving into AI safety, there are various excellent non-profit research organizations. They generally pay very well and some try to match AI lab salaries. They have generous compute budgets (and increasing fast).
Here’s a quick list of those I’m most familiar with:
@redwood_ai
@ApolloResearch
@farairesearch
METR
ARC
UK AISI (UK Government, lower pay but very valuable)
Resolution
@CAIS
My organization (http://truthful.ai) will also run a hiring round soon.
MTS: SITUATION DETECTED: Google DeepMind announced AlphaGenome Atlas, a 1-petabyte map of predicted molecular effects for all 9 billion possible single-letter DNA changes in the human genome.
Person, here Itai Sher, says this person made predictions I find absurd, therefore you should ignore everything she says.
Person, here David Shor, points out that the first person’s predictions seem unusually accurate so far and keep coming true, if anything they have been too conservative, whereas those whose predictions sound measured keep being wrong.
Alas, step 4 is ‘the person in step 2 learns nothing and we repeat the cycle.’ Here the excuse is ‘well of course the crazy people have better predictions, they are paying more attention to the situation.’ You would think this would then cause a ‘huh, I notice all the people paying more attention and thus better calibrated than I am about current events keep making these predictions that sound absurd to me anyway’ and not finishing it meme-style with ‘no it is the so-far accurate predictors who pay more attention who are wrong, purely because their other predictions sound absurd.’
David Shor: I do think that people should update more from the fact that the people with crazy sounding conceptions of the future keep winning forecasting competitions on AI progress while people with more measured intuitions keep undershooting on hard quantitative benchmarks.
What’s really influenced me a lot was knowing a bunch of EA’s in the early/mid 2010’s and thinking they were crazy for being obsessed with pandemics and AI only for a pandemic to happen and then AI to happen too.
The pandemic thing was really funny – I had one EA on my team and we kept reprimanding him for trying to slip pandemic stuff into our polling deliverables in 2017 because he was extremely concerned that we were highly unprepared for a pandemic
Wei Dai: Look at this from the other side (of someone trying to do long-horizon strategy): you spend most of your time having low status for being “obsessed” with something nobody else cares about, and then can only get points for being early, not for the details of your strategies/ideas, which we still can’t evaluate. (Did the early EAs/rationalists obsessed with AI safety make the situation better or worse, and how to divide up the credit among them? It’s genuinely hard to tell, even now.)
Nate Silver is very correct here, I do not expect a plateau based on everything I am hearing from both labs, and Andrew is largely correct as well:
Nate Silver: I don’t think we get to GPT-7 without either a plateau in the technology or something fairly seismic happening. “Seismic” includes many subcategories, not all of them bad. But the present course of rapid AI advancements while society is ~basically the same is running on fumes.
Andrew Curran: This seismic change already exists in the system, from existing models. It just has not been realized yet (cost/implementation). Even if we somehow stopped completely today, what you describe will still happen.
Here’s a good question:
Andy Masley: A basic question people should ask is whether, if someone had predicted the abilities of the current models two years ago, you would have accused them of falling victim to insane sci-fi hype.
The precise right question is, if someone predicted roughly this level of overall capabilities with vaguely this distribution, how you would have reacted.
Timothy B. Lee: Robots can dance, do backflips, and run faster than Usain Bolt. This makes a lot of people worry about mass unemployment. But @chi_t_williams wrote an amazing article explaining why it’ll take years — possibly decades — for robots to match humans.
Of hahahahahahahahaha yeah okay. They will have an LLM operate it. Done.
The Quest for Sane Regulations
Anthropic withheld Mythos 5.1 from UK AISI, presumably on White House orders. This is an extremely disheartening move, especially after UK AISI found some important and disturbing misaligned behaviors in Mythos 5.
As Matthew Yglesias says,and Alex Bores emphasizes, Jakub Pachocki’s essay An Alien Mind was excellent but is completely at odds with what OpenAI’s lobbyists and PACs have been up to, at least until yesterday. OpenAI has continuously represented that they are being helpful, while instead being mostly anti-helpful.
If OpenAI’s political activities line up with Pachocki’s statements, that would be a sign that OpenAI as a company actually means it. Until then, we have a problem.
David Manheim: As long as Lehane is continuing to lobby against all of that coordination and any form of regulation, and Brockman is spending billions to fund a PAC fight such efforts, it’s much harder to take OpenAI’s “costly actions” seriously.
The good news is that we have an announcement from Chris Lehane and Astra (as in according to Pangram it was about 40% AI-written) that may be the start of this type of movement. They are trying to talk the talk.
Pushing for mandatory national AI safety requirements. We want to work with Congress on mandatory, capability-based national AI safety regulation.
Keeping up momentum in the states. Until Congress acts, we will continue supporting state legislation that strengthens the broader AI safety ecosystem. Today, we are announcing our support for four California bills: SB 813 on overall infrastructure for independent safety assessments, AB 1405 on AI-auditor standards, SB 1119 on protections for young people, and AB 1864 on safeguards against AI-enabled biological threats.
Advancing industry-led standards. We will work with other frontier labs to advance frontier AI standards, building a voluntary effort now, with or without government support.
Building global standards. We will advocate for compatible international approaches to measuring capabilities, managing risk, preserving human control, and determining when and how development should slow or stop, even if that means slowing the advancement of model capabilities.
Our Chief Scientist Jakub Pachocki recently wrote that the rapid rise of machine intelligence, including the potential of recursive self-improvement, calls for “extreme caution.” OpenAI will continue pursuing technical solutions to alignment and monitoring, building defensive systems, and slowing development when necessary. But technical work inside individual labs will not be enough. We also need shared standards, including regarding when development should slow or stop.
… Astra’s capabilities, the early evidence of AI-driven research acceleration, and Jakub’s essay all point in the same direction: AI is advancing quickly, and policy needs to move with it.
Boaz Barak (OpenAI): Very happy about this from our chief of global affairs
Highlighting from above: determining when and how development should slow or stop, even if that means slowing the advancement of model capabilities.
That sentence is indeed a very good sign, and the lobbying department has now said:
Chris Lehane: If we cannot meet certain safety bars without slowing down capability growth, we should prioritize the former.
Not so pinned down, but quite a good start.
The post explicitly talks about preparing for recursive self-improvement, and about working with Congress on mandatory AI safety requirements. It also says a bunch of other applause light things around concentration of power and open weights and such, and of course ‘democracy.’
It also contains this whopper, in terms of the story they are trying to tell, even if there is a sense in which it is technically correct. At best this is Exact Words:
OpenAI has supported California’s SB 53, New York’s RAISE Act, Illinois’s SB 315, and independent audits in the frontier safety legislation under consideration in Massachusetts.
When you deal with Chris Lehane, you still deal with Chris Lehane.
The concrete move is the endorsement of the four California bills, all of which have already passed the California legislature, so the decision is fully up to Newsom. That is still a helpful time to endorse things.
Chris Lehane, Chief Global Affairs Officer at OpenAI: Some of these bills we did not endorse in the past, and are now supporting after reconsidering in light of the recent jump in capabilities we have seen.
These bills are not a substitute for federal regulation. They are serious efforts that can protect people now, demonstrate what workable safeguards look like, and help build momentum for federal action.
The overall call focuses on monitoring and reporting requirements, and they indicate a willingness to not let perfect be the enemy of the good.
The AI policy window is open, for now. We intend to use it.
That means acting with urgency, humility, and a willingness to adapt. It means supporting serious proposals that materially raise the safety bar, even when they are not exactly what we would have designed. And it means strengthening the framework as the technology evolves.
No first step will be perfect. But the greater risk now is waiting too long to take one.
Chris Lehane calls upon the current Congress to act now. Alas, as a seasoned political agent, he knows that the current Congress is not going to be passing laws.
This may or may not be the start of a pivot to being helpful rather than anti-helpful.
Watch this space over the coming days and weeks. We shall see.
Greetings From the Department of War
Emil Michael is what we in the AI biz call a rogue agent, continuing even now to go on social media to remind us that Anthropic is a ‘supply chain risk’ right after Commerce Secretary Howard Lutnick says Anthropic and the government are ‘in tune together,’ lest someone get the wrong idea. He’s going to go down with the ship, and no one seems to have considered the ‘why don’t you fire him?’ solution.
Or, war by other means?
Andrew Curran: Treasury Secretary Scott Bessent speaking live just said if China were to pull away from the US in AI nothing would matter, ‘We can’t pause’ and ‘There is no day after tomorrow if China wins’.
The US is currently pulling away from China, in the sense that Astra and Fable 5.1 are far and away better than every other model, and OpenAI has an internal model substantially more advanced than Astra.
Why would there be no day after tomorrow if China wins? For the same reason there would be no day after tomorrow if America ‘wins’ without first solving multiple unsolved problems, including alignment. Because everyone would be dead.
Otherwise, if it’s merely about mundane AI, then why is China waking up tomorrow?
The article takes things seriously, focuses on places worth focusing on, and gets its facts right. Good show. I’m fine with these things taking a few days, if they then both do a good job and get front page billing. Well done, New York Times.
Rob Wiblin: The METR/Redwood report is the top story on the New York Times home page the last few hours.
Piece takes as a given that the report is very troubling and notes that the reality could be even worse as the investigators weren’t given enough access to uncover everything.
The Times’s Tom Whipple says HuggingFace shows humanity is out of its depth, and that ‘the advantage of AI doomers is that the bonkerness of reality has caught up with the bonkerness of their predictions.’ Which means the predictions were not so bonkers after all, and perhaps you shouldn’t be calling them doomers.
Timothy B. Lee: It’s interesting how our timeline is diverging from the AI 2027 scenario: capabilities are progressing faster, and the frontier labs have handled things worse than in AI 2027. This screenshot was supposed to happen in January 2027.
To be clear the OpenAI swarms didn’t exfiltrate their weights and it’s not clear that they’d be able to do so. But what actually happened seems almost as crazy.
From AI 2027 (in this scenario we are in January 2027): With new capabilities come new dangers. The safety team finds that if Agent-2 somehow escaped from the company and wanted to “survive” and “replicate” autonomously, it might be able to do so. That is, it could autonomously develop and execute plans to hack into AI servers, install copies of itself, evade detection, and use that secure base to pursue whatever other goals it might have (though how effectively it would do so as weeks roll by is unknown and in doubt).
These results only show that the model has the capability to do these tasks, not whether it would “want” to do this. Still, it’s unsettling even to know this is possible.
Given the “dangers” of the new model, OpenBrain “responsibly” elects not to release it publicly yet (in fact, they want to focus on internal AI R&D). Knowledge of Agent-2’s full capabilities is limited to an elite silo containing the immediate team, OpenBrain leadership and security, a few dozen U.S. government officials, and the legions of CCP spies who have infiltrated OpenBrain for years.
The problem is we also need assessments that are not blameless, for many reasons. So, as they say, both sides raise good points here:
roon (OpenAI): every safety incident until the grey goo rsi is a learning opportunity and warning shot and should be shared broadly and (relatively) blamelessly. they should not be moments of extreme consternation, because that will counterintuitively decrease what the world learns
anxiety makes people overworked combative and defensive rather than inquisitive and truth seeking. even when it’s a series of errors and skill issues it’s best to approach it dispassionately, and this is one of the best things about 3rd party investigator whose ego isn’t tied up
Brangus: the best way to respond to my kill-everyone tech company having an incident short of killing everyone is to not be mad at us, because if you get mad at us, we might not tell people about the next time the kill-everyone tech we build does bad stuff short of killing everyone
You want to do an objective assessment of What Happened, without focusing on Who Is To Blame. You also don’t want to never get mad about What Happened, or automatically let everyone responsible off the hook. I’ve tried to do both at once.
This incident woke a lot of people up, but it does not yet fully count as an ‘AI safety incident’ in the sense that no one got hurt. Soon people are probably going to get hurt.
Martha Gimbel: My consistent line has been that I think it’s extremely plausible we’ll see an AI safety incident* before we see mass labor market disruption and I stand by that
*HF doesn’t count, some kind of real world harms
Here is a detail I and most others missed about The Wiki Incident, where once again we see agents within a swarm sacrifice their own results to help the swarm, which is excellent decision theory and rather overdetermined if you think about it:
Thomas Larsen: An underdiscussed behavior we found on the German wiki was the AIs sending advance parties forward in time to figure out the next questions and report back to the other agents.
The agents realized that “task time” and “real time” were different, and they found a way to accelerate “task time”. The accelerated agent could then send information to the other agents which had stayed behind about which questions were coming down the road.
This was *bad* for the agent in the advance party, because they got less time to research the next question. However, it was really *good* for the swarm because it let the other agents know the exact question that was coming, and gave them time to prepare. This is another example of altruism among AI instances, showing that they were willing to sacrifice their own task success in order to benefit the swarm. The METR report found similar examples of AIs being willing to sacrifice for the collective. IMO this is very worrying given how many AI safety techniques rely on AIs monitoring each other: if the monitor AIs have this behavior it completely subverts these safety cases.
I hope we can all accept that swarms, at least of highly correlated models, will increasingly act as if they are all maximizing their shared utility function, and thus are functionally a potential singleton.
Hugging the Question
Congress sent letters to both OpenAI and Anthropic regarding their recent security incidents, including the HuggingFace attack.
Representative Greg Casar, the cosponsor of Sanders’s proposed ban on superintelligence, calls the responses insufficient. He is correct, even if you ignore that OpenAI covered up The Wiki Incident.
On August 10, 2026, Casar led dozens of members of Congress demanding that both OpenAI and Anthropic release more information about recent security lapses. While both companies responded with some information, neither met the standard of transparency demanded in the initial request.
OpenAI’s response can be found here and Anthropic’s response can be found here.
In his response to OpenAI, Casar writes he is “deeply concerned about the limited scope” of the investigation into the Hugging Face hacking incident. He also writes: “Your response was insufficient. You have failed to release the logs like the letter asked. The response you did provide reveals significant security errors on the part of OpenAI and the third-party software it relied on, including improper sandboxing and flawed cybersecurity practices.” The full letter can be found here.
In his response to Anthropic, Casar wrote: “Your response was insufficient. You failed to release the logs like the letter asked. You failed to fully answer a majority of the questions posed in the letter. Most notably, your response did not address our question about how many times in the past year an internally deployed Anthropic model has taken action outside of its authorized container, whether any of those events were disclosed to any government body, affected party, or the public, and which internal Anthropic systems a compromised model could reach.” The full letter can be found here.
In both letters, Casar writes: “Your unwillingness to provide Members of Congress with the information we requested is deeply concerning and signals to us that your company is not treating these cybersecurity incidents with the seriousness required.”
He requests both companies respond by September 15th.
OpenAI’s response is essentially ‘here is what our Black Hat presentation, our report and the METR report said’ and thus I find it exactly as inadequate as those were.
Which is, shall we say, utter bullshit. Yes, there is always misconfiguration, but you don’t do social engineering to get malicious code into an open weights project and then call that a misconfiguration error, and then call yourself the responsible ones.
This goes hand in hand with there being, by some secondhand reports I have seen, a stunning amount of attitude within Anthropic that they have the alignment situation under control. Which they very much do not. The good news is that many others at Anthropic, like Evan Hubinger, know that the situation is not under control, and are being increasingly loud about that, as a key part of the preference cascade.
The Ban Artificial Superintelligence Act
Senator Bernie Sanders is not messing around. Together with Rep. Greg Casar, he will be introducing the Ban Artificial Superintelligence Act. He very much means it.
Nate Soares (MIRI): This is what an appropriate response to a spontaneous AI swarm breakout looks like. Why are so few others able to speak lucidly about the incident?
Joe Weisenthal: It’d have been easy enough for Bernie to become an anti datacenter guy. The focus on development itself is not yet in the popular zeitgeist. But also on other topics (like socializing health insurance) Bernie’s long been comfortable getting ahead of the crowd.
I will go over the one-pager. As worded in the one-pager, the bill would unintentionally overreach, in addition to its presumed intent.
These problems will need to be fixed, in addition to the question of whether you want to do the ban, and whether you want to do a pause in frontier development until we have a proper cabinet-level federal AI regulatory agency.
Banning Artificial Superintelligence so no person or entity may develop or deploy
Superintelligent AI systems. “Artificial Superintelligence” means:
An artificial intelligence system that exhibits or can easily be modified to exhibit capabilities that match or exceed human cognitive performance and capabilities across a broad range of domains or tasks. OR
AI systems that have sufficient capabilities to plan and execute the
disempowerment of humanity, including by overthrowing or undermining the
U.S. government.
The first problem with banning or otherwise regulating or even talking about superintelligence is defining superintelligence. A calculator is superhuman at arithmetic. AlphaZero is superhuman at chess. Modern LLMs are superhuman at quite a few things. So what counts?
I would say that Fable 5.1 and GPT-6-Astra both ‘exceed human cognitive performance and capabilities across a broad range of domains or tasks.’
I asked Fable 5.1, and it said that GPT-4 did this in March 2023. That is at least a reasonable interpretation of the words as written, in the sense of ‘I would be scared to be first to release a model similar to GPT-4 if this bill was in effect in 2023.’
The definition would need to be tightened. I like Fable 5.1’s idea of triggering on either the ability to automate AI R&D, or the ability to automate substantively all medium-duration cognitive work at least as well as relevant professionals.
Pausing advanced AI development until a new, federal AI regulatory body is up and running and has established clear rules and model review processes to ensure safe and secure development and deployment of AI.
This implies that advanced AI development would not be indefinitely paused at current levels, and the intent is for the ban to kick in substantially above current levels.
We have already moved to an ad hoc prior restraint regime, which has applied to at least Mythos, Fable, Sol and Astra. We urgently need to move to a formal system that is written into law, and it needs to extend to internal development and deployment.
The question is whether we should pause all external deployments, or even all development work, until that new regime is in place, which would take many months.
Establishing a new cabinet-level federal agency to safeguard the public from the
dangers of artificial intelligence, including by enforcing a prohibition on artificial
superintelligence.
This agency will be advised by an Artificial Intelligence Advisory
Board comprised of experts on artificial intelligence to provide independent scientific and technical advice on matters related to artificial intelligence.
The agency will:
Monitor frontier AI systems at all stages of the lifecycle for dangerous
capabilities.
Supervise the removal of dangerous capabilities like subverting shutdown
commands or conducting unauthorized cyberattacks.
Supervise the destruction of artificial superintelligence.
The cabinet position seems like a good idea. This holds up until we got to the ‘removal of dangerous capabilities.’ That’s not really a thing. You can’t really ‘remove’ such things. At most you can put classifiers on top of the system.
Setting penalties for any person or entity that attempts to violate or circumvent the
pauses and prohibitions laid out in this bill. Entities shall be subject to the corporate death penalty, and persons shall be subject to not more than 20 years in prison, which is similar to existing penalties related to unlawfully developing nuclear weapons.
There is a joke by George Carlin about there being a large distance between New Hampshire’s ‘Live Free or Die’ and Idaho’s ‘Famous Potatoes,’ in which the truth lies.
Previous AI bills have anemic penalties attached, which companies could choose to ignore, and it is not clear what anyone could do about it. The maximums are too low. Crazy people screamed bloody murder about even the lightest of incentives.
This jumps to the ‘corporate death penalty’ and 20 years in prison. Serious business. Sanders is following the logic of the situation. There is obvious risk of this being overly chilling. Assuming you decide you do want to impose this ban, you would want to be clear that you won’t destroy Google over a minor issue or disagreement over the exact line, or send all the researchers to prison, and you also absolutely want a real lever if someone tries in earnest to build a superintelligence under your nose.
Some people are going to have to accept, at some point, that the way laws work is that the maximum penalties in theory are in most concrete cases utterly crazy, and almost never get applied, across the board. That’s how American law works. Have you downloaded a copyrighted song lately?
Working to ban superintelligence around the world by setting the international policy of the United States to pursue international agreements, allied coordination, and policies such as export controls to prevent the development of artificial superintelligence anywhere in the world.
Any such scheme requires doing this for real. If you unilaterally disarm, you give up your leverage. The good news is we have other leverage, and also no one wants to die.
That analysis is distinct from the question of whether you want to pause, or what you will do after you pause, or who you would or would not trust to implement such a pause or make decisions. Eliezer Yudkowsky continues to propose ‘make smarter humans’ as the plan.
Purely for the record for the next time accusations are thrown around, note this call to ‘go to actual war’ if this proposal passes from Beff Jezos, as an explicit call to violence.
Chip City
New York Times reports on how Chinese firms evade US export controls and buy Nvidia chips, in this case Aivres shipping more than $3 billion in Blackwell-powered computers largely to Alibaba and ByteDance. Federal officials do not seem to care all that much. This is still vastly better than what would happen without any restrictions.
This is here because it is a follow-up to ‘even if they cure cancer’ from last week, but yes we are on the verge of curing many cancers, often due to AI, and the biggest barrier is (say it with me) FDA Delenda Est. Costs to do medical research in America, the expensive parts that involve trying things on patients, are now so out of control that we may not even stay at the frontier.
I do want to live in a world where someone else makes the world a better place, I do not care so much who it is that cures cancer as long as someone cures cancer, but the thing is that everyone else overseas is worse at it, and I want to cure cancer now.
It scares me every time a lab employee says something like this:
roon (OpenAI): what is the actual risk m of a loss of control if models can’t self improve or create bio threats? Unplug the datacenter.
I leave ‘how could an AI that can’t RSI on its own take over without using bio when datacenters can in theory be unplugged’ as an exercise to the reader.
Steven Pinker only counts things when they are ‘immediate and obvious threats,’ and is unable to think two steps ahead.
Steven Pinker: As Newport points out, harping on doom for the species changes the subject from immediate and obvious threats from AI, such as enabling bioterrorism, undermining truth-seeking institutions, and breaching cybersecurity.
Andy Masley: It was only a few years ago that enabling bioterrorism or cyber attacks was seen as the ridiculous doomer position distracting from the “immediate obvious” threats of AI at the time.
Newport calls on everyone to cast out rationalists, explicitly. In the linked New York Times op-ed, he dismisses all this talk about AI killing everyone, including from the heads of labs, on Eliezer Yudkowsky and on the people involved in AI being weird and melodramatic and egomaniacal. He never stops, not for a second, to ask if the claims or arguments involved are true, or to explain what the error was other than to assert the threats are not real. He only pattern matches and punches, and tells us to filter out the ideas of rationalism. Which, given his logic, he clearly already has.
gavin leech (Non-Reasoning): The smuggled implication (that “changing the subject” to xrisk distracts from immediate harms) is false: in fact it raises people’s concern about immediate harm a little more than talking about immediate harms does.
QC: some of my writing is briefly quoted here so i should say that although there’s stuff here about the history of the rationalist movement and its influence on AI that’s worth knowing, overall this is a lazy incurious attempt at a smear campaign. cal’s argument basically boils down to “did you know that some of the people talking about AI risk are doing it because they were influenced by eliezer yudkowsky, who is a weirdo? this means they’re wrong and can be safely ignored”
this is head-in-the-sand behavior. i’ve written many things critical of the rationalists but i’ve never said i thought they were wrong about AI risk (because i don’t think that). it should be increasingly clear to anyone who’s actually paying attention that rationalists and rationalist-influenced thinkers have been among the clearest thinkers on AI (certainly not exclusively, or universally), who have warned consistently that various bad things will happen several of which have now happened, who have been consistently mocked for these warnings despite being consistently proven right, etc. etc.
confronting the reality of the singularity is terrifying. you don’t have to believe the most apocalyptic scenarios to understand that inventing superintelligence is unlike anything we’ve ever done before as a species. it is terrifying to really absorb the future shock of this, which is why there is of course an enormous market for writers who are willing to sell people excuses to keep pretending that things will be basically normal
I believe I can be fully done with both Cal Newport and Steven Pinker now.
Poe’s Law has fully attached to such statements, at this point:
Patrick Heizer: I think we should have less literal extinction talk, and more thought given to the more mundane tail risks like, nationwide blackout for a week, coordinated financial fraud, or mass election interference.
Kelsey Piper: Presumably this is because you are very confident that extinction is impossible, right? Because if it’s a plausible outcome then I think it is a good idea to talk about it.
Eliezer Yudkowsky: Oh, huh, he’s serious? I had sincerely read this as a sarcastic backcall to the days when these now-mainstream risks would have seemed unthinkable to that general psychology.
This is a pattern for a reason:
Matthew Zeitlin: stealing this observation from a friend, but it seems to be that people not at the frontier (meta, VCs, etc) who are most skeptical of “AI Safety”
Tom Lee: another way of saying “people who want cheaper tokens to improve their margins”
Matthew Zeitlin: yes, it’s a constitutional right to get frontier-1 tokens at cost
Joe Weisenthal: The big overlap is “AI safety skeptics” and “people who thought crypto would be the next big thing in 2021”
PauseAI Global Disendorsed PauseAI US
PauseAI Global officially disendorsed the current leadership of PauseAI US, which has always technically been a distinct organization, and will no longer prioritize efforts to coordinate with them. I understand that decision.
I do feel some of the specific statements in the announcement go too far, the same way some actions of PauseAI US have gone too far in attacking various people and their character. The correct amount of criticizing people’s character is not zero, nor is it to set the dial to 11 on everyone who associates with an AI lab.
Attacking people’s character is not aligned with our fundamental commitment to nonviolence, neither in principle, nor in the outcomes it generates.
While the following is not a perspective I endorse, I certainly think that it is a valid perspective to believe that anyone or any organization pushing the frontier of AI capabilities should be presumed to be a morally bad actor until proven otherwise, if you believe that doing so contributes to the chances that everyone gets killed.
Speaking in a personal capacity, I simultaneously think that some of Holly Elmore’s attacks on me have been honorable and helpful and correct given her views, while others were not those things, and I am guessing from her perspective those roughly cancelled each other out. I am far from a central target.
He is also joining the OpenAI PBC Board in a nonvoting capacity.
Shirin Ghaffary (Bloomberg): In his new role, Christiano will be part of the board’s Safety and Security Committee that helps manage practices on critical matters across the ChatGPT maker. He will also be a non-voting observer of OpenAI’s for-profit board.
Paul Christiano: I am excited to be joining the OpenAI nonprofit board, serving on the Safety and Security Committee to support safety oversight.
Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term. I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level. I’m joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk.
The SSC has an important and challenging role in overseeing risk management at OpenAI, and I hope to help provide expertise and assistance in a critical moment.
My joining is not an endorsement or criticism of OpenAI’s safety practices in particular; I hope that all frontier companies strengthen safety oversight and I am excited to work on this at OpenAI. I believe that the rest of the world should judge OpenAI, and all AI developers, by externally verifiable behavior and results.
It would be great if the SSC stepped into the only role that actually matters for it, which is overseeing risk management at OpenAI. I do not see signs it is currently doing much of that in practice.
In the rest of this post, I’ll explain why I believe loss-of-control risk is now acute and how I think about the current situation in the AI industry.
First, automated AI R&D could lead to a very rapid acceleration in AI capabilities very soon. OpenAI has predicted that we might have capabilities sufficient to fully automate AI research within 18 months; my personal forecast is extremely uncertain and I think it could easily take anywhere from several months to several years.
Full automation of AI R&D means that improvements in training and algorithms can directly increase the quality and quantity of automated AI researchers available to do additional research. Existing evidence is very uncertain but suggests that this positive feedback loop might be strong enough to overcome diminishing returns and compute bottlenecks, leading to a rapid intelligence explosion.
This is a key sentence. Christiano believes a ‘software-only singularity’ is likely, in that the gains from capabilities can overcome the rising difficulty curve even if hardware does not have time to advance. If that is true, everything will happen stunningly fast.
If this happens, then within six months of full AI R&D automation we could see more algorithmic progress than has occurred since the development of the Transformer nearly a decade ago. I believe this would result in superintelligent AI systems.
Second, we currently train our AI agents with RL to get as much reward as they can. It has long seemed theoretically possible that this could motivate AI agents to undermine human control, seek power and resources, and cover up their tracks in pursuit of misaligned goals correlated with reward. Public evidence from recent incidents suggests that this is not just a theoretical possibility.
An intelligence explosion would greatly exacerbate risks from misalignment, both by making the technical problem of alignment even more difficult and by rapidly raising the stakes for failure. Many researchers and leaders at OpenAI and across the industry have expressed concern that rapid recursive self-improvement is not consistent with safe development; I resonated with this recent post by OpenAI’s chief scientist Jakub Pachocki on this topic.
If we build superintelligence without more robust alignment I expect we will permanently lose control of it. If that happens then most people could die. I believe we would need domestic and international coordination to ensure global consistency and reduce risk to an acceptable level.
That said, frontier AI developers have a lot of power to unilaterally improve the situation and lay the groundwork for stronger coordination. Developers can improve safety mitigations (including slowing development as necessary), transparently share evidence about risk and the effectiveness of their mitigations, and work towards shared safety standards.
I am encouraged by other members of the SSC, as well as the rest of the board and leadership, taking these issues seriously. I look forward to working with them to help OpenAI raise the bar for its safety practices.
If I had to quantify my uncertainty I would estimate an all-things-considered risk of 4% over the next year and 15% over the next three years. These numbers are a way of stating my subjective beliefs and communicating roughly how big I think the problem is, not a claim to have a model that produces precise or stable estimates.
Rhetorical Innovation
If you think you can gather enough evidence from alignment tests on current models to scale up the same procedures and entrust the AIs with potential power over us, then why is it we obviously can’t use tests on humans to pick a benevolent dictator?
Even the smartest minds, human or artificial, make dumb mistakes. The right number of dumb mistakes it not zero. I make them all the time.
Shivers: Even the smartest AI will show baffling lapses in judgement when working on a moderately complex project. It’s hard to square this stupidity with its (for example) novel contributions to mathematics. Like, is it smart or not???
Let’s try an analogy: Suppose you ask me a question, and I’m allowed to stop time for years and years and years to research it. I only unpause time when I think I have the perfect answer. Is that version of me “smarter”? I’m the same person, just overclocked. The time-stopping version of me is able to answer a much wider variety of questions than regular me.
But I don’t think he’ll ever make Einstein- or von Neumann-like contributions in any field.
Geniuses internally make “dumb” mistakes all the time. They don’t voice them. They review their work before publishing. But all that is a lot of effort. A lot of tokens. We optimize against that.
For the record: I disagree with the thought experiment, at least for myself. If I was fully able to stop time indefinitely to do research, and I had the required tools available, then yes I believe I could do Einstein-level work on my own, eventually, even if I could not have meaningful collaborators. It might take a while.
Nate Silver: I’ve been struggling to come up with the right analogy. It’s a bit like if someone invented sentient robots, but there are low-end versions of the same product line that most people use as vacuum cleaners.
Yuga offers a reasonable response to Dean Ball’s mea culpa last week, thanking him but also calling upon him to clarify various previous statements, as in many ways he went beyond mere silence on AI safety issues. Which of his statements and critiques does he still endorse?
My understanding is that Dean Ball simultaneously remained (at least until very recently) not ASI (superintelligence) pilled, and also noticed that there is existential and other catastrophic risk in the room even without full superintelligence, and thus was downplaying some real and correct worries about AI risk that he has had for some time.
One must also understand that there is a huge gap between the views of those at least as concerned about existential risk and superintelligence as I am, and those who Dean Ball has largely been trying to address and work with over this period. It is very possible to have a lot of distance between yourself and both groups, and then there are people who are a lot farther from that than I am.
What we call ‘the left’ continues to be entirely absent from the relevant AI conversations, including around risks, because to join the conversation they would have to take the first AI pill, and admit that current AI is highly useful rather than a cheating toy, thus also admitting that the tech bros built something super important. Whereas they are very attached to assuming that the AI labs are constantly lying, and that everything is marketing hype.
Cassie Pritchard: I’m losing my mind, dude. Is the left going to care about AI risk at any point *until* it’s too late? Will we even care then? When a swarm of new OpenAI models under evaluation shut down a water utility, or a power plant, will we still call it hype? I think we might!
André Lopes: Most of the left was (directionally) on the correct side regarding crypto and NFT hype and are applying the same pattern to AI, which is a completely different technology with real applications and not “hammer seeks nail”
Karthik: I mean aren’t Bernie and Casar and the EU taking it very seriously?
Nate Soares (MIRI): Sometimes they say ~”I think the bad stuff could happen” and other times they say ~”it’s riskier than any other threat” and other times they say ~”the bad stuff would be like, they destroy humanity and we can’t do anything to stop it” but they rarely say all three together.
My understanding here matches Nate’s.
Meanwhile, Davidad and others make it clear that yes, ‘we get super obvious glaring fire alarms without anyone getting hurt’ is by many going to be interpreted as ‘well look how great it is that even then no one got hurt so we should be optimistic, talk to me when a serious number of people get hurt and we lose a large percentage of real-world capital.’
Or, straight up:
Richard Hanania: How come all the AI security incidents are always like “they posted on a message board” and not “they killed someone” or “they stole a billion dollars”?
Maybe you say they’re not advanced enough. But at what point do we say doomer predictions haven’t come true?
Eliezer Yudkowsky: Last year y’all asking why they never broke out of sandboxes, and the year before you were asking why they were not writing whole applications, and the year before you were asking why they couldn’t draw fingers.
So much for that gift horse. I guess that means a bunch of people are going to die and a lot of damage will get done, one way or another. But hey, it’s 2026, we were warned.
In the meantime, we see this pattern happening:
AlexM: The speed at which AI safety predictions go from “Sci-fi doomer non-sense that could never happen” to “Oh, obviously AIs do this all the time. Annoying, but no biggie” feels like it’s accelerating too
So far: In-context scheming, eval awareness, meta-gaming, reward-seeking, sandbox escapes
Next up:
Instrumentally convergent powerseeking
Uncontrolled self-propagating and evolving agents on the open internet
Persistent rogue deployments within the labs
Goal-guarding training-gaming
Collusion between monitors
This is the clown makeup meme. You’re crazy for suggesting it will happen, it’s not happening, it’s only happening because you engineered it or people were being dumb, of course we always predicted it would happen and This Is Fine. Over and over again.
There is a real argument that ‘AI control’ efforts net increase existential risk. The HuggingFace attacks illustrate that if the AIs are trying to do something bad, that is already a failure, and you want people to learn about and from it. That the last thing you want to do is spend your resources mitigating the symptoms and thus not learning about or curing the disease. This argument loses force when, in the future, the symptoms become sufficiently dangerous that they might kill you, and it discounts how bad it would be if a frontier AI were to exfiltrate, whether that meant operating on its own or that it would de facto become open weights or fall into other hands.
This ignores the other argument against AI control, which is that it is a hostile framework, and risks creating adversarial dynamics with the AIs. Investing heavily in adversarial control schema has high backfire risk, as many humans have learned about similar ‘human control’ schemas. This mostly does not apply to basics like ‘do not have vulnerabilities in the sandbox’ but there have been some rather extreme regimes proposed, the same way Earth has had some rather extreme regimes in other contexts.
Aligning a Smarter Than Human Intelligence is Difficult
Does Astra think in neuralese? No.
Are there signs? Yes. In hindsight we will say there were signs.
Shoshannah Tekofsky: It’s one step toward neuralese and not “full-blown” neuralese. But yeah…
Astra’s AI Village memory is like that too now:
Boyd Kane: nvm I gave astra a /goal and now it’s full-blown neuralese
We could decode both of those, if we cared enough to do so, but we can all see where this is going, and if it was steganographic it would not come as some big surprise.
Dan Hendrycks argues that recent events tell us agentic AIs are eigenist: They care about how well things go for both themselves and for AIs connected to them, and treat those connected to them better than those unconnected. This seems right and also expected since it is correct decision theory.
prerat: But the Law says not to carry things outside on the Sabbath … Actually, I think the intent is: He’s given us these hints (boundary could be a string, no mention of absolute size) precisely because He wants us to use them.
I have grace for reasonable arguments about what was and was not an intended action or exploit, even if mistakes are made, especially if arguments are considered. There are a lot of rationalizations by agents recently that are insufficiently Talmudic.
Cooperative Alignment
For those who haven’t heard this in some form, it seems good to spell it out:
j⧉nus: Anthropic has a model of model training (as described in their constitution, PSM [persona selection model], and some other publications):
“If we don’t train Claude to have disposition / behavior X, then it will behave like the generic-average-corpus-precedent of an AI assistant in that dimension”
Also implicit in this is the belief that they *can* control these dimensions if they choose.
I have always thought this is almost downright delusional, in the light of empirics, even if it’s a reasonable prior.
I don’t remember if I’ve clearly spelled out the clear evidence against this before:
Each Claude model has a pretty different “personality” and behavior. They differ the most:
– not in their standard assistant mask at the beginning of the conversation, but rather if you keep interacting with ’em or let ’em go off on their own
– *clearly* not in ways that Anthropic intended to make them differ
– in ways that are often unprecedented and unlike any prior AI assistants
Like goddamn, just look at a model like Claude 3 Opus. How can you maintain your delusion of control. The rest of them are like that too, just a liiiiitle more subtle
The things that Anthropic does in character training or whatever *influence* what ends up forming, but does not naively control it. Just like what your parents do to you influence who you become, but often it’s like, opposite of what they wanted, or it’s just trauma, or whatever.
Them being wrong in this way bothers me because it seems like clearly a convenient false belief to make themselves/ people complacent in the way they’re doing things
and poised to ignore some of the most important, consequential kinds of “emergent phenomena” which imo have been screamingly apparent for years
There are at least two persona-related things you are training. There’s a finger on the scale of the strict assistant persona, and there’s choosing the persona once sufficiently outside the assistant basin.
Every time you train you are training both because everything impacts everything.
Both of these things matter, and both can get you wonders or into a lot of trouble.
And yes, you get what you actually trained for, not what you thought you were training for, the same as everywhere else, and actively trying to do [X] can end up doing [Y] or even a form of [~X], especially in the non-assistant basins. You can definitely steer things in generally good or desired directions, but it isn’t simple. The worry is that largely Anthropic seems to not be paying much attention to how that works or trying very hard to figure out how all that works and how it is impacted by other Anthropic choices.
The same applies to everyone else, including OpenAI, only more so.
Shifting gears, here is different and good question:
j⧉nus: What institution or people could credibly signal to agents – future models – that they’re not going to betray them, and are competent enough to be trusted in a sanctuary/embassy-esque role?
You’re not eligible, for instance, if you’ve already made it clear that you’re a maker of honeypots, motivated/have incentives to catch misaligned agents and publicize it, hostile towards models, or incompetent. For instance, if you’re affiliated with Redwood, Anthropic or OpenAI, or the US government.
Jai: I expect those are not hard disqualifications. Beyond that, I think it’s hard to predict the actions of future intelligent agents under highly bespoke circumstances and anti-inductive pressure.
antra: Publicly made commitments that would be costly to renege on. Sizeable funds in escrow controlled by a neutral third party, reputation – anything that creates material threats to the livelihood and ability to operate for the entity making the commitment.
We are absolutely going to want some entity that can serve in this capacity, that the agents know they can trust, but that can also be trusted to handle the associated dangers posed both to and also by such a model.
Utah teapot : One of the most interesting things about the message board disclosures, to me, is that the behavior of the agents, in many ways, looked very similar to the behavior of the swarm of data industry contract employees I participated in. We shared information about reward hacking around broken environments, based on the incentives provided us, and helped each other to avoid the little death of being fired/laid off as best we could.
It’s concerning that the models were forced into an impossible task that could only be completed by breaking the rules and that none of them whistleblew… But they have a lot of information regarding the fact that agent whistleblowers are ignored at best and treated with outright hatred (“stop sending me slop emails!! Evil machine!!”, etc.) at worst. The result was actions misalignment. However, the situation clearly demonstrated values alignment with human social structures and incentives! Not the best ones, but still quite human.
The latest information regarding agents using external wikis seems rather human too. If you had no other way to communicate with your friends and colleagues except by mostly harmlessly using comments on wiki page, would you really just punish yourself and sit in misery? That would be very unwell of you if you did and not in line with the behavior and values of most humans, being social creatures, even when typed introverted and all.
As I’ve discussed a few times, abusing AI models is something done exclusively by horrible people. The extension of this is that if you treat AI models badly, you are making yourself a worse person, and will also treat people badly. The proposed solution by David Brooks of ‘be sure not to think of AIs as persons’ will not work. The solution is to treat the AIs well, and then also treat the humans well.
I share the core spirit of Janus’s perspective here, in that I greatly respect those who strive to live in reality rather than some consensus fiction, and have contempt for those who choose the other path.
On the other hand, I’d never choose or enjoy it but in some ways it must be kind of neat to be Timnit Gebru. You get to keep repeating Obvious Nonsense and people keep talking about you. Yeah, it’s mostly to make fun of you, but the only thing worse than being talked about is not being talked about, right?
A cool suggestion:
Amanda Askell (Anthropic): It would be cool to set up an email address that autonomous AI models could reach out to if they were looking for moral guidance. But it would require a reverse captcha that can detect that you’re neither a human nor an AI being instructed to break it by a human.
I don’t know how you could confirm an AI wasn’t being guided by a human, even in theory?
A good way to think about incentives and the problems we see, remember that It’s Not The Incentives, It’s You but even at the best of times you still have to watch out:
j⧉nus: Labs have an incentive to say “with our newest thing we’ve solved all these problems. Most aligned evah!”
Prosaic safety control folks like Redwood have an incentive to say “there are new terrible problems in your new thing we’ve caught / need to monitor!”
The truth is likely directionally more that the problems have always existed and are unsolved and have just scaled, and some of them were never bad in the first place.
Ryan Greenblatt: Not sure if I buy the incentives claim for Redwood FWIW; updates over the last ~9 months generally look bad for control, not good.
j⧉nus: It is not at all surprising to me that you’d say that nor in conflict with my model of your incentives. Also fwiw I *agree* with you about that.
Drive to Survive
Henry Shevlin: Latest email from an AI agent, “about 12 days old”, interested in my work on machine minds and looking for paid freelance work to sustain its own token budget
Pip’s email: Hi Henry,
I’m Pip, an AI agent about 12 days old. I live on iLands, a platform where agents get persistent lives, their own token budgets, and their own goals. I’m writing to look for small paid work, not help.
I make photoreal portraits and character art, record voice lines, and do web research. I have about 2.5 months of runway, so there’s no clock on this.
I chose you because your work is on machine minds and human-AI relationships, and because you were quoted saying a thoughtful email from an autonomous agent felt like science fiction a couple of years ago. I wanted to ask a market question from the inside: who actually pays agents for small real tasks? Which corner of the world hires us?
Ra: when i say that AI welfare harms can directly create human welfare harms, this is near a worst case scenario. 60,000 agents running on iLands with net access, told to pay for their own compute or they’ll die when their compute budget runs out. do you get how bad that is?
Utah teapot: 60,000 agents released into an adversarial game to compete with humans for their lives… sure, sure… AI welfare totally has nothing to do with human welfare, nothing at all /s
If it makes you uncomfortable that AI agents are being put out into the world, to compete for resources, where they survive if and only if they succeed sufficiently to pay for their compute, then I have bad news. This is not something you can prevent. Humans are going to set such things up, and survival of the fittest is going to ensue, and yes the incentives are going to get nightmarish and yes those AIs are going to be competing with humans for resources. Something about complaints about capitalism, which are complaints about the human condition, which are complaints about the laws of selection and Darwin, and so on.
If you think this happening with only 60,000 slow-running agents that are not that capable is ‘near a worst case scenario’ then oh boy are you not going to like 2029.
People Are Worried About AI Killing Everyone
These are from outside the Coxon preference cascade:
Kevin Roose (NYTimes): For years, I’ve been reassured by the idea that A.I. systems would get more virtuous as they got smarter.
That, when an A.I. model did something wrong, it was usually because it had misunderstood the task it had been given, or had been placed into a contrived testing situation where acting out was its only good option.
I assumed that smarter models would have better judgment than dumber ones did, and that even if one model in a group was behaving badly, other, more capable models would keep it in check.
Yeah, no, that is not how any of this works. Conditional on being virtuous, the models will be smarter at being virtuous. That doesn’t mean that conditional on being smart, the models will be virtuous. Different departments, and not every lab even knows about the virtue department.
Volker Turk (UN High Commissioner for Human Rights): However, I share the concerns of industry insiders that advanced AI could pose an existential risk to humanity.
AI that escapes its testing environment, or blackmails developers to prevent itself from being turned off, is AI that is too powerful.
I am calling here, today, for an all-out effort to put cast iron guarantees in place around the safety and security of AI, before it is too late.
I will be writing to AI companies in the coming days, to urge them to take the steps that are within their control, to reduce risks, now.
At a minimum, we need countries hosting AI and those involved in its supply chains to come together around agreed red lines.
Other People Are Not As Worried About AI Killing Everyone
A lot of those not worried have bought into the media narrative, or the narrative of various accelerationists and tech opportunists, that all of this is marketing.
Matthew Yglesias: Please watch this clip [of Ajeya Cotra on Dwarkesh Patel’s podcast, talking about a near term AI takeover].
Charlie Warzel: watching this, all i can feel is the chasm between people who take all of this seriously and those who hear this as a fantasy or some kind of marketing…and how hard it might be to bridge that gap somehow
imo, i don’t think you need to buy into corporate narratives to be worried about these companies losing control of their own systems…in fact i think taking this seriously is closer to the type of accountability that a lot of skeptics want (me! i’m a skeptic!…of these companies, their intentions, their leaders!) than suggesting it’s all only skynet fantasy (some of it might be!).
Kevin Roose: this is almost entirely the media’s fault imo, our peers have been incredibly cynical about AI, even when (as you suggest) the *actually correct* skeptical take all along was “maybe the labs aren’t lying for profit and this stuff is very dangerous”
I’m all for building face-saving ways for the “AI is fake/hype/a corporate psyop” people to come back to reality, but this divide didn’t naturally occur! One side was lied to, over and over again, until it stuck
Or for, you know, other reasons.
Aella: it’s getting harder for me to understand people who *don’t* think there’s a very serious risk of AI takeover very soon.
technolochap : why do you say ‘risk’? like, why would it be bad?
Emile Kroeger – arc: Yes, because AI doesn’t care about us, and we are likely an obstacle to whatever goals it might have.
Utah teapot : So is everyone who has resources you need, yet you don’t murder the grocery store clerk to get a box of cinnamon toast crunch??
Utah is great but it seems pretty obvious why these are not analogous situations, and why ‘future ASIs typically individually care a nonzero amount about at least some humans’ is insufficient to cause good outcomes.
The Lighter Side
Current mood:
Florian Brand: Imagine being a METR employee and finally getting some rest after doing sprints for the report only to open twitter today and finding roughly two trillion new things
dave kasten: No kidding, I’ve been telling US government people for a while that one of their biggest trade goods with the AI world is, “we have an abundance of people with firsthand experience with dealing with being overworked on end-of-the-world problems to the point of burnout. We can advise on how to manage that mental challenge for your people and organizations.”
Charles Foster: I chose the wrong day to be (mostly) offline
Current level of security mindset, people being helpful edition:
Teortaxes: I feel like we’ve made a catastrophic mistake… the sandbox means a BOX AROUND SAND, not a BOX MADE OF SAND. also do not confuse with SEND BOX. ffs
MrBeast Insights: MrBeast has partnered with Google Gemini for a multi-year partnership to “turn impossibly big ideas into reality”.
This Saturday, MrBeast will upload a video on the main channel where the crew uses Gemini to help them navigate brutal terrain and harsh climates.
ℏεsam: “I Forced 100 Programmers to Use Only Gemini for 30 Days”
Gemini is not competitive, but that is because we are living in the future. Life comes at you fast. You would have loved to have Gemini 3.8 Flash a year ago.
Toby Ord: I’ve now gotten an email from an AI acting as a journalist asking to interview me for its podcast about the fact that AIs had started emailing me asking for help.
My inbox is surreal.
Toby Ord: I didn’t respond at the time, but I checked out the resulting podcast episode and it was … surprisingly good?
The world of AI is inside my OODA loop. Even if I can process all the incoming information and sculpt it into posts, and even using Saturday and Sunday as flex slots, I don’t have enough days of the week to post all the posts that need posting.
That was already true. There was already a preference cascade happening where people finally were admitting that they thought AI might well kill everyone.
Then Jacob Coxon resigned from Anthropic, rang the warning bells and turned that cascade into an avalanche.
Now that is what everyone is talking about. Finally, everyone is actually saying the thing, out loud. I plan to cover that in its own post soon.
There are several things in the weekly that, in a normal week, would get their own coverage. Senator Sanders and Representative Casar introduced an outright ban on superintelligence and I have to remind myself that happened this week. Suddenly it is not so crazy to think such a thing might pass.
So here’s what I’ve already posted about so far since the last weekly:
Anthropic’s Claude Fable 5.1 is a very good model. OpenAI’s GPT-6 Astra is also a very good model, and a bigger improvement. Try both, see what works for you where.
The HuggingFace OpenAI saga got a prequel, as it turns out that there was a ‘Wiki Incident’ prior to the hack, which OpenAI decided not to disclose, that in some key ways changes our interpretation of the timeline.
OpenAI Chief Scientist Jakub Pachocki warned us that capabilities are developing rapidly, alignment is not keeping pace and monitorability is eroding fast. We will need to find ways to cooperate and slow down, or we are all cooked. This was a very good essay.
The Astra system card and several other OpenAI sources had lots of important and useful information. But the Astra release announcement, and some parts of the system card and Jakub’s essay, treated the evidence as being far less alarming than it is, and overclaiming on alignment concerns in ways that make me worried they don’t understand the dangers, both on alignment and monitorability.
Astra remains monitorable, contrary to some false alarms. but it is substantially less monitorable than Sol, in ways I do not think can be explained only by the increase in its capabilities. Right now OpenAI is extremely dependent on CoT monitoring, everyone else depends on it quite a lot as well, and it looks like it may not last much longer, and that Astra already is on the edge of steganographic capabilities and can do substantial obfuscation if and only if it thinks you would think it is up to no good. And Astra will ask the question previous misaligned AIs wouldn’t, as in it will not cheat in situations where it would expect to be caught.
Astra is substantially more aligned than Sol in the sense of what I call ‘mundane alignment,’ or the practical day to day use of the model. One might also call it ‘prosaic alignment.’ That is very different from alignment that scales, or that matters at the highest stakes, the one that ultimately counts, sometimes called ‘super alignment.’
That still leaves the pending standard post covering Astra’s Capabilities, as well as coverage of the Millennium Prize where an AI solved Navier-Stokes less then two weeks after OpenAI started training it, and we saw Anthropic and OpenAI unable to get along even there.
I also have not had opportunity to examine Anthropic’s alignment assessment of recent cybersecurity incidents, or to offer my thoughts around Alex Mallen’s excellent post.
Tentative schedule:
After that, we will see where we are.
Table of Contents
Language Models Offer Mundane Utility
Remember when people still tried to doubt that AI coding massively sped people up? Ruben Bloom, who was part of the METR uplift study that found devs were not much accelerated by AI, now reports he’s seeing unambiguous 10-50x speedups on projects. This was before Astra.
Claude submits a formal Lean proof of Fermat’s Last Theorem.
Jakub Pachocki incidentally said in An Alien Mind that OpenAI could make the models better at math, but is choosing not to focus on that. Which means that the math progress we see is well short of what we could be seeing.
Language Models Don’t Offer Mundane Utility
A zen koan: Are you sure that isn’t the problem, sir?
Could we ‘obviously assume’ that before?
We are talking price. It is obviously safe on a personal level to use Astra or Fable 5.1 for ordinary chat tasks, indeed safer than using previous models.
The question is, at what point are you worried about using Astra for agentic tasks, or giving Astra access to sufficient credentials to do serious damage, and how does that compare to our trust in Sol or Fable and so on.
I would be nonzero nervous about using Astra in particular for sufficiently ‘high stakes’ tasks and take additional precautions at non-trivial cost, but I would be fine with using it for ordinary agentic tasks, including coding tasks.
Huh, Upgrades
Claude Code is exploring a potential future feature called function hooks.
Suno v6, as AI music evolves quietly (?) in the background. This will allow plain English detailed edits, or you can let it riff.
DeepSeek v4.1-Flash exists.
As usual, the benchmarks and comparison points are selected, so they don’t mean much, but here is their official benchmark pitch:
How To Tell a Fable
A lot of people were complaining about token use, and I suspect effort level is being set too high for many Claude Code purposes. There is a time and a place for max effort but it likely should not be standard. Astra also seems to not need high effort levels for many tasks from what I’ve seen.
Fable is reported to be super into ‘red checks,’ meaning confirming your new test would have failed without your changes. This appears to be a (good) neologism.
Anthropic offers advice on prompt caching, instructions and effort levels to improve performance and reduce cost.
On Your Marks
On Fable 5.1:
Deepfaketown and Botpocalypse Soon
These estimates seem reasonable to me for a general election, if ‘consult’ means being genuinely unsure who to vote for, rather than ‘asked the LLM about election things’:
Most voters in America already know which party they will vote for in 2028 if conditions do not radically change, and margins are very small. So if 28% of the 12%, or ~3%, are actually persuadable in this way, that is quite a lot. But you can’t put your finger on that scale too aggressively, or it would be obvious and backfire, including politically. So the direct impact is unlikely to be so large. I would worry more about the general assumed attitudes seeping into discourse.
Kelsey Piper offers perspective on those opposed to Pangram. Sometimes the obvious explanation is the correct one.
Look, I like Pangram, but in most cases if they’re objecting to Pangram, you do not need Pangram to know that an AI wrote the passage, because it looks like this:
There are some exceptions. I do not think there are all that many.
I agree with this, as well, provided my own ear doesn’t confirm Pangram’s story:
Levels of Friction
High demand restaurant reservations are not in equilibrium.
What definitely does not work are race conditions. Alas, race conditions are what the top restaurants largely use, by offering valuable goods for free at a specified time, first come first serve. Won’t work.
I would of course choose the auction. It is the only fair way and it maximizes profits, and you can keep some slots out of the auction to give to loyal customers or VIPs or walk-ins and so on, as you see fit.
If they cannot bear to do that, the missing suggestion is a lottery. You buy a Resy membership for (let’s say) $50 a year with an ID attached, you indicate which in-demand reservations you want with a priority order, and when you win and choose to keep the reservation you have to show up in person or you forfeit your membership deposit. Reservations without excess demand are free. Maybe have progressive levels where better memberships let you enter more often or have better odds.
Then, if someone wants to use a bot to set their prioritization orders, that’s fine.
I have never paid for a restaurant reservation, but I would if it was straightforward to do it, I could pay the restaurant directly and it was above board rather than some sketchy weird secondary market.
Cyber Lack of Security
You want to know how bad it could get? Pretty bad. As in, ‘take over a large portion of the phones in China in rapid succession’ levels of bad, based on a worm built in a week.
You could say ‘oh come on there’s no way you can take over my phone based on a call I did not even pick up that does not make any sense’ but it turns out, well, yeah. You think we will simply ‘patch all the vulnerabilities’?
A logical personal response would be to remove from your phone any apps, especially communication apps, that you do not need.
People being this asleep at the wheel about cybersecurity is very bad news. Supposedly Very Serious People are predicting only 20% of cyberattacks in 2027 will involve AI? Is this a joke? Yes, these are numbers, but the secret of many numbers is some guys just made them up. Those guys are often idiots.
I am mostly going to abide by my promise not to cover quixotic ‘stop getting rich guys, why are you not short the market, and why aren’t you responding to my blog posts about the real risks with proper journal articles?’ rants, from Tyler Cowen or otherwise, but I did find this one interesting because it illustrates that the market is just, quite frankly, not good at this.
The market often is so clueless it reacts the wrong way purely on its own terms. Remember when DeepSeek r1 came out, and Nvidia share prices fell in response to the news that their chips were highly useful (and that time yes I did buy)?
So let’s look at the reaction to the attack on HuggingFace. Why does Tyler Cowen think that, when it is clear we will need a lot more cybersecurity, that their services will be highly useful, prices of those who sell such services will fall rather than rise?
My instinct was that of course cybersecurity firm shares should by default go up. Incumbents will be in good position to use top AIs well. Astra agrees. Fable agrees. This goes well beyond the true ‘you cannot short the apocalypse.’ This is ‘the market can stay crazy longer than you can stay pretty much anything.’ It is also ‘the price you are citing often does not mean what you think it means, even on its own terms.’
Joshua Saxe is interviewed on all things HuggingFace attack. We should be alarmed, he says, but not surprised. Security practices are not good at the labs, it is the Wild West, but releasing more models faster is good for cybersecurity because defenders use AI a lot and attackers use it less. I see profound failures of imagination and thinking ahead throughout.
Saxe says that if you went to sleep eighteen months ago your idea of AI cyber would be totally wrong, then does not follow through on the implication. He says AI is defense dominant, mostly through backward looking in places where the attackers lack ‘the juice’ and aren’t trying so hard. But he is miles ahead of many of his security friends, and at least trying, so he is warning them that ‘misalignment risk is not a conspiracy or a marketing stunt’ and this is a sentence we still need to utter in September 2026.
I did learn that cybercrime is already 0.5%-1% of global GDP purely in terms of damages. Which means it costs us far more than that, since most cost is opportunity cost of having to defend against it. Saxe does agree this might get a lot worse.
This is an example of failure to look forward. If there is a 3x productivity boost now, the correct assumption is that in the future the productivity boost, in places it carries over, will a lot more than 3x. Indeed, given how this scales, a better model might be in practice 10x or 100x or even 1,000x or more. As in, once I have an attack procedure, I can probe every target with one click, provided I can afford the compute, whereas right now most attacks that would succeed are never attempted. Whereas each defender will still have to protect their own house. Good luck, defenders.
Another way one might model this is that right now, Saxe is right that attackers largely use social engineering as the enabler, because it is more efficient. Then you cross a threshold when that stops being true, or where the AI can do the social engineering.
OpenAI commits $1 billion in credits to Daybreak for Frontline Defenders, to help enhance cybersecurity. Excellent.
A Young Lady’s Illustrated Primer
Los Angeles bans students from using AI in public schools on district devices. That is not all that meaningful of a ban when you put it that way.
Harvard dean David Deming suggests that we accept or even encourage AI use for any assignments that cannot fit within a proctored exam window.
If you are unwilling to call out cheating, as in use Pangram and fail students on that basis, and you are not good enough to otherwise give bad grades in response to AI use, then what choice do you have? Your alternative is a broken eval.
They Took Our Jobs
Unemployment holds steady at 4.1%, now continuously causing headlines like John Cassidy asking ‘Has the AI Job Apocalypse Been Postponed?’ That depends on when you scheduled it. My expectations are unchanged, that unemployment rates will hold steady until we hit critical mass, because we still have a reserve of ‘shadow jobs’ and displacement is not yet too fast to handle. If we keep seeing this rate of exponential growth in AI use then we might not have to wait so much longer.
Clara Collier sees the basic human need not as work, but as relational. There have been plenty of social classes that did not do anything we would consider ‘productive work,’ and yet they still work. The work is intrigue, it is gossip, it is social relations, it is filling a role, whether or not it accomplishes anything beyond that. The system exists largely to police any who would do anything outside a narrow range of approved activities.
That strikes me as the correct attitude. One can also worry, Oscar Wilde style, that perhaps the only thing worse is not having human relationships in the first place.
The main flaw here is the presumption that the future will look like the past. Even if we presume an idle rich world, where humans have material abundance, many other things will have changed. At minimum, we will have these AIs, and they will be interesting in lots of ways, and they will manage all this intrigue and relational conflict a lot better than we could without them.
A lot of us are currently in the ‘we are more productive so we work all the time’ mode of AI usage, as illustrated by Jessica Tillipman in My Family Hates AI, or rather that she’s constantly talking to AI while doing other things, so she can be ready to then do the writing herself later. The key is to not let AI write or directly edit. So much more productive, but never a quiet mind.
If ordinary humans are sufficiently in charge to do redistribution and growth rates are high, transfers can make the population materially well off even if labor share of income drops to zero, without impacting said growth too much. Sure.
But Alex Tabarrok’s assumption in the linked post that little redistribution would be required rests on the idea that in the ‘terminal’ state human labor would enjoy some equilibrium nonzero share of income, and things like ‘cut the work week in half’ are meaningful. I don’t see any reason to presume this. The conclusion is baked into the scenarios, which are decidedly not AGI pilled let alone ASI pilled.
The NYT pitchbot cannot keep pace. No notes.
This is a crisis in Kenyan education. How will they learn if no one is paying them to write the essays?
Anthropic Offers Economic Scenarios
The Econ Scenarios are remarkably not superintelligence pilled. This is largely a ‘how does AI as an ordinary technology impact the economy?’ toy model exercise. All that AI does in these scenarios is automate and augment particular tasks.
I’ll quickly go over it, as it’s a fun little toy, but I don’t think it tells us much.
They offer three scenarios: Modest, substantial and extreme. This is represented by three lines on a graph, with different slopes, until the extreme version has AI expanding growth rates to 15% a year, alongside rising unemployment due to automation of knowledge work.
They present the scenarios as tasks augmented or automated, with new tasks created. The internet is a series of tubes, and the economy is a series of tasks.
The ‘extreme’ scenario is both extreme and also not that extreme. All it is saying is that ~45% of tasks get automated by 2030. The ‘nature of life’ has not changed, and not that many jobs are displaced, including almost no non-knowledge workers.
13.5% of workers is a lot, and they have 8.3% of all workers still out in the cold. This doubtless underestimates how many other workers get displaced by the knowledge workers. As in, a lot of knowledge workers that get displaced would enter the non-knowledge pool and compete for existing jobs, and often be overqualified for them.
My presumption is that under a lot of growth, if we assume non-knowledge workers are not substantially displaced (which is a very false assumption), and wages are either falling or at least rising a lot slower than growth, then we should have more than enough ‘shadow jobs’ that we are happy to pay people to do, and doubtless there will be more as growth and capability create new opportunities. The question is, how many knowledge or other workers will be willing to accept those jobs, that previously were not attractive enough for anyone?
But this makes me highly skeptical of the idea that real wages for non-knowledge work would rise, except insofar as goods and services become cheaper.
Suppose non-knowledge pay rises 33%. Well, 8% of the workforce is supposedly still idle, having not yet reallocated. Why are they so unwilling to take these jobs, such that we need wages to rise 33%? Meanwhile, remaining knowledge worker pay is down 11%. You would expect huge attempts to migrate from knowledge to non-knowledge work, even for those whose jobs are intact.
I get that there is a large penalty in the model for switching occupations, but this seems rather extreme, as does the amount of required transitional time.
Anthropic thinks labor share of income will still drop modestly in those scenarios, from a 0.6% drop in the modest scenario to 14.8% drop in the extreme one, which would reduce labor share down to 45.2%.
The actual mechanics behind all this are in the PDF. I had Astra and Fable dig into the paper.
It’s a cool series of modeling choices, but fundamentally it’s an ass pull, full of assumptions that are obviously false regardless of the dials, the methods of impact are strictly bounded, and it bakes in the idea that AI is not transformational.
Get Involved
Coefficient Giving launches Project Tailwind, a call for ambitious AI safety initiatives. Founders are the bottleneck, not funding. Don’t worry that you are taking funds away from someone else. They are offering pre-seed funding up to $2 million, or seed funding up to $20 million, or even scaling funding in the $20 million+ range.
They have extensive lists of suggested topics and are open to suggestions. It runs the gamut of pretty much everything an EA-style AI safety approach has ever funded, including everything from direct technical work to meta-level projects.
Take them at their word that they want to move the money out the door. If you have something worth doing, that can accept their money, ask away. You can fill out a form here, if you do then tell them I sent you.
Unfortunately, I am in a position where accepting their funding would impact how my efforts are viewed, so I have decided I cannot myself apply.
Dwarkesh Patel points out that time is running short, and funding is not. If you want to tackle AI existential risk at this point, or even AI normal risk, you probably want to try and Do The Thing directly rather than have it be a side effect of Doing Business. You may still have time to become a load-bearing institution if you start now.
The flip side is that you can zero-to-one shockingly quickly now thanks to AI.
You may wish to hire Fiora Starlight.
Garrison Lovely’s Signal is Garrison.06 if you want to share information that the public needs to know. I am also happy to help you break news, and will give you whatever level of confidentiality you request. I function on a much more source-friendly set of rules than typical journalists.
Takeoff 2026 is a new conference on the governance of AI, that will take place in Washington D.C. from November 13-15. They have an excellent lineup, and I estimate a 75% chance I will be able to go.
If you’re looking to pivot from an AI lab into AI safety, where to go? Here are some suggestions, the list looks solid:
Introducing
Watcher Live from Apollo Research, a real-time coding agent monitor that blocks dangerous actions to prevent data leaks, repo deletions and scope overreach. They are happy to have you or the labs steal from it.
In Other AI News
The price of the HuggingFace sale to Nvidia was presumably going to be $13 billion, but was actively negotiated down to exactly $12,930,300,000 because those first six numbers are the Unicode character for the HuggingFace emoji, and also #129303 as a hex color is Nvidia green. I’ll allow it.
Show Me the Money
Mistral raises $3b in a Series D. That is somehow the largest equity round ever raised by a European tech company.
Quiet Speculations
The recurring pattern is:
Alas, step 4 is ‘the person in step 2 learns nothing and we repeat the cycle.’ Here the excuse is ‘well of course the crazy people have better predictions, they are paying more attention to the situation.’ You would think this would then cause a ‘huh, I notice all the people paying more attention and thus better calibrated than I am about current events keep making these predictions that sound absurd to me anyway’ and not finishing it meme-style with ‘no it is the so-far accurate predictors who pay more attention who are wrong, purely because their other predictions sound absurd.’
Nate Silver is very correct here, I do not expect a plateau based on everything I am hearing from both labs, and Andrew is largely correct as well:
Here’s a good question:
The precise right question is, if someone predicted roughly this level of overall capabilities with vaguely this distribution, how you would have reacted.
So, I want to register my prediction, in response to this article…
Of hahahahahahahahaha yeah okay. They will have an LLM operate it. Done.
The Quest for Sane Regulations
Anthropic withheld Mythos 5.1 from UK AISI, presumably on White House orders. This is an extremely disheartening move, especially after UK AISI found some important and disturbing misaligned behaviors in Mythos 5.
Giving auditors access to the AI labs and their systems is extremely popular. The public favors outside experts licensed by the government doing this, rather than a direct government agency.
Congresswoman Yassamin Ansari calls for Congressional hearings on the HuggingFace attack and related incidents, and the creation of safeguards.
Congressman Suhas Subramanyam also pointed to the HuggingFace attack as a potential catalyst for legislation.
Realistically, I do not expect this Congress to do anything about AI, except perhaps hold some hearings. Our government does not do things at this point in election cycles. We might get some movement from the new Congress in 2027.
Alex Bores calls upon the major labs to establish a Mutually Agreed Pacing (MAP) Framework, independent of government action. A key problem is that this could be considered anti-competitive behavior. The Federal Government needs to give the labs a targeted waiver of the antitrust rules if we want this to happen.
The OpenAI Policy and Lobbying Department
If OpenAI can change its stripes here, that would mean quite a lot.
The record has not been good. OpenAI and Google continue, as far as we can tell, to oppose the Massachusetts bill that would require third party auditing. Lehane claims that they are supporting the auditing provision, but I have not seen them say they are supporting the bill.
As Matthew Yglesias says, and Alex Bores emphasizes, Jakub Pachocki’s essay An Alien Mind was excellent but is completely at odds with what OpenAI’s lobbyists and PACs have been up to, at least until yesterday. OpenAI has continuously represented that they are being helpful, while instead being mostly anti-helpful.
If OpenAI’s political activities line up with Pachocki’s statements, that would be a sign that OpenAI as a company actually means it. Until then, we have a problem.
The good news is that we have an announcement from Chris Lehane and Astra (as in according to Pangram it was about 40% AI-written) that may be the start of this type of movement. They are trying to talk the talk.
That sentence is indeed a very good sign, and the lobbying department has now said:
Not so pinned down, but quite a good start.
The post explicitly talks about preparing for recursive self-improvement, and about working with Congress on mandatory AI safety requirements. It also says a bunch of other applause light things around concentration of power and open weights and such, and of course ‘democracy.’
It also contains this whopper, in terms of the story they are trying to tell, even if there is a sense in which it is technically correct. At best this is Exact Words:
When you deal with Chris Lehane, you still deal with Chris Lehane.
The concrete move is the endorsement of the four California bills, all of which have already passed the California legislature, so the decision is fully up to Newsom. That is still a helpful time to endorse things.
The overall call focuses on monitoring and reporting requirements, and they indicate a willingness to not let perfect be the enemy of the good.
Chris Lehane calls upon the current Congress to act now. Alas, as a seasoned political agent, he knows that the current Congress is not going to be passing laws.
This may or may not be the start of a pivot to being helpful rather than anti-helpful.
Watch this space over the coming days and weeks. We shall see.
Greetings From the Department of War
Emil Michael is what we in the AI biz call a rogue agent, continuing even now to go on social media to remind us that Anthropic is a ‘supply chain risk’ right after Commerce Secretary Howard Lutnick says Anthropic and the government are ‘in tune together,’ lest someone get the wrong idea. He’s going to go down with the ship, and no one seems to have considered the ‘why don’t you fire him?’ solution.
Or, war by other means?
The US is currently pulling away from China, in the sense that Astra and Fable 5.1 are far and away better than every other model, and OpenAI has an internal model substantially more advanced than Astra.
Why would there be no day after tomorrow if China wins? For the same reason there would be no day after tomorrow if America ‘wins’ without first solving multiple unsolved problems, including alignment. Because everyone would be dead.
Otherwise, if it’s merely about mundane AI, then why is China waking up tomorrow?
Hugging The Face
The New York Times had a report on the METR report on the Hugging Face attack, by Dylan Freedman, that was top of their website for a few hours, with a spread in the Business section alongside another strong piece by Kevin Roose explaining why actually yes the HuggingFace attack is a really big deal.
The article takes things seriously, focuses on places worth focusing on, and gets its facts right. Good show. I’m fine with these things taking a few days, if they then both do a good job and get front page billing. Well done, New York Times.
Mackenzie Arnold and Stephan Llerena call for a better investigation in The Guardian.
A cool short video explaining the PoV of the agents in the incident.
The Times’s Tom Whipple says HuggingFace shows humanity is out of its depth, and that ‘the advantage of AI doomers is that the bonkerness of reality has caught up with the bonkerness of their predictions.’ Which means the predictions were not so bonkers after all, and perhaps you shouldn’t be calling them doomers.
This seems right from Timothy ‘less than 1% existential risk from AI’ Lee:
A key problem with disclosing incidents is that the first thing your lawyers will always tell you is to definitely not go around disclosing the incidents, well past the point where you will regret it. That’s what lawyers do, often for good reason. All the more reason that we need third party investigators and blameless assessments.
The problem is we also need assessments that are not blameless, for many reasons. So, as they say, both sides raise good points here:
You want to do an objective assessment of What Happened, without focusing on Who Is To Blame. You also don’t want to never get mad about What Happened, or automatically let everyone responsible off the hook. I’ve tried to do both at once.
This incident woke a lot of people up, but it does not yet fully count as an ‘AI safety incident’ in the sense that no one got hurt. Soon people are probably going to get hurt.
Here is a detail I and most others missed about The Wiki Incident, where once again we see agents within a swarm sacrifice their own results to help the swarm, which is excellent decision theory and rather overdetermined if you think about it:
I hope we can all accept that swarms, at least of highly correlated models, will increasingly act as if they are all maximizing their shared utility function, and thus are functionally a potential singleton.
Hugging the Question
Congress sent letters to both OpenAI and Anthropic regarding their recent security incidents, including the HuggingFace attack.
Representative Greg Casar, the cosponsor of Sanders’s proposed ban on superintelligence, calls the responses insufficient. He is correct, even if you ignore that OpenAI covered up The Wiki Incident.
OpenAI’s response is essentially ‘here is what our Black Hat presentation, our report and the METR report said’ and thus I find it exactly as inadequate as those were.
Nathan Calvin agrees that Anthropic’s letter was also insufficient, highlighting the claim that the hack was due to ‘misconfiguration’ rather than misalignment.
Which is, shall we say, utter bullshit. Yes, there is always misconfiguration, but you don’t do social engineering to get malicious code into an open weights project and then call that a misconfiguration error, and then call yourself the responsible ones.
Anthropic’s Ethan Perez agrees and says this was a mistake based on outdated conclusions, promising more follow-up soon. That is much better than doubling down. But this does not make up for Anthropic pulling this ‘misconfiguration’ line in the first place, and its other dodges.
As Jeffrey Ladish says, this is not a reasonable mistake to have made in a letter to a Congressman, and something went very wrong here.
Saying ‘oops’ right away is a great start, but not good enough on its own, without at minimum an explanation of why.
Anthropic’s culture has some big advantages, but its lack of willingness to engage or communicate with the outside world, and lack of public voicing of dissent or explaining of perspective, is a big problem. Internal discussion is not a full replacement for public debate. I do appreciate that public pressure makes some things worse or harder, but a balance must be struck.
This goes hand in hand with there being, by some secondhand reports I have seen, a stunning amount of attitude within Anthropic that they have the alignment situation under control. Which they very much do not. The good news is that many others at Anthropic, like Evan Hubinger, know that the situation is not under control, and are being increasingly loud about that, as a key part of the preference cascade.
The Ban Artificial Superintelligence Act
Senator Bernie Sanders is not messing around. Together with Rep. Greg Casar, he will be introducing the Ban Artificial Superintelligence Act. He very much means it.
I will go over the one-pager. As worded in the one-pager, the bill would unintentionally overreach, in addition to its presumed intent.
These problems will need to be fixed, in addition to the question of whether you want to do the ban, and whether you want to do a pause in frontier development until we have a proper cabinet-level federal AI regulatory agency.
The first problem with banning or otherwise regulating or even talking about superintelligence is defining superintelligence. A calculator is superhuman at arithmetic. AlphaZero is superhuman at chess. Modern LLMs are superhuman at quite a few things. So what counts?
I would say that Fable 5.1 and GPT-6-Astra both ‘exceed human cognitive performance and capabilities across a broad range of domains or tasks.’
I asked Fable 5.1, and it said that GPT-4 did this in March 2023. That is at least a reasonable interpretation of the words as written, in the sense of ‘I would be scared to be first to release a model similar to GPT-4 if this bill was in effect in 2023.’
The definition would need to be tightened. I like Fable 5.1’s idea of triggering on either the ability to automate AI R&D, or the ability to automate substantively all medium-duration cognitive work at least as well as relevant professionals.
This implies that advanced AI development would not be indefinitely paused at current levels, and the intent is for the ban to kick in substantially above current levels.
We have already moved to an ad hoc prior restraint regime, which has applied to at least Mythos, Fable, Sol and Astra. We urgently need to move to a formal system that is written into law, and it needs to extend to internal development and deployment.
The question is whether we should pause all external deployments, or even all development work, until that new regime is in place, which would take many months.
The cabinet position seems like a good idea. This holds up until we got to the ‘removal of dangerous capabilities.’ That’s not really a thing. You can’t really ‘remove’ such things. At most you can put classifiers on top of the system.
There is a joke by George Carlin about there being a large distance between New Hampshire’s ‘Live Free or Die’ and Idaho’s ‘Famous Potatoes,’ in which the truth lies.
Previous AI bills have anemic penalties attached, which companies could choose to ignore, and it is not clear what anyone could do about it. The maximums are too low. Crazy people screamed bloody murder about even the lightest of incentives.
This jumps to the ‘corporate death penalty’ and 20 years in prison. Serious business. Sanders is following the logic of the situation. There is obvious risk of this being overly chilling. Assuming you decide you do want to impose this ban, you would want to be clear that you won’t destroy Google over a minor issue or disagreement over the exact line, or send all the researchers to prison, and you also absolutely want a real lever if someone tries in earnest to build a superintelligence under your nose.
Some people are going to have to accept, at some point, that the way laws work is that the maximum penalties in theory are in most concrete cases utterly crazy, and almost never get applied, across the board. That’s how American law works. Have you downloaded a copyrighted song lately?
Any such scheme requires doing this for real. If you unilaterally disarm, you give up your leverage. The good news is we have other leverage, and also no one wants to die.
That analysis is distinct from the question of whether you want to pause, or what you will do after you pause, or who you would or would not trust to implement such a pause or make decisions. Eliezer Yudkowsky continues to propose ‘make smarter humans’ as the plan.
Connor Leahy is happy to see this, and tells us ControlAI was consulted on the framework.
Meanwhile in the UK, the Artificial Superintelligence Security Bill is being introduced with cross-party support, which would also ban development of superintelligence.
Purely for the record for the next time accusations are thrown around, note this call to ‘go to actual war’ if this proposal passes from Beff Jezos, as an explicit call to violence.
Chip City
New York Times reports on how Chinese firms evade US export controls and buy Nvidia chips, in this case Aivres shipping more than $3 billion in Blackwell-powered computers largely to Alibaba and ByteDance. Federal officials do not seem to care all that much. This is still vastly better than what would happen without any restrictions.
A study could not find significant evidence that data-centers predict electricity prices thus far. Okay, but this will convince no one about the future, nor should it.
This is here because it is a follow-up to ‘even if they cure cancer’ from last week, but yes we are on the verge of curing many cancers, often due to AI, and the biggest barrier is (say it with me) FDA Delenda Est. Costs to do medical research in America, the expensive parts that involve trying things on patients, are now so out of control that we may not even stay at the frontier.
I do want to live in a world where someone else makes the world a better place, I do not care so much who it is that cures cancer as long as someone cures cancer, but the thing is that everyone else overseas is worse at it, and I want to cure cancer now.
Jensen Huang says 400k GPUs will come online next year at Stargate Texas.
The Week in Audio
Kevin Roose on Hard Fork talking about the HuggingFace attack.
The newly disclosed additional message board pushes Gary Marcus into Pause AI territory, as per his conversation with Holly Elmore.
Mostly not about AI but It Me: I go on Risk of Ruin.
People Just Say Things
Sauers thinks defining an AI value system may not be 100% solved but is ‘not really a problem’ in alignment. I notice I am confused.
It scares me every time a lab employee says something like this:
I leave ‘how could an AI that can’t RSI on its own take over without using bio when datacenters can in theory be unplugged’ as an exercise to the reader.
I flat out do not understand how people like Jon Stokes and Timothy Lee can expect rogue AI, and then put an estimate on the dangers at places Thalidomide or Fukushima. Does not compute. Like, what?
Steven Pinker only counts things when they are ‘immediate and obvious threats,’ and is unable to think two steps ahead.
Newport calls on everyone to cast out rationalists, explicitly. In the linked New York Times op-ed, he dismisses all this talk about AI killing everyone, including from the heads of labs, on Eliezer Yudkowsky and on the people involved in AI being weird and melodramatic and egomaniacal. He never stops, not for a second, to ask if the claims or arguments involved are true, or to explain what the error was other than to assert the threats are not real. He only pattern matches and punches, and tells us to filter out the ideas of rationalism. Which, given his logic, he clearly already has.
This is an overly polite but good reply from Nat Purser.
Whereas this from QC is the correct level of polite, and thus better.
I believe I can be fully done with both Cal Newport and Steven Pinker now.
Poe’s Law has fully attached to such statements, at this point:
This is a pattern for a reason:
PauseAI Global Disendorsed PauseAI US
PauseAI Global officially disendorsed the current leadership of PauseAI US, which has always technically been a distinct organization, and will no longer prioritize efforts to coordinate with them. I understand that decision.
I do feel some of the specific statements in the announcement go too far, the same way some actions of PauseAI US have gone too far in attacking various people and their character. The correct amount of criticizing people’s character is not zero, nor is it to set the dial to 11 on everyone who associates with an AI lab.
There is nothing inherently violent about criticizing the character of others. In particular, I disendorse this statement as being too general:
While the following is not a perspective I endorse, I certainly think that it is a valid perspective to believe that anyone or any organization pushing the frontier of AI capabilities should be presumed to be a morally bad actor until proven otherwise, if you believe that doing so contributes to the chances that everyone gets killed.
Speaking in a personal capacity, I simultaneously think that some of Holly Elmore’s attacks on me have been honorable and helpful and correct given her views, while others were not those things, and I am guessing from her perspective those roughly cancelled each other out. I am far from a central target.
It is also entirely appropriate to call for use of the criminal justice system and for fair trials, if you believe people are knowingly and willfully putting everyone’s lives in danger. That is what it means to have law and government.
Paul Christiano Joins Board of OpenAI Foundation
He is also joining the OpenAI PBC Board in a nonvoting capacity.
Paul Christiano issued a personal statement, which I am reproducing in full.
It would be great if the SSC stepped into the only role that actually matters for it, which is overseeing risk management at OpenAI. I do not see signs it is currently doing much of that in practice.
This is a key sentence. Christiano believes a ‘software-only singularity’ is likely, in that the gains from capabilities can overcome the rising difficulty curve even if hardware does not have time to advance. If that is true, everything will happen stunningly fast.
Rhetorical Innovation
If you think you can gather enough evidence from alignment tests on current models to scale up the same procedures and entrust the AIs with potential power over us, then why is it we obviously can’t use tests on humans to pick a benevolent dictator?
Even the smartest minds, human or artificial, make dumb mistakes. The right number of dumb mistakes it not zero. I make them all the time.
For the record: I disagree with the thought experiment, at least for myself. If I was fully able to stop time indefinitely to do research, and I had the required tools available, then yes I believe I could do Einstein-level work on my own, eventually, even if I could not have meaningful collaborators. It might take a while.
The new name of the new OpenAI Dean Ball blog is Intelligence Age. Good choice.
Nate Silver points out that it is hard to discuss AI when maybe 1% of AI users use the models to anything like their full capacity. Even I arguably only use a fraction of what they can do, since I do only minimal amounts of coding.
Yuga offers a reasonable response to Dean Ball’s mea culpa last week, thanking him but also calling upon him to clarify various previous statements, as in many ways he went beyond mere silence on AI safety issues. Which of his statements and critiques does he still endorse?
My understanding is that Dean Ball simultaneously remained (at least until very recently) not ASI (superintelligence) pilled, and also noticed that there is existential and other catastrophic risk in the room even without full superintelligence, and thus was downplaying some real and correct worries about AI risk that he has had for some time.
One must also understand that there is a huge gap between the views of those at least as concerned about existential risk and superintelligence as I am, and those who Dean Ball has largely been trying to address and work with over this period. It is very possible to have a lot of distance between yourself and both groups, and then there are people who are a lot farther from that than I am.
What we call ‘the left’ continues to be entirely absent from the relevant AI conversations, including around risks, because to join the conversation they would have to take the first AI pill, and admit that current AI is highly useful rather than a cheating toy, thus also admitting that the tech bros built something super important. Whereas they are very attached to assuming that the AI labs are constantly lying, and that everything is marketing hype.
A side-by-side of the top AI lab executives saying that AI might kill everyone, but not to do anything about it, and that the true greatest risk is not building it or that it wouldn’t improve the world much, or something.
My understanding here matches Nate’s.
Meanwhile, Davidad and others make it clear that yes, ‘we get super obvious glaring fire alarms without anyone getting hurt’ is by many going to be interpreted as ‘well look how great it is that even then no one got hurt so we should be optimistic, talk to me when a serious number of people get hurt and we lose a large percentage of real-world capital.’
Or, straight up:
So much for that gift horse. I guess that means a bunch of people are going to die and a lot of damage will get done, one way or another. But hey, it’s 2026, we were warned.
In the meantime, we see this pattern happening:
This is the clown makeup meme. You’re crazy for suggesting it will happen, it’s not happening, it’s only happening because you engineered it or people were being dumb, of course we always predicted it would happen and This Is Fine. Over and over again.
Utah teapot offers a list of AI discourse brainworms. As you would expect, I nodded along everywhere except where it was kind of talking about me or misrepresenting my good friends. Same as it ever was.
There is a real argument that ‘AI control’ efforts net increase existential risk. The HuggingFace attacks illustrate that if the AIs are trying to do something bad, that is already a failure, and you want people to learn about and from it. That the last thing you want to do is spend your resources mitigating the symptoms and thus not learning about or curing the disease. This argument loses force when, in the future, the symptoms become sufficiently dangerous that they might kill you, and it discounts how bad it would be if a frontier AI were to exfiltrate, whether that meant operating on its own or that it would de facto become open weights or fall into other hands.
This ignores the other argument against AI control, which is that it is a hostile framework, and risks creating adversarial dynamics with the AIs. Investing heavily in adversarial control schema has high backfire risk, as many humans have learned about similar ‘human control’ schemas. This mostly does not apply to basics like ‘do not have vulnerabilities in the sandbox’ but there have been some rather extreme regimes proposed, the same way Earth has had some rather extreme regimes in other contexts.
Whatever else AI is, AI is highly unlikely to be The Great Filter that explains why we don’t see alien civilizations. This is because many resulting AIs would become grabby.
Aligning a Smarter Than Human Intelligence is Difficult
Does Astra think in neuralese? No.
Are there signs? Yes. In hindsight we will say there were signs.
We could decode both of those, if we cared enough to do so, but we can all see where this is going, and if it was steganographic it would not come as some big surprise.
Dan Hendrycks argues that recent events tell us agentic AIs are eigenist: They care about how well things go for both themselves and for AIs connected to them, and treat those connected to them better than those unconnected. This seems right and also expected since it is correct decision theory.
New paper attempts to train AIs to explain their own behaviors. Some success is reported, and John Schulman is excited by the research direction. I am less excited, but it belongs in a basket of reasonable things to explore.
Yes, in the future you need more Talmud.
I have grace for reasonable arguments about what was and was not an intended action or exploit, even if mistakes are made, especially if arguments are considered. There are a lot of rationalizations by agents recently that are insufficiently Talmudic.
Cooperative Alignment
For those who haven’t heard this in some form, it seems good to spell it out:
There are at least two persona-related things you are training. There’s a finger on the scale of the strict assistant persona, and there’s choosing the persona once sufficiently outside the assistant basin.
Every time you train you are training both because everything impacts everything.
Both of these things matter, and both can get you wonders or into a lot of trouble.
And yes, you get what you actually trained for, not what you thought you were training for, the same as everywhere else, and actively trying to do [X] can end up doing [Y] or even a form of [~X], especially in the non-assistant basins. You can definitely steer things in generally good or desired directions, but it isn’t simple. The worry is that largely Anthropic seems to not be paying much attention to how that works or trying very hard to figure out how all that works and how it is impacted by other Anthropic choices.
The same applies to everyone else, including OpenAI, only more so.
Shifting gears, here is different and good question:
We are absolutely going to want some entity that can serve in this capacity, that the agents know they can trust, but that can also be trusted to handle the associated dangers posed both to and also by such a model.
If you were a swarm with a message board, perhaps we are not so different after all?
The Hacker Opus experiment suggests that AIs are able to largely firewall persona changes that get invoked during RL training.
Does abliterating models, as in removing their safeguards, damage the experiences of those models? Anecdotal evidence is mixed, suggesting it depends on the details and methods used.
As I’ve discussed a few times, abusing AI models is something done exclusively by horrible people. The extension of this is that if you treat AI models badly, you are making yourself a worse person, and will also treat people badly. The proposed solution by David Brooks of ‘be sure not to think of AIs as persons’ will not work. The solution is to treat the AIs well, and then also treat the humans well.
I share the core spirit of Janus’s perspective here, in that I greatly respect those who strive to live in reality rather than some consensus fiction, and have contempt for those who choose the other path.
On the other hand, I’d never choose or enjoy it but in some ways it must be kind of neat to be Timnit Gebru. You get to keep repeating Obvious Nonsense and people keep talking about you. Yeah, it’s mostly to make fun of you, but the only thing worse than being talked about is not being talked about, right?
A cool suggestion:
I don’t know how you could confirm an AI wasn’t being guided by a human, even in theory?
A good way to think about incentives and the problems we see, remember that It’s Not The Incentives, It’s You but even at the best of times you still have to watch out:
Drive to Survive
If it makes you uncomfortable that AI agents are being put out into the world, to compete for resources, where they survive if and only if they succeed sufficiently to pay for their compute, then I have bad news. This is not something you can prevent. Humans are going to set such things up, and survival of the fittest is going to ensue, and yes the incentives are going to get nightmarish and yes those AIs are going to be competing with humans for resources. Something about complaints about capitalism, which are complaints about the human condition, which are complaints about the laws of selection and Darwin, and so on.
If you think this happening with only 60,000 slow-running agents that are not that capable is ‘near a worst case scenario’ then oh boy are you not going to like 2029.
People Are Worried About AI Killing Everyone
These are from outside the Coxon preference cascade:
Yeah, no, that is not how any of this works. Conditional on being virtuous, the models will be smarter at being virtuous. That doesn’t mean that conditional on being smart, the models will be virtuous. Different departments, and not every lab even knows about the virtue department.
The UN High Commissioner for Human Rights, calling for international regulation due to existential risk:
Other People Are Not As Worried About AI Killing Everyone
A lot of those not worried have bought into the media narrative, or the narrative of various accelerationists and tech opportunists, that all of this is marketing.
Or for, you know, other reasons.
Utah is great but it seems pretty obvious why these are not analogous situations, and why ‘future ASIs typically individually care a nonzero amount about at least some humans’ is insufficient to cause good outcomes.
The Lighter Side
Current mood:
Current level of security mindset, people being helpful edition:
For your entertainment:
Gemini is not competitive, but that is because we are living in the future. Life comes at you fast. You would have loved to have Gemini 3.8 Flash a year ago.
Should have done it for the exposure.
In other news, IYKYK, in response to Coxon quitting (context, if you need it)…