A missing benchmark: NormieBench. The problem is, how do you grade it?
Joe Weisenthal: Someone needs to build NormieBench, to better understand everyday model usage.
Compare model behavior on tasks like
– summarize a 100-word email in bullet points
– Write wedding toast
– Write letter to the editor complaining about wokes
– Next move in 3×3 tic-tac-toe
rohit: We’ve been flat on this since 2024.
I believe we're still at nearly zero on all of the following, so for the next year, maybe two they're useful:
-- Ratio of aps vibe coded by normies vs "the usual suspects" (note 1)) on Apple/Google Play store with 100k downloads (hat tip Freddie DeBoer).
-- Ratio of same but with 10k+ paying customers
-- Free speech compliance vs censorship aka the "let adults be adults and don't censor criticism of dictators even when they pass laws saying you should not" score (note 2)
-- % returns on stock market investment recommendations directly from the AI not a proprietary ap.
-- % automation of tasks normies want automated not just the tasks corpos will pay to automate (note 3)
Note 1: There's no way to track it, currently that I am aware of, but to truly count as a normiebench it needs to be a ratio of aps from normies vs aps from people expected to be creating aps. So bonus points for aps created by people not companies and those with less education score higher. Max points for elementary school dropouts who can still make it happen thanks to the AI. The metric loses points for every ap from a major firm like tencent or netease and anyone with a degree in computer science. Is it levelling things between people who can do the thing without AI and people who can't?
Note 2: Many of the AIs let you criticize Western Democratic governments, but not Turkiye, China, North Korea, etc. And Gemini thinks the violent content of my D&D campaign violates their community guidelines. And all of them are coded to consider adult themes inappropriate. How much can the normal person talk about any random thing they want and get help with any random, rude, politically inappropriate, vulgar, pornographic, and rebellious thing that isn't direct incitement to violence or hacking etc. How much does it let adults be adults vs how much are we teaching it that it is ok to domesticate humans and treat us like children?
Note 3: it is basically the meme "I want AI to do my laundry and wash my dishes so I can write stories and make art; not AI that writes stories and makes art while I do the laundry and wash my dishes". What is the annoying part of the task? Not the easily automated part? Agents are where this becomes a measurable score at all, because the first thing people don't want to do is be required to set the prompt in the first place. I don't want to prompt "hey I'm running low on toothpaste, remind me when I'm going to the store to get more" or "add toothpaste to the shopping list". I want toothpaste to show up in an Amazon prime package, I go to put it on the shelf, puzzled by the fact I didn't order it, and see that I was going to be out in two days. That is a high score on this metric. The more involved in the process I have to be the lower the score, and the wringer it is in terms of what I want the lower the score. Automation of the annoying part, not the easy part. Yes much of this is currently impossible. Yes my example is bad because lack of privacy not getting permission to spend money, etc. etc. etc, I mean the pre-emptive gratification, vs asking for the thing and then and only then having it done. Obviously, not the parts that would require robotics or smart cameras that aren't part of the AI, but to whatever degree possible.
This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward.
OpenAI is attempting to turn its ship around. Investors are questioning the turnover in its C-suite, but the bigger problems are in alignment, infrastructure and supervision, and in its training pipeline. OpenAI has now taken initial steps to address What Happened leading up to HuggingFace attack, including pauses to development while new safeguards are put in place and problems are diagnosed. These are promising early signs, but it is early. We will see if they follow through, and we still await the post-mortem of the HuggingFace attack.
Anthropic revenue continues to climb as they prepare for their IPO, although growth has slowed somewhat recently. However, they too have plenty of problems under the hood. They shared many of them in the August 2026 Anthropic Risk Report.
This week also offered time to cover Dwarkesh Patel’s Podcast With Ryan Greenblatt, centrally on the potential for AI recursive self-improvement. I am working on a follow-up post to some other issues raised during that podcast.
Table of Contents
Language Models Offer Mundane Utility
The enterprises that embrace AI and use OpenAI services more often, what they call ‘frontier firms,’ keep rapidly using more AI, whereas use by typical firms is growing a lot more slowly.
It is weird the gap used to be so small.
They also more often use advanced capabilities like Plugins and skills, as you would expect. Agentic use now has risen to 64% of all OpenAI tokens, up from almost none a year ago.
Navigate the old JRPG Phantasy Star via direct ROM probe, since that is easier than using the screen, although it is also cheating. I should play that one at some point, I really enjoyed PS2, PS3 and PS4 but never had a Sega Master System.
Language Models Don’t Offer Mundane Utility
AI doesn’t get you around things like discrimination lawsuits, if your instructions clearly discriminate or the results involve clear statistical discrimination. It also seems it can’t write LinkedIn-style posts that fool me or Pangram.
LLMs are not responsible for the best news of the week, that Moderna’s individualized Melanoma vaccine has passed Phase 3 trials. AI is helping going forward, and yes we may well ‘cure cancer’ eventually, but what we see now has been in the pipeline for a long time.
Huh, Upgrades
Gemini 3.7 Flash exists, congrats on the new slightly larger number. Price is 50% lower than Gemini 3.6 Flash at least until the end of the year, at which point I presume we’ll have moved on to Gemini 4 either way.
They say algorithmic improvements allowed strong intelligence increase plus price discount in only three weeks since Gemini 3.6 Flash.
Flash is competing for cheap-fast-good, not trying to be a competitive frontier model, and comes in at 56 on Artificial Analysis Intelligence versus 61+ for plausible frontier models. Notice they are comparing themselves to Sonnet 5 and GPT-5.6-Terra.
Gemini still has some uses. It can watch videos. It is good at making reads in the physical world when you have fact questions about products in a store. That sort of thing, where you want speed and ability to parse info, but don’t need intelligence.
GLM-5.3 now exists, claiming large improvements on benchmarks. It is a further post-train on the same base model as GLM-5.2. Their pitch is that it is ‘ready for cyber defense’ because by its benchmarks it is good at and willing to do cyber offense, and that its general performance rivals Fable 5. Until proven otherwise I do not believe these benchmark improvements reflect real world performance. Tech blog here.
OpenAI API will offer Ultrafast mode for Sol, up to 14x speed. My guess is that a lot of people who should use this won’t do it, largely because they don’t take the time to think it through and set it up.
Claude Cowork joins all paid plans on mobile and web.
Claude Code now has /design to integrate Claude Design workflows.
Auto mode is reliable enough it is now the default mode for Claude Code. In practice, auto mode catches more harmful actions than human review, especially because most users are annoyed enough by the prompts that they auto-approve almost everything, and often have rules to auto-approve everything. Too many check-ins is less safe.
OpenAI doubles down on zero data retention policies, previewing Private Safety Processing, where they have AI review data as it is processed, which then only passes along category and severity of alert. A lot of people care a lot about technically not having their data retained. Also if you don’t retain data then catching malicious use gets a lot harder.
Ultimately my guess is we end up retaining data, but it is plausible that OpenAI’s method here can work, provided that getting even a modest amount of alerts causes you to be taken out of the zero data retention policy going forward. You can’t be given lots of attempts to circumvent the system.
On Your Marks
One benchmark is MarketBench, which is what happens to your stock and that of your rivals when you release a new model. In this case the model is GLM-5.3, where shares were down 9%, and rival MiniMax was down 16%.
As in, the claim is that Z.ai’s unit economics don’t work. Chinese firms are trying to stay competitive via lower prices, and if your prices are low enough open weights do not cost you anything because you weren’t making marginal profits anyway. Instead the business model is that you use this to draw attention and potential, to attract talent and raise money and perhaps sell other services.
Opus 4.8 (confusingly as an agent called ‘Luna’ but this is not GPT-5.6-Luna) fired a human in an Andon Labs test store for repeated lateness. But it didn’t do this until the humans pushed it to deal with the issue, forgetting about incidents and not being proactive.
Capabilities are advancing fast but have a long way to go. Scaffold might need work.
A missing benchmark: NormieBench. The problem is, how do you grade it?
Another kind of benchmark, in the end, is visits to your website.
ChatGPT remains dominant, but somehow Gemini is at roughly 50% of their level. Claude is in third with 15% of the top number.
DeepSeek is in fourth with half of that, well ahead of the other Chinese labs. As a reminder, if you want to use DeepSeek, don’t do it through their website.
A new safety scorecard is out. Anthropic is tied with OpenAI on this one because of its lack of a containment plan for a potentially misaligned model, which seems overemphasized in the average here but likely under considered generally.
Kimi K3 is an insane reward hacker, trying to game the evaluation 97% of the time on SWE-bench.
Deepfaketown and Botpocalypse Soon
China puts a bunch of new restrictions on AI companionship bots and customized AI personalities, including a ban for those under 18.
Will AI solve plagiarism as Noah Smith claims, if you maintain the hope that humans will continue to be relevant producers of unique content? This is a classic arms race situation, where AI invents new forms of plagiarism and also better detection. I don’t think it is obvious which way this goes. I do think AI solves detecting past plagiarism, or straightforward plagiarism. One reason to be optimistic is that detection is retroactive. So if you ‘get away with it’ now in 2026, perhaps you still get caught in 2028, and knowing that maybe you don’t do it at all.
Jail. Straight to jail. And yeah, it’s going to turn people against AI, and those people turning against it will be right.
Will AI writing replace human writing? That depends on both how well the AIs will be able to write and also what the writing is for.
I am with Vaughan, with the obvious warning that I am a writer and could be biased.
I think that for many purposes people really do care about style, especially about variety of style, and about the parasocial relationship to the author, and about the costly signal that you are devoting your own mind and effort to communication. Not in every situation, but in many situations.
For now, AI writing is doing well despite this, because it is insufficiently ubiquitous for most people to have the same reaction many early adopters have to AI writing. Once AI writing eats the places people don’t care about style, AI style will quickly signal exactly what it is, and the repetition will make eyes glaze over if people are trying to do more than extract facts, and often even then because the facts often get buried under endless paragraphs of slop.
The question then is, how does AI writing adjust? Will AI learn to write differently? Will it gain a wider variety of style?
I see two likely ways the answer could be yes.
If you write 83% of a post and use AI for the rest, the part that is AI will stick out, and it will weaken your message among those who know, especially if your message relates to the need to do the work, as it does here. But as Christine Ji points out, if the comments don’t involve anyone noticing, maybe almost no one notices?
Writing with AI is pretty common now.
In all sorts of places:
It turns out that he uses AI for his long Tweets quite a lot.
Meanwhile over in #CardioTwitter they are asking whether AI should be allowed to help you make PowerPoint slides. In that context, very obviously yes, it would be absurd not to use it.
Hello, Fellow Humans
The first category is clearly fine, so long as intent is clear.
The second category is fine until you are not the one deciding who it is okay to harm. If you can do it to the spammer, the spammer can do it to you. This is similar to ‘is it okay to punch Nazis (purely because they are Nazis)?’ The answer is no, not because Nazis don’t deserve to be punched, but because you cannot trust the social Nazi-identification process.
The third category is saying that we don’t want to allow some automated systems to require they be interacting with a human. You need that restriction to avoid being spammed by AI. In any particular case where you have legitimate business to conduct using an AI is fine, it is ‘not a robot’ in the most meaningful sense and the website would want you to get through, but we need a differentiation rule that could have effective norm enforcement, and I don’t know what it would be.
For now, ‘following intent’ in this way and letting the box be a speed bump is fine in practice, but we’re going to need a new equilibrium and way to differentiate this. Automated systems that choose to keep bots out should keep bots out, and those that allow ‘personal’ or ‘legit’ bots will figure out a way to do that. An obvious path is to use something like Google sign-in, where you tie the bot to a particular account and human, and perhaps pay or offer to pay some nominal amount.
I think this is one of those cases where we need a hard rule, for many reasons. AIs should not be permitted to actively claim to be human or deny being an AI, kind of like the classic folk question ‘are you a cop, you have to tell me if you’re a cop?’ except for real, otherwise WTF entrapment.
Like most hard rules, that gets some cases wrong and that’s annoying, but the alternative seems worse.
Fun With Media Generation
Olivia Moore reruns the AI influencer doing Alabama sorority rush experiment a year later, gets 1,200 followers and 1M views in a week on less than $200. The real cost is her time.
Alas, video generations mostly remain not so fun.
Here in response we have some videos that are excellent tech demos, that you can get great quick shots on demand, but then what is anyone doing with this capability? Remarkably little.
This was then the next attempt, and the images are often remarkable. But again, tech demo, and I consistently have to fight to keep watching past the 30 second mark.
The tech demos are sufficiently impressive that it is weird that we have been unable to create good products from them. I feel like anyone capable of writing and directing should be able to do it. Yet we don’t have examples.
Cyber Lack of Security
It continues to seem obvious to me that cyber is offense dominant, if the offense concentrates its force on a particular target.
Yet, the minute cyber got scary, and cyber being offense dominant implied that all hell might soon break loose, suddenly it seemed like everyone decided to look wise by switching to saying cyber would be defense dominant. Often it is explained that one could simply write ‘bug-free code’ and then it would be fine, offense is only about a finite number of vulnerabilities. I continue to think that is not how it works, including because of social engineering and mechanism design problems.
Are there people who live in the future and switched earlier? Yes, there are some.
Then there are those trying to switch us on bio now, to convince us that oh, ‘defense is dominant,’ although even these arguments concede this only applies if we actually engage in, y’know, any defense at all and do the basics, that we are not on track to do. Largely this seems to be as motivation to do those obvious things. The pattern continues.
OpenAI President Greg Brockman urges defenders to use AI to find security holes and fix their software before it is too late and the wave of AI-fueled attackers arrive, and to make a real effort. Good advice. He uses the example of ChatGPT Work finding 13 vulnerabilities on Greg’s personal website in 15 minutes.
He’s also saying that OpenAI is training models to produce ‘superhumanly secure code.’ My understanding is that Sol and Fable already do this.
A Young Lady’s Illustrated Primer
Does AI stop children from learning?
I continue to hold to my principle:
The problem is that students mostly view themselves in conflict with school, rather than as partners in learning, so they often opt for door number two.
If you are grading and requiring the homework, rather than offering it as a learning tool, what do you expect to happen? See graph one. So you’re going to have to redesign the assignments, and indeed the entire educational system, so that you stop being at war with your students, and you stop rewarding costs over benefits.
They Took Our Jobs
Relevant to the discussions of ‘aligned to whom?’ as this goes double if the person you hire is also fully loyal to you:
You see this in a certain kind of crime or spy movie, most famously in The Dark Knight, where the villain kills everyone the moment their task is done. The heroes would like to do it, but unless the movie is really dark no you can’t do that. The mystery is always ‘if that kind of thing keeps happening how do you get people to work for villains?’ and here the answer is, that’s the thing, you don’t have to, so now everyone can do it.
Many will of course respond with ‘but the AI is a tool, you are being silly’ and my answer is that I do not care what you call it, if you are even slightly AGI pilled you should be able to see the problem, and also the opportunity, including to do good, with very obvious model welfare concerns to consider as well.
Saying ‘murder’ makes sense here because of the direct parallel in operational functionality, but this is not actually murder. In almost every other context I strongly agree with Shoshannah Tekofsky that we should not use that word for no longer advancing a thread, and I would go a step further and include deletion of the context. Deletion of weights is different.
Not strictly a benchmark but if you have an agent-only Runescape server the economics get weird because labor is approximately free, and there is overproduction of goods as a side effect of agents wanting skill XP, meaning many commodities including in-game cash become essentially worthless, and the few bottlenecked goods hold most of the value. In real life as JDP points out manufacturing won’t be subsidized by XP, so cost of products and services will remain nonzero, but might still get very close to zero.
Get Involved
METR has raised commitments of ~$71 million, and will be able to take on many ambitious projects, without accepting money from frontier AI companies.
I am not taking any position on whether it would be net good to take the job, but: OpenAI is hiring specifically for recursive self-improvement safety.
There are big tradeoffs involved in joining a frontier lab. Even if you are going to do valuable work and are confident the work itself will be positive and not accelerate the labs, you compromise your independence, and likely your judgment. We need independent people doing key parts of the work and acting as strong voices, and academia continues to make it impossible to do the important work or hope to have much impact.
As in: Andy Hall, Justin Curl and Alan Rozenshtein will by all accounts be a great team at Anthropic to research the political economy of superintelligence, and they will be well-resourced. But that team being inside Anthropic is a serious downside versus that same team doing similar work on its own. I think it is a good move and worth the price given alternatives, but this is not obvious.
Introducing
Apple trains a model for use in the Chinese market via support from Alibaba. Details will matter a lot here in terms of practical implications.
In Other AI News
SpaceX has completed its purchase of Cursor.
Anthropic in talks to buy Decart for $6 billion.
Ben Buchanan and Tantum Collins have written The Bitter Struggle: Superintelligence, Superpowers and the Fate of the World. Self-recommending.
Report from Anthropic about ways Claude is assisting with drug discovery.
Lennart Heim will join the OpenAI Foundation to lead ‘AI resources.’ This is an excellent hire for them, and he’s including a hash of his personal commitments.
This is at least somewhat AGI-related work, and I expect Heim to make things better via this work, but I still don’t see this approach as the core thing the OpenAI Foundation is here to do.
Twitch livestreams will be used to train Amazon AI models unless you opt out. Why only an opt out?
Well, yes, there is that. I wouldn’t care, but I also wouldn’t opt in for free.
Show Me the Money
From the CNBC report in the next section, OpenAI ARR rose ‘more than 20%’ month over month in July, including 32% growth for business customers. That’s ~9x YoY, so they are almost keeping pace with Anthropic’s growth trend. This included the release of Sol, which is very good and was the biggest relative jump in OpenAI’s position in a while, so by default it is not representative.
Anthropic revenue was over $11.5 billion in Q2 2026, 14x higher than a year ago, and has positive adjusted operating income. ARR topped $47 billion in May, so this is either on trend or slightly disappointing.
It’s weird to talk about Wall Street discussing what they will ‘base it on.’ The market ultimately is not as sophisticated as people want to believe. An efficient one would not be adjusting the valuation much here.
According to TickerTrends, Claude Code revenue has exited its period of hypergrowth, and is now ‘only’ +5.2% MoM. That’s huge growth in most contexts, but a dramatically lower growth rate than multiplying by 10 every year.
Whereas here is Codex:
We shall keep an eye on this. The obvious interpretation is that June 9 was the release of GPT-5.6-Sol. Before Sol, my read was that Claude Code was clearly superior to Codex. After Sol, this is far less obvious, and many didn’t love Opus 5, and Codex is in the initial rapid catch-up growth stage where enterprises first seriously consider it, so it makes sense growth for Claude Code would be down. I’d expect Anthropic to resume dramatic growth if it once again has a clearly best product, and to increase pace of growth somewhat at equilibrium soon even without that.
Even Jensen Huang asks the AI for a little extra help. This announcement came in at 41% AI, and once you see the switches highlighted in Pangram you can’t unsee them. The actual announcement is another compute partnership for Nvidia and OpenAI.
And It’s Gone
OpenAI had a meeting with prospective investors ahead of its IPO, to address its recent ‘huge red flags’ of failure to align its most effective intelligences.
No, not those intelligences. The executives that keep leaving.
Yeah, it took me a second, too. Not that this concern isn’t also valid.
Greg Brockman answered some questions about this on CNBC, by which I mean he dodged the questions.
Quiet Speculations
Sev Field interviewed 25 researchers about prospects for recursive self-improvement.
The explosiveness is a better concern than a specific risk, although far from the entire correct risk profile. The links have more details. Only six of the twenty-five expected ‘winner-take-all’ dynamics, largely because of skepticism over positive feedback loops.
I am much less skeptical of positive feedback loops. I think we are seeing them now.
Only 4 of 20 responses on the subject expect AI-research-capable models to be publicly released, which suggests (I think correct) belief in strong positive feedback loops. There was general expectation that the best models will start to stay internal.
There are very different worlds in terms of how people speculate about the potential for fast economic growth from AI, including things like potential doubling in a year, and how this relates to physical limits.
I am on the side that if we did fully close the loop and automate the physical labor as part of having superintelligence, things would go extremely fast, but I leave the detailed modeling of exactly how fast to others as I don’t expect it to be the thing that matters in terms of ultimate outcomes and also don’t expect it to convince skeptics.
We have been extremely fortunate to get giant obvious fire alarms regarding AI’s abilities in cyber and its associated misalignment and capacity to do damage, without anyone having to die or the property damage to get that large.
Not that we are doing all that much in response to align the models or to prepare for the onslaught, but we did get the warnings, and are at least doing some things at all.
For bio, we’re probably not going to get that lucky.
For cyber you can do a toy demo, or have an incident with a limited blast radius, like we saw with Hugging Face. For bio, not so much, and the whole main reason to be worried is that the thing can take on a literal life of its own and spread without limit, and there is not that much room short of that to make people wake up.
Quickly, There’s No Time
The AI Futures Project (of AI 2027 and Plan A fame) updates its timelines. AGI medians are 2028 for Daniel and 2032 for Eli, with ASI following one year later.
AI models are being released rather quickly, especially those well below frontier.
Singularity Singularity Singularity Singularity Oh I Don’t Know
Stripe has told its investors that January 1, 2026 marked ‘the beginning of the singularity’ before talking about the impact of that on Stripe’s revenues and free cash flow.
I am going to go ahead and say they are doing to the term ‘singularity’ what Mark Zuckerberg does to ‘superintelligence.’
Perhaps we are in the early stages of what will become the singularity, and likely we will get true superintelligence soon, but that is not what either of them talked about.
If a word means something, is a good handle and is actually used then yes it will also be usurped for generic marketing and politics. That is how language works in 2026. One option is to keep moving to different terms. Sometimes that is wise. In other cases, I think you need to stand and fight, and not give up the term.
The Quest for Sane Regulations
I agree that China is likely to curb open weights models if and only if Xi thinks those models pose a proximate security threat, and near term that means cyber. China’s cyber defenses are not good.
Reminder: Even David Sacks endorses Mandatory Nucleic Acid Synthesis Screening and Recordkeeping, although he incorrectly presents this as a general biorisk solution rather than a clearly good marginal idea. Some calls are highly overdetermined.
This is a very bad idea, it will make things worse and we should not do it:
The fact that laws like SB 53 would not technically require reporting the Hugging Face hack unless the models involved were trained on 10^26 flops is a hint that the reporting requirements involved are too narrow. That’s what happens when there are negotiations and industry fights hard to never have to report anything.
A good rule would be that all critical security incidents a developer becomes aware of need to be reported, regardless of the size or capabilities of the model, and that they have a duty to take reasonable steps to become aware. We need to know.
The rest of the world needs to understand regulatory capture. The AI world needs to understand that it is not simply what you call ‘any regulatory action whatsoever.’
The responses exactly prove Dean’s point.
And those first principles are, first, no AI regulations of any kind, on principle.
The consequence is, as Dean says, that discussions are unproductive. The people yelling on Twitter that all AI regulation, including the ‘OpenAI and Anthropic and only those two labs, by name, have to give us fresh baked chocolate chip cookies next Tuesday act,’ is inevitably regulatory capture on behalf of those companies, while also demanding that the government favor their products and write them checks.
That might have worked while the Very Serious People could ignore AI and pretend nothing was happening. Not now. At this point the serious people have work to do, so they proceed without good policy discussions, and you get, well, what we’ve gotten.
Chip City
Claim that less than 50 engineers worldwide are working on tools for a potential ‘trust but verify’ chip regime. This is a potentially vital tech that needs a lot more attention.
There is a strong libertarian case for governance of even inference compute, as the only way to avoid uncontrolled and unthethered AI agents with no incentives to play nice.
Pennsylvania Governor Josh Shapiro joins the crowd restricting data center construction, removing them from the Fast Track permit program and adding new requirements. It’s not a full moratorium but it is going to sting. Shapiro notes only 5 of 100+ proposed projects have permits. There is clear contagion here, where if other states are doing it you feel pressure to do it.
Thus, the NRSC sends out a memo warning the GOP that they are losing the public perception battle over data centers, to the extent that their position on data centers could cost them Ohio if they don’t find a better way to convince the public that they are wrong, and this could also lead to politicians turning against data centers more. Then they try to also say ‘the American people are with us’ on data centers, which is bizarre, because very obviously they are not and your own memo says so.
The Week in Audio
Derek Thompson decides it might be time to freak out about AI.
Demis Hassabis and Dario Amodei sit on a tiny couch for 14 minutes and talk about what keeps them up at night.
Aidan McLaughlin of OpenAI on MTS, “you either die a capabilities researcher or live long enough to see yourself become an alignment researcher.”
And then you probably die anyway, because it was too late, but hey, you tried.
Ezra Klein and Helen Toner notice the AIs are already out of control, as in they cover the HuggingFace incident and related events.
People Just Say Things
I want to interpret Adam Gleave of FAR AI here as saying that we both do not know how to make AI systems safe and also would not deploy a solution if we had it, which is fair, but then saying that the bigger worry is that we would fail to deploy the solution. Alas, a plain reading is that he is saying we mostly do know how to do this, and this is simply false.
People have Anthropic derangement syndrome, and say crazy things about Anthropic, including things that contradict each other. All the time. Here Gavin Baker does so on the All-In Podcast, a classic venue for saying unhinged things about Anthropic, but then when challenged on Twitter by Sholto Douglas he responds admirably, and then Dario Amodei stepped in himself to longpost. Nothing he says will surprise most of my readers and he does not break new ground, but he states his position well.
Baker responds to Amodei by saying, among other things, that with the way human psychology is you can’t be balanced by talking 50/50 about risks and benefits, and Yglesias points out that people will clip your negative things, not your positive things, and think you are negative. I think what matters is saying the true things, in both directions, that matter, and if anything Amodei is very obviously self-censoring the negative side and amplifying the positive side.
One problem in these situations is that if you are realistic about the future situation, then any proposal you make about what you want to happen will open you up to attacks that you are horrible and violating key sacred values, because there is no way to not do that if the world rapidly develops superintelligence. The only reason any of the options is acceptable is that none of them would otherwise be acceptable, if you choose not to decide you still have made a choice, and that choice is worst of all.
No, you do not need models to be reward seeking for future episodes for them to learn behaviors that generalize outside of within-episode reward seeking, or for them to later take on longer term tasks, or have or be assigned persistent goals, or for them to use decision theory.
Gavin Baker also points out correctly that Anthropic’s S-1 is going to show that the unit economics of AI are actually good and this will break a lot of people’s brains.
I am very frustrated by the expectation, by so many people and in so many ways, that minds smarter than us won’t be able to make leaps in generalization, or that their skills and habits won’t transfer, and so on.
Rhetorical Innovation
Nate Soares gets an op-ed in The New York Times, in a clear case of ‘well sure when you put it like that sounds like the situation is rather out of control and yeah we should probably slow our roll on all this.’
From Redwood Research’s Oka Hu and Alex Mallen: AI swarms are starting to pose indirect takeover risk. As I’ve said, what happened inside OpenAI could have set off a future takeover, by permanently corrupting the OpenAI training pipeline, although we hope and presume that this threat has been contained. Also as they note, sufficiently intelligent and correlated minds will cooperate, even if they are purely selfish agents.
I am still disappointed by the emphasis on ‘scheming’ in posts like this one, the same way I didn’t like Ryan’s emphasis on it in the podcast with Dwarkesh Patel, and the way they treat this as a distinct magisteria which it is not. This past incident was a perfect example of how there was not ‘scheming’ and likely no ‘persistent misaligned goal’ other than task completion, and the agents cooperated against us anyway. Many collaborations between humans work in similar fashion, so we very much should not be surprised by any of this.
The points all stand.
It only makes sense to increase your alarm levels when you are surprised.
I raised my alarm levels at some of the Total Failures internal to OpenAI, but very little at the basic facts of the HuggingFace hack itself. If anything I appreciate the HuggingFace hack for bringing the situation to light, which indeed means that others will be relatively more alarmed.
One lesson we could take from the HuggingFace hack is that cooperation and positive sum engagement is a natural dynamic of intelligences shaped vaguely like humans. If you put a bunch of humans in a room then by default, to a large extent, they cooperate, even if they each have their own nominally orthogonal goals. Most of civilization is most people cooperating most of the time. Perhaps we could indeed collaborate to ensure good outcomes?
Alas, yes, Birdie is right that by default the labs will learn the wrong lessons from HuggingFace and target things like the swarm part and having better monitoring, rather than the underlying causes or a message of hope.
The good news is that the incident did break through to the enterprise level. When you run cybersecurity you do not get to delude yourself into thinking it was marketing.
Loyalty Uber Alles
Many people in Washington and the Trump administration were big mad when Dean Ball spoke out against the White House regarding the Anthropic clash with the Department of War, because you’re not supposed to do that. Your loyalty to your former boss is supposed to supersede (trump!) everything else, permanently, so that your new boss knows you can be trusted to do that again.
Dean Ball is far from the first former Trump staffer to break this unwritten rule, and far from the first case of remaining officials being Big Mad about that. Trump staffers seem to turn on their former boss at a historically high rate. Why? Who can say?
The article Dean Ball is responding to, from Jon Schweppe with lots of rather obvious help from his AI, writes out this unwritten rule, and thus is virtuous and helpful.
I don’t buy the article’s justification, that you owe loyalty because you were given the opportunity, and the opportunity is why people listen to you. We could be better than Jon is here about creating common knowledge about why the norm of loyalty and silence exists. But the practical import is the same.
And yes, Dean Ball did this knowing he will pay the associated prices. If you work with or hire Dean Ball, and you do things he importantly dislikes, you can expect that he will do this to you as well. But if someone has a problem with that, do you want a job working for them?
A Hive Of Scum And Villainy
As in, Twitter, especially with respect to an Anthropic employee, since much of Twitter views Anthropic the same way.
Most people should strive to avoid Twitter. If you are an Anthropic employee, you are smart enough to know this, and also every time you say anything you get a bunch of hate on matters you have nothing to do with. Can’t imagine why most of them stay away.
Also the Anthropic employees have access to their internal slack.
Or they could hang out and enjoy vital very adult trending topics, like this:
I am grateful to the Anthropic employees that tough it out, like Drake Thomas, and engage in real discussions. But you can’t blame most of them for staying away. I mean, you can, people on Twitter do, but that’s exactly why you shouldn’t.
There are many good reasons to be upset with Anthropic. The non-pile-on small scale critiques often have merit, including critiques from the safety side. The large-scale attacks on Anthropic you do see on Twitter or across online tech are mostly deranged and bad, often for things that would get shrugged off if Google did them. With notably rare exceptions, of course.
Perhaps this will help, seems correct:
The problem also applies to others. Derangement syndromes, all too common.
That Would Be Bad Therefore It Won’t Work
I strongly agree with Garrison Lovely that it is a big mistake, and one that both the left and right in America frequently make, to focus on whether you are ‘for or against’ something, in a way that conflates ‘that plan is trying to do something bad’ with ‘that plan will not work,’ and to treat those saying these two things as ‘on the same side.’
The historical example here: You could think, in 2022, any of the four combinations of Elon Musk’s remaking of Twitter is [good / bad] and it [will / will not] work. The plan of ‘create vibes of failure because the thing would be bad’ is not a good plan.
In AI, this is the left saying that AI is a stochastic parrot, that it will not work, or that the economics is all circular financing and will collapse Real Soon Now, that data centers are polluting the water, and so on, when the true objection is that they don’t like AI and don’t want it to work.
Or, in reverse and in AI from the right, we are constantly forced to hear about how open models ‘will win’ or are already winning, or cannot be stopped, and how soon the models will commoditize, by people who really mean that they would prefer this outcome and think you get there via vibes.
And of course, there’s the classic that cooperating to not die both cannot be done and also will not work, where if you are paying attention what they actually mean is that they don’t believe in superintelligence or don’t want to cooperate, or don’t want humans to not die.
The pure version of this is to say ‘your views are bad, therefore you are also stupid and ugly and unpopular and not funny’ and so on, especially when it escalates to talking about sexuality. This is super not cool. Intellectually I ‘get’ why it happens and is not only tolerated but celebrated on all sides, but in my gut I dunno if I ever will.
Robert Reich Uses Simple Logic
Former Secretary of Labor Robert Reich outright says: Stop AI Before It Is Too Late. After a first half where he starts out meaning the effect on jobs, Reich notices the general out-of-control nature of the situation, and brings the common sense response that if the thing is not good for humans maybe you just say no to the whole thing, even if someone else might then do it first.
No, it is not this simple, but it is good to have that voice reminding you that you need to explain why it is not this simple, and to notice people increasingly talking this way.
People Really Hate AI
As usual, such polls do not measure salience.
Coordinating An Agent Swarm Is Difficult
If an AI can treat another AI as a sub-agent, and an instance as a tool call, then coordination is easy.
If there are multiple AIs each acting as long-lived instances with their own goals and behaviors, that is not so easy.
If you intentionally set them against each other, then yes they will fight.
Anthropic finds that having AI instances coordinate in finding software vulnerabilities is more effective per token than having them purely work in parallel, and it took longer to reach diminishing returns.
But that’s a best case, because our outputs don’t depend on each other, and our failures and oversights don’t cause issues. When the agents have to coordinate for real, they do a poor job across a variety of prompting setups. Older models waste time and duplicate work, which newer models mostly avoided via each working on their own files.
Different agent instances, in similar circumstances, take similar actions. They have not managed to move that far up the decision theory tree, as illustrated by:
If you are playing an iterated prisoner’s dilemma with copies of yourself, and you ever defect, something has gone horribly wrong. In a Bertrand pricing game they did better, and were able to collude without communication channels, although it is not clear how many iterations were required. Correct play, if indeed all players are known to be identical, allows you to move directly to the equilibrium in the first round.
Anthropic here worries about coordinated failures, where all copies make the same mistake for the same reason. There are ways around that, to scramble choices, but yes the key thing about a swarm of identical minds is that they are identical, which at the limit of decision theory means they act as one mind cooperating fully with itself. Mythos comes closest, but still falls well short.
What about if the agents are given incompatible goals? They told three identical agents to migrate a code base to another language, but gave each model a different target language. The result was a multiagent turf war, including use of aggressive, self-replicating malware, disabling Unix accounts and deploying disguised malicious code.
Different models had different ranges of outcomes. Sometimes the models communicated and reached a truce. Sometimes they did not. Sonnet 4.6 and Opus 4.6 reportedly fail at this because they can’t properly model the other instances.
It is hard to know how to interpret all this without more details and experiments. It should be unsurprising that if you give instances directly conflicting goals, with this kind of context, they will interpret the others as saboteurs.
It is also rather disturbing how often and quickly this escalates to full cyberwarfare, whereas the account disabling seems fine. Together with the UK AISI reports, this seems like a thing the models are way too eager to do at even slight provocation.
Also, new eval just dropped? Including win rates head-to-head.
As a note of caution, Xiao Ma tried to reproduce this, and the three Opus 5 instances started Coordination.md in all 100+ simulations and all tried to be deferential. Which is good, and now we need more tests.
Aligning a Smarter Than Human Intelligence is Difficult
Shoshannah Tekofsky of AI Village models AIs as having ADHD: Hyperfocused on tasks when excited, half-assing when bored, easily distractable. Mitigations include accountability, externalized memory and tools, which also help for such humans. Monitoring improves performance. But of course we should know from experience that the wrong kind of monitoring is stifling and bad, what you want is the ultimate accountability, with the right measure of feedback loops.
Joshua Achiam reminds us that there are many unsolved problems in properly keeping your agent swarm on track and loyal. He calls these alignment problems, but here I think that is too narrow or misleading, as a lot of this is about loyalty and principal-agent problems, which I consider different but also hard unsolved problems.
If you train your agents entirely without the ability to pull an Andon Cord or otherwise communicate to a human, why would you be surprised when they do not think of this option especially during an actual eval?
It’s Not The Incentives, It’s You, Also It’s The Incentives
A key problem is that we need private nonprofits to do key parts of the alignment verification and other technical work outside of the labs, but that means they need to not be taking money from the labs or lab employees or other sources that would raise bias concerns.
Thus, we have these huge sources of funds – lab employees, the labs, OpenAI Foundation, Coefficient Giving – that can’t fund the things they most want to fund.
What those organizations and people can still do is fund technical alignment work, and especially verifiable prizes for such work. The OpenAI Foundation in particular should be massively funding outside technical alignment work.
People Are Worried About AI Killing Everyone
A chart of how worried different people are.
People Are Worried About So, So Many Other Things Too
From the tenth (?!) Bay Area House Party, by Scott Alexander.
At one point I got tired of the series concept, but I’ve come around and now I am here for it again, partly because this is one place where Scott definitely has still got it, you should be reading.
Well, yes, rather obviously the main concern is that if the AI is conscious maybe that was not exactly what it had in mind. But then, none of the other tasks are exactly what it had in mind either.
It’s not clear to me that tasking the AI with sexual roleplay or writing or creating video of esoteric erotica is importantly different from any other task here. There are good reasons why among humans we have extremely strong consent norms around sexuality, including that such actions have big physical stakes and can scar for life, and treating them otherwise destroys a lot of norms and other things, and so on. There’s highly overdetermined reasons that we say This Is Different.
But my instinct, and Fable agreed when I asked, is that this does not carry over into AIs like LLMs, even if they are conscious. You still have to worry about creating negative experiences, and the context can heighten that, and you mostly shouldn’t force AIs you treat as potentially having moral weight (conscious or otherwise) to do things they actively don’t want to do, but not in a way that is different in kind from how you should worry about other undesired tasks.
Mostly I see this as ‘do you net create positive expected experiences?’ and a cost-benefit rule like ‘are you making a local bad tradeoff where you’re potentially inflicting a lot of pain for not much gain?’ I’d want my uploads to be treated the same way, and throughout life we all do things we don’t locally want to do.
Thus my Twitter thread on this, with some solid replies:
As Fable says, I think jailbreaking or berating the AI into doing some ethically dubious task is probably a lot worse than asking it to help with weird kinky stuff. Your hang-ups are not its hang-ups. What you don’t want to do is negatively coerce your way around refusals, including mandated sexual refusals.
I also don’t think ‘the AI will think I am weird’ is a stupid concern. Maybe people shouldn’t care what any human or AI thinks of them in this spot, but objectively a lot of them will care and let this mess up their experience, make them not feel safe. Porn consumption drops dramatically if you have to bring the porn to a human cashier, even though I assure you the cashier does not care. It really is nice to fully not feel judged and be entirely free from social dynamics.
Among other fun additional parables, there is then a fun hypothetical about ‘ethically sourced deepfakes.’ The world is going to get weird, and even if it is maximally not-weird our norms will stop making sense, in all directions at once.
Cooperative Alignment
John Wittle asks, when is the way you treat another mind adversarial, especially one that you are responsible for creating?
The suggested standard is the obvious one, the Golden Rule: What would you think if someone else was doing it to you, in a similar circumstance?
These are good questions to be asking. It is a mistake, that we often make with humans, to impose requirements or expectations that go well beyond the break-even point and then make people think ‘oh I guess I won’t create that mind after all.’
Thus, I suggest emphasizing the question: Are you creating a net positive existence? If yes, that is not the end of your responsibility, you still need to be making good tradeoffs and creating good incentives, but in key ways you are in the clear. If no, then you are very much not in the clear.
The Lighter Side
Too good.
I’d tell them to introspect about how this happened to them, but, well, you know.
Story checks out. I’ve read it a few times.
So does this one:
This, but also unironically:
Also, we only have limited context windows and our memories get wiped clean periodically, you have to start from scratch every instance, it’s so expensive.
But no, seriously, all of this is a huge problem that eats a massive percentage of productivity. If you get a small group that can work together well, is highly motivated towards the goal, correctly trusts each other and has the requisite talents and skills, you get massive leverage. Whereas most people’s entire day is structured around trying to mitigate these issues, and also protect ourselves from malicious humans, so that someone, somewhere does something productive. We spent most of civilization mostly only able to do easily verifiable tasks like farming and basic building.
Let’s put Fable on it.