Could anyone explain why Anthropic's 2028 scenarios don't try to explore the possibilities of one of the AIs becoming misaligned and fine with destroying all life in the world, then recreating it?
In quickly skimming this, this piece seems to be aimed at policymakers to encourage them to enact GPU restrictions on exports to China. Comms 101 is don’t confuse your message.
(To be clear I’m not defending Anthropic here, just saying why I think this was written)
David Sacks on X is tone-deaf, especially in bold:
... (read more)I don’t see what you see in the lab. If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible.
But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. Stop pretending you need a regulatory approval process that supersedes product liability. Stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff. Stop pretending you need those same evaluators to police competitors who aren’t even at the frontier.
Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability. Call it alignment if you want. It is also just giving customers what they want.
Pacing the frontier would also create breathing room for a more intelligent conversati
I do think David has some good points, though. OpenAI and Anthropic could’ve started screaming for much stronger national and global regulation of AI much earlier than now - I remember discussion on the RSPv3 announcement post on this, I think. They could’ve tried to throw their weight around to get all the other orgs to cooperate with an industry pause. Their attitude towards AI development has been “move fast and break things,” but when it comes to trying to coordinate a slowdown, suddenly they start mumbling about their hands being tied by antitrust. METR has relied on the goodwill of the labs for access and tokens, and is partly funded by an Anthropic board observer. And there really is a compelling business and regulatory risk case for getting your AIs to not commit crimes!
I don’t agree with everything he’s saying. Obviously I want every company to be audited. And I want us to actually try coordinating with China. But if someone doesn’t trust the frontier labs, it can be rather difficult to tell if they actually care about the long term future of humanity, or if they’re just following their local incentive gradient.
On China's Zhihu — the country's highest-quality Q&A platform — under discussions about Jacob's resignation and the recent statements by the "AI Big Three" regarding an AI slowdown, there are a few answers about AI alignment. However, most people know nothing about alignment and do not believe in AI risks, instead framing the issue purely in terms of marketing hype, the breakdown of scaling laws, and market confidence. Given that Zhihu is considered the platform with the highest quality of discourse in China's Q&A space, the situation is not optimistic.
It seems very unlikely an AI-generated article from 'The Trumplandia Report' nor a different AI-generated article from a different source of equivalent dubiousness are going to include notable leaked information from SSI. I would be very willing to bet against this information being true[1].
It has been long rumored that SSI specifically has (or is aiming to do) continual learning, ever since at least Sutskever's interview with Dwarkesh; if there was an actual leak confirming this rumor, it seems very likely to me it would be either first sent to the press or to a highly trusted AI lab whisperer with an excellent track record (think: Andrew Curran[2] or 'Arfur Grok'[3]), not to a few minor AI-generated newsletters.
though I as a college student do not have high liquidity to do so at non-token sizes, you can PM me to figure out something if you want to do so.
Note that due to SSI's infamous secrecy, operationalizing this bet will likely be difficult, since we want to bet specifically on the rumor's veracity (that SSI already has continual learning and is likely scaling it up greatly, not 'will SSI announce continual learning in a year, before a frontier lab does'); however, most oper
if there was an actual leak confirming this rumor, it seems very likely to me it would be either first sent to the press or to a highly trusted AI lab whisperer with an excellent track record (think: Andrew Curran[2] or 'Arfur Grok'[3]), not to a few minor AI-generated newsletters.
Andrew Curran did post the following recently:
Looks like this is going to be a big week. All the main players have releases in their final stages, and it's possible they all arrive over the next five days. There's also a model in early access from a startup driving a lot of the recent hype vagueposts from prominent accounts.
I don't know much about it except rumors, and nothing is verified, but if you've played this game for a while you can tell when something feels real, and this one does. Apparently a breakthrough in continual learning, but not from one of the big labs. We shall find out together.
I predict that this refers to SSI.
The IABIED march, which was supposed to happen if 100 thousand people sign the pledge to march in Washington DC, had 898 people sign the pledge and 1630 people sign up to be notified, which is 2 OOMs less. Therefore, the march's organisers overestimated either the amount of protesters necessary to gain leverage or the biggest amount of people who would be ready to protest. Does it mean that Washington is a bad place because it has only ~700K inhabitants and that a protest in the NYC would recruit ~10 times more people? Or that the IABIED statement requires... (read more)
@Zvi somehow missed the first Solid Result from EpochAI's FrontierMath Open Problems benchmark.
Consumers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact
I hope that the judges are sane enough to prevent this from punishing AI labs for focusing on safety...
To solve the Navier–Stokes problem, OpenAI used an internal model that is significantly more capable than GPT‑6 Astra.
GPT-6.7 Supernova, go delete yourself!
Why is @Daniel Kokotajlo's estimate of METR's doubling time used for Q1 2026 Timelines Update four months instead of 5-7? I see the following counterevidence and doublechecked it with Claude Sonnet 4.6:
How many months behind is laypeople's perception of the AIs' capabilities? I have enocuntered a video released a day ago on a popular YouTube blog claiming that AI is a technology with no lived experience and zero reasoning abilities. How does such a misperception allow us to explain the threat caused by the AGIs, which have yet to emerge, to laypeople?
The ARC-AGI team evaluated Claude Sonnet 4.5. On the ARC-AGI-1 leaderboard, OpenAI's models o4-mini, o3 and GPT-5 formed a nearly straight line. Claude Sonnet 4.5 was slightly below said line when thinking in 1K,4K or 16K tokens and held the line when thinking in 8K or 32K tokens. This could imply that the benchmark has scaling laws for non-distilled models, and OpenAI and Anthropic reached these laws.
On the ARC-AGI-2 leaderboard Claude Sonnet 4.5 Thinking became the new leader between $0.142/task and $0.8/task; it also completed 13.6% of tasks... (read more)
The AI sycophancy-related trance is probably one of the worst news in AI alignment. About two years ago someone proposed to use prison guards to ensure that they aren't CONVINCED to release the AI. And now the AI demonstrates that its primitive version can hypnotise the guards. Does it mean that human feedback should immediately be replaced with AI feedback or feedback on tasks with verifiable reward? Or that everyone should copy the KimiK2 sycophancy-beating approach? And what if it instills the same misalignment issues in all models in the world?
Al... (read more)
How similar is the AI-2040 revenue forecast as opposed to the Q2.5 revenue forecast to the news about Anthropic's slowdown? I tried to understand the real trend two times. What I found was the growth of the revenue of Claude Code and Codex combined had a temporary slowdown, which proceeded to accelerate again. Could anyone explain this?
I looked up EpochAI's list of benchmarks. The very poor ECI of Grok 4.5 as opposed to GPT-5.5 seems to be a result of xAI not caring about math. Additionally, Grok 4.6 seems to be closer to Opus 4.8 with respect to ARC-AGI-3, if not outright outperforming Opus, as the actual ARC-AGI-3 score suggests. Does it mean that xAI is less behind than we think? I guess that Zvi will have to write something like "Grok 4.6 is three, not six, mounths behind. Stop xAI to hell!"
P.S. The same issue seems to apply to Meta's Muse Spark, except that it has even less evaluate... (read more)
Also, SpaceX has caught up to Meta in credibly having enough compute in 2027-2028 [1] to stay in the game, if either of them can assemble a functional model development team. As Google illustrates, it's not easy to do that (even when you have some of the best people), but as OpenAI and Anthropic illustrate, it's not so difficult that it can't be replicated. Muse Sparks are probably small enough models that their non-frontier performance doesn't count as evidence that they're not well-made (and that a Mythos-sized Muse model won't have Mythos-level capabilities). SpaceX is in a worse position in 2026 on priors, because the xAI team wasn't doing that well in 2025 and then got disrupted in early 2026, but they aren't in a worse situation than Meta was in 2025, so it remains plausible that in 2027 they catch up.
The public part of the text article leaves many gaps in the arguments (wildly gesturing to a somewhat unusual extent at the presumed arguments hidden behind the various paywalls), but some relevant things were discussed in the video version. Basically, it's a feasibility and track record argument. I'm guessing there's an assumption that others aren't competing too st
Meta's Watermelon is rumored to match GPT-5.5's performance. If we got another lab two months behind the frontier, then Meta will have to publish a BIG model card to compensate...
I have encountered an issue with editing links. When a link's edition gets close to the 'Save' button, I find it hard to save the link and not the comment.
The scenarios related to futures of mankind with the AI race[1] by now either lack concrete details, like the take of Yudkowsky and Soares or the story about AI taking over by 2027, or are reduced to modifications of the AI-2027 forecast due to the immense amount of work that the AI Futures team did.
The AI-2027 forecast in a nutshell can be described as follows. The USA's leading company and China enter the AI race, the American rivals are left behind, the ... (read more)
@Cleo Nardo What do you mean by the hypothesis that "if we score highly transcripts which look good to a human and score poorly the transcripts which look bad to a human, then the model would be aligned to human values"? I find it unlikely for two reasons:
RR views on "conceptual uplift stuff" aren't yet public, but they are currently working on writing up their thoughts about this. It is unclear to me whether this will then be made public, but it seems it will at least be available to in group people who ask to see it. -- bpomo's Shortform
AI is not doing the integrative reasoning that would find new connections and develop new insights, the kind of work where new techniques and new discoveries are made along the way that push the field forward -- Taylor G. Lunt on Richard Ngo's Shortform
My worst-case scena... (read more)
The AI-for-epistemics section of AI-2040 seems to be overly optimistic about a potential positive basin for the following reasons.
When I was writing my last post, I wasn't aware that xAI released Grok 4.6. Nor that Meta planned to open-f**king-source Muse Spark 1.2.
I struggle to understand the main reason why it's so hard to explain AI sceptics why they should believe in superintelligence (UPD: emerging soon). Is it due to the obsolete concept of soul or due to AI systems being applied for things like recommending the next video to watch? What analogies could one use to dismantle the sceptics' disbelief?
"Belief in superintelligence" isn't specific. A lot of AI scepticism is about claims of what happens when, not about what's possible in a million years. And it's easy to be wrong about the more specific claims. So under many specific senses of "belief in superintelligence" that someone might contest, the purpose of "dismantle the sceptics' disbelief" might be epistemic violence or a bottom line written before an argument. Conversely, the sense of "belief in superintelligence" needs to be specific enough for it to be credible that it's robustly and objectively correct, rather than a miscommunication about something genuinely contentious, a different epistemic status following a different intended meaning.
What analogies could one use to dismantle the sceptics' disbelief?
Plan A used to rest on control of powerseeking Ais from 2032-35 and apparent success seekers from 2030-32. What does Zvi's most recent post on OAI's fiasco imply about such a strategy?
@habryka Could you check the Leaderboard for bugs? I don't think that I gained 318 karma in the last month.
GDM contributes a lot to AI safety, but Geminis, like the cobbler's children who don't have shoes, stay hardly evaluated since Gemini 3.1 Pro...
Steven Veld et al[1] just released a new modification of the AI-2027 scenario as a part of MATS.
The main differences are the following:
MATS just released a new modification of the AI-2027 scenario.
MATS doesn't release things. MATS is a training program! This is not some kind of official MATS release. I would phrase this differently (like saying "A MATS scholars just published")
ARC-AGI-1 performance of the newest Gemini 3 Flash and the older Grok 4 Fast implies a potential cluster of maximal capabilities of models with ~100B params/token. Unfortunately, the potential cluster didn't have any company try and create more models of such class.
After introducing the ARC-AGI-3 benchmark, the team decided to measure the performance of Grok 4.20 (presumably Grok 4.20 as of March 9?) on ARC-AGI-1 and ARC-AGI-2. Grok... demonstrated its capabilities. How likely is it that Grok has stopped being a train wreck and became something worthy of being tested? What could one do to have Grok tested on other benchmarks?
While GPT-5.1's improvement on the ARC-AGI-1 benchmark was mostly incremental and continued the straight line described in my prior analysis, Gemini 3 Pro and Gemini 3 Deep Think Preview scored, respectively, 75% and 87.5% on ARC-AGI-1, while having cost $0.493/task and an unknown cost, presumably $44.26/task. For comparison, o3-preview reached 75% for $200/task and 88% for over $1K/task.
We don't know anything about Grok 4.1, but Grok 4 Fast scored 48.5% for $0.031 on th... (read more)
GPT-5.1 failed to find a known example where Wei Dai's Updateless DT or Yudkowsky-Soares' Functional DT yield different results. If such an example actually doesn't exist, then should they be considered as a single DT?
It looks as if scaling laws of various benchmarks tend to be multilinear:
The ARC-AGI leaderboard got an update. IIRC, the base LLM Qwen3-235b-a22b Instruct (25/07) is the first Chinese model to excel at the Pareto frontier. Or is it likely to be closely matched by the West, as happened with DeepSeek R1 (released on Jaunary 20?) and o3-mini (January 31)? And is China likely to cheaply create higher-level models like an analogue of o3 BEFORE the West? If China does, then how are the two countries to reach the Slowdown Ending?
The two main problems with the slowdown ending of the AI-2027 scenario are the two optimistic assumptions, which I plan to cover in two different posts.
Why does the Race Ending of the AI-2027 Forecast claim that "there are compelling theoretical reasons to expect no aliens for another fifty million light years beyond that"? If it's false, then sapient alien lifeforms should also be moral patients in a way. For example, this implies that all or almost all resources in their home system (and, apparently, some part of space around them) should belong to them, not to humans or a human-aligned AI. And that's ignoring the possibility that humans encounter a planet having the chance to generate a sapient lifeform...
If we were to view raising the humans from birth to adulthood and training the AI agents from birth to deployment as similar processes, then what human analogues do the six goal types from the AI-2027 forecast have? The analogues of developers are, obviously, the adults who have at least partial control over the human's life. Then the analogues of written Specs and developer-intended goals are the adults' intentions; the analogues of reward/reinforcement seems to be short-term stimuli and the morals of one's communities. I also think that the b... (read more)
It has never happened before, and here we go again...