Thank you for the summary,
But where is mention of world models in AI 2027 and 2040 or in your summary? Don't all the robots and AIs driving them need to understand the physical world before taking it over?
Thanks in advance!
What exactly do you mean by world models, and why do you think they're necessary? I have some initial thoughts on what I think you mean (e.g. I can imagine a future in which an AI-assisted coder or AI coder develops a program that can model external environments, which can then be used by LLMs that themselves don't have world models) but I want to make sure we're talking about the same thing first.
My thought is that LLMs can probably program physical "muscle memory" for themselves with basic multimodal inputs + tool calling capability then control them just fine. The muscle memory themselves would perhaps need a world model to execute akin to a brain's motor cortex but those models probably don't need any capability beyond functioning as a motor cortex, i.e. "follow simple instructions and forward important surprises to the LLM", so there's not really much doubt that they can be developed.
This post will only make sense if you already know about AI 2040. If you don't, consider reading the authors' announcement post (Substack, Less Wrong), or reading the full scenario.
Alternatively, if you want to read a longer unofficial summary of AI 2040: Plan A, consider reading Scott Alexander’s introduction and reaction post or Zvi Mowshowitz’s introduction and reaction(s) post.
Intro
I did not like AI 2027 and expected to feel similarly about AI 2040: Plan A. Having now read Plan A, I can safely say I was wrong: I vastly prefer it to AI 2027 and sincerely hope more people read it. However, Plan A is long and inaccessible to all except the most committed of readers, in large part because the authors decided not to include an executive summary.
The authors of AI 2040 explain:
I disagree. I think many people will not read AI 2040, including some who probably should, precisely because it is so long and inaccessible. I also think those people would benefit from reading an executive summary-style post describing it. This post is my attempt at writing that summary.
The Argument
The AI Futures Project authors believe that the world should agree to an international AI-race slowdown treaty—basically just an arms control deal—that still advances research aimed at aligning and controlling AI. They call this Plan A.
Their argument is as follows:
This very helpful image comes from Astral Codex Ten, not AI 2040.
The Plan
Executed perfectly, Plan A looks something like this:
The authors explain this diagram in more detail here. Author (and diagram creator) Daniel Kokotajlo further explains the diagram here.
from AI 2040
Roadblocks and Failure Modes
The authors are acutely aware of Plan A's many downsides. Below are seven potential roadblocks and failure modes.
Coda
I plan to write another post focusing on the potential cruxes, points of disagreement, and strongest criticisms of Plan A. For example, I wish that they’d explained their robotics automation forecasts in more detail. I also wish they’d talked a bit more about what happens if AI capabilities progress is slower than they anticipate, which would significantly affect the dynamics at play.
But overall, I think Plan A is excellent. No other plan comes close to being as comprehensive, and I’m confident that the authors’ work will be put to good use even if the plan falls through.
The START treaties were a series of nuclear weapons disarmament treaties between the U.S. and Russia.
That being said, the authors are generally sympathetic to Plan S. Excerpted from them:
Do note that their preferred versions of Plan S strongly resembles Plan A:
While Plan A strictly applies to the whole world, the specific scenario primarily revolves around the U.S. and China, the two leading players in AI. Specifically, Plan A assumes that the U.S. is acting optimally in the best interests of itself and the world, while China is acting as it actually would in real life. This is not a safe assumption for a purely probabilistic forecast, but Plan A is more of a positive vision/wish list anyway.
In their own words:
That being said, Total Research Transparency is more restrictive than full open sourcing, as it does not require model weights to be open (and actively advises against it):
Mutually Assured Compute Destruction is necessary to punish defection and/or roll back capabilities if something goes horribly wrong. From AI 2040:
For more on how a Citizen's Dividend might work, consider The Windfall Clause (short talk, long paper).
This is by far the most speculative and exotic part of the scenario: the authors confine it to an epilogue section. I'm not terribly interested in it, largely because I think it’s far too early to predict this kind of thing, so I won't get into it here.
The authors expect America to maintain its current AI capabilities lead, leaving the rest of the world scrambling to accept a mutually beneficial deal. Quoting them:
It would thus be harder to establish a treaty if America lost its capabilities lead.
Another concern they mention is the risk of America not agreeing to Plan A:
However, because Plan A implicitly assumes that the U.S. is acting optimally in the best interests of itself and the world, the authors do not discuss this concern further.
I think this is a fair choice to make—it’s hard to write a positive vision for the future if you can’t assume that someone is going to make the right decisions—but for those of us in the real world, it does add an extra barrier to implementation (”even if I think Plan A sounds fine, how would we get and sustain the political will required to make all these drastic changes?”)
I keep repeating this phrase for a reason; if we had viable plans that did better than the status quo on any of these metrics, I'd link to them, and presumably the AIFP authors would too. But I'm unaware of any such plans, and so I have nothing to compare Plan A to.
From Scott Alexander's response post:
In the Plan A Assumptions supplement, author Thomas Larsen argues that the main downsides of building out data centers are:
He concludes that building out data centers is a good idea if and only if:
If any of those three premises are false, the argument becomes invalid.
For more on these concerns, particularly regarding the feasibility of Mutually Assured Compute Destruction, consider reading Plan A's problem with dry tinder by Tom Davidson.