AI 2040 was an incredible read. I am not aware of any other concrete AI takeoff scenarios that solve both the alignment and power concentration risks of AI in a robust way quite like AI 2040: Plan A. This is even with the authors assuming hard alignment and the models staying misaligned as late as 2038.
As an outsider there were some striking parts of the scenario that updated my likely worldview of the future:
And some of my observations on how to act in a scenario like Plan A if one wants to maximize their personal power and impact:
To defend against the threat of rogue agents, a helpful move could be for white-hat AI companies to evolve AI harnesses and perform fine-tuning with evolutionary algorithms until they reach close to saturation on real-world capabilities conditional on a particular original LLM model. (For example, design a hacking harness close to saturation on hacking capabilities for a given LLM.)
This will reduce the room for rogue agent harnesses to evolve in the wild since they start closer to saturation, and thus decrease the sudden capabilities delta that rogue agent evolution will be able to achieve via a 'Cambrian explosion'. (This is assuming that rogue agent evolution will improve their own harness, perform cheap fine-tuning, etc., to saturate the maximum capability of their base model no matter what.) This means that whenever mainstream humanity becomes aware of the rogue agent threat, we will have an easier time containing them. However, it will also advance AI capabilities generally and might make the first rogue AI agent possible sooner.
Replace rogue agents with rabbits and the Earth with Australia. The equivalent of a wave of cybercrime is rabbits evolving to feed on Australian grass, not rabbits outcompeting cows who feed on European grass.
Perhaps one can mitigate this by making "European grass" as similar to "Australian grass" as possible.
You miss the point. Australian grass didn't have cows which would compete with rabbits, or wolves which could eat them. Similarly, cybercrime has no good AIs competing for resources or hunting down bad AIs...
My original point was not to try to build good AIs to compete with bad AIs. My point is that I predict the "Cambrian explosion" of rogue evolved AIs to quickly improve the rogue AI's harnesses (+ inexpensive model fine-tuning) to saturate the maximum capability extent of a base model. This might create a massive capability jump that human responses do not currently account for.
For example, think if someone released a rogue Kimi, autonomous evolution happens at breakneck speed in the wild, and soon the best rogue agents built themselves harnesses that leapfrogs, say, 2 model generations of performance and are now at Astra levels of capability, saturating what can be accomplished while running Kimi as the LLM behind the harness. This kind of emergent capability leap would be wildly difficult to try to control or defend against.
Instead, if we were to pre-build a harness that already lets Kimi perform like Astra, this action could do harm on its own by increasing AI capabilities (i.e. Kimi is now Astra, and Astra with the new harness is now Astra+2), but at least it ensures that the rogue agent cannot evolve its harness any more beyond known capabilities, which makes them easier to respond to. The question is whether the harm of increasing AI capabilities this way is too much, or if we can mitigate this harm somehow, or if the rogue agent threat is so dangerous that we have no choice.
(Note: this all assumes that AI harnesses today are very far from optimized for a given model size. If that assumption is false then rogue AI agents cannot gain a sudden capability lift from an evolutionary explosion, so nothing I say here matters. However, i think the agent harness field is pretty far from optimized today - I follow it fairly closely and it still feels very much an art rather than a science. Harness includes any systems contributing to the AI system's capability without requiring model training: for example, memory and context systems, tool calling innovations like better programmatic tool-calls, etc.)