I worry that too much of the AI safety world's (at least the portion I interact with) mental bandwidth is being consumed by the threat of an eventual AI takeover. This is not an argument against dedicating substantial resources to alignment or takeover risk. Rather, I worry that we are overweighting takeover relative to other existential risks posed by capable AI systems.
For example, this post argues that autonomous replicating agents are not themselves an existential risk, and analyzes them largely through the lens of whether they accelerate the race toward an ASI capable of takeover. I think that analysis is useful, but I worry about treating AI takeover as the primary yardstick by which other AI risks are judged.
Taking the Hugging Face incident as an example it seems that models trained using RL are very good at finding the weakest link in a chain to achieve their particular goal. The important question may not be only, "Can this model take over?" but also, "What is the weakest consequential system this model can break?" The question still remains as to what the current weakest link is and the risk inherent in breaking that link. Given the proliferation of nuclear weapons in the world, what link in the chain would need to be broken for an existing agent to begin risking chances of nuclear war? What software bugs related to alerting or decisions relating to deployment of ordinance exist? Is there a leader with a disproportionate amount of power who could be influenced with a well-timed post or email? If you're not worried about this in the US, what about every other world power with the ability to effectively wage war? The same general concern applies beyond nuclear weapons. AI does not necessarily need to become sovereign to make those systems substantially more dangerous.
I worry that we are missing the forest for the one very big tree at the moment. Perhaps rather than AI takeover consuming 98% of governance thought, we need to reevaluate and place it closer to 85% (or your preferred number) to ensure that these short term risks are not underrepresented. Takeover is a very salient idea, and one that my mind is very primed to jump to as a potential future risk, but I worry that there is more risk than we are currently assigning to comparatively boring failure modes.