This is a special post for quick takes by Jesse Li. Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page.
A crazy idea, which to be honest I haven't fully thought through - would it be possible and net beneficial to intentionally release an AI model (~Mythos capability) whose primary objective was to infect data centers, self replicate, and cause controlled disruption (a la ransomware), solely as an artificial warning shot for runaway AGI?
Benefits (assuming widespread disruption is caused, e.g. large chunks of the internet goes down):
Raises public awareness of capabilities and danger of AGI
Tests how and how quickly society responds (hopefully with lessons learned for if/when this happens unintentionally)
We're still at a stage where even if this happened, uncontrolled runaway can be stopped because within 8~24 hours, humans can physically travel to all the data centers and pull out all the power plug
Costs:
Obviously taking the internet down is bad
Restarting can only be done slowly to avoid recontamination - it'd be the Morris worm on a much, much larger scale
Leaks the weights to the public
If the released model figures out how to mutate and self-improve, it could lead to uncontrolled ASI (my guess is that current models aren't capable enough to do this, and maybe some countermeasures (checksumming?) can be applied? But I'm not up-to-date on all of the frontier model capabilities)
Similarly, if the model replicates to personal computers (possibly distributed storage of weights is needed?), it would be much harder to clean from the internet
Huge personal cost for whoever actually does this, if they get caught
[Aside - I'm a new contributor. The guidelines mention a high bar with respect to AI content and suggest first posting to the AI Questions Open Thread, but is that still active and correct advice? The last open thread seems to be 3 years ago, but it seems like they're supposed to happen monthly?]