One of the most frequently suggested ways AI could wipe out humanity is via an engineered virus.
I have no doubt that a sufficiently intelligent AI given sufficient time could socially-or-otherwise engineer his way to have a whole team of people working on developing super-viruses to wipe out humanity, and could also kill us in any one of a dozen other ways.
However much more important for very short term risks is whether a slightly more intelligent LLM, with spiky abilities, could do such a thing.
The argument against is that designing complex viruses requires iteration. It's insufficient just to trick a single lab into making a virus for you once, you need to have a lab where you can repeatedly iterate till you get it right. That's a far higher bar to clear, and requires a correspondingly more advanced AI, giving us more time to get alignment right.
Our earliest evidence for this will be AIs abilities in materials and biological sciences. Achievements here are in some ways similar to the Navier Stokes counterexample, in that they just require finding a single example that fits a bunch of complex criteria to achieve specific desired results. However, unlike NS, testing them requires significant interaction with the physical world and complex equipment.
If frontier LLMs are able to one-shot room temperature superconductors, or cures for cancer, we should be extremely worried about what they do with a virus. If they cannot[1], that suggests that we can focus less on "AI emails a DNA synthesis company the plans for a supervirus" as a threat model, and focus instead on other threat models, and ensuring they never get long term access to a lab[2].
I am a materials scientist by training and while I haven’t spent much time specifically on superconductivity, my impression is that the existing materials such as hydrides under extreme pressures are already pushing what is physically possible. While a discovery of such a material would undoubtedly be massive, my current best estimate is that room-temperature superconductors at ambient pressure probably (70%) don’t exist in our universe. In any case, this seems like a significantly harder problem to me than designing viruses; ab initio approaches struggle to accurately predict high-temperature superconductivity, so AI would have to first make significant theoretical progress.
Cancer cures, on the other hand, I don’t see any reason for not being feasible, but I’m out of my depth there.
One of the most frequently suggested ways AI could wipe out humanity is via an engineered virus.
I have no doubt that a sufficiently intelligent AI given sufficient time could socially-or-otherwise engineer his way to have a whole team of people working on developing super-viruses to wipe out humanity, and could also kill us in any one of a dozen other ways.
However much more important for very short term risks is whether a slightly more intelligent LLM, with spiky abilities, could do such a thing.
The argument against is that designing complex viruses requires iteration. It's insufficient just to trick a single lab into making a virus for you once, you need to have a lab where you can repeatedly iterate till you get it right. That's a far higher bar to clear, and requires a correspondingly more advanced AI, giving us more time to get alignment right.
Our earliest evidence for this will be AIs abilities in materials and biological sciences. Achievements here are in some ways similar to the Navier Stokes counterexample, in that they just require finding a single example that fits a bunch of complex criteria to achieve specific desired results. However, unlike NS, testing them requires significant interaction with the physical world and complex equipment.
If frontier LLMs are able to one-shot room temperature superconductors, or cures for cancer, we should be extremely worried about what they do with a virus. If they cannot[1], that suggests that we can focus less on "AI emails a DNA synthesis company the plans for a supervirus" as a threat model, and focus instead on other threat models, and ensuring they never get long term access to a lab[2].
Either because they're simply not capable of it, or it requires significant iteration.
For example - make it illegal to have an autonomously operated wet lab.