One of the most frequently suggested ways AI could wipe out humanity is via an engineered virus.
I have no doubt that a sufficiently intelligent AI given sufficient time could socially-or-otherwise engineer his way to have a whole team of people working on developing super-viruses to wipe out humanity, and could also kill us in any one of a dozen other ways.
However much more important for very short term risks is whether a slightly more intelligent LLM, with spiky abilities, could do such a thing.
The argument against is that designing complex viruses requires iteration. It's insufficient just to trick a single lab into making a virus for you once, you need to have a lab where you can repeatedly iterate till you get it right. That's a far higher bar to clear, and requires a correspondingly more advanced AI, giving us more time to get alignment right.
Our earliest evidence for this will be AIs abilities in materials and biological sciences. Achievements here are in some ways similar to the Navier Stokes counterexample, in that they just require finding a single example that fits a bunch of complex criteria to achieve specific desired results. However, unlike NS, testing them requires significant interaction with the physical world and complex equipment.
If frontier LLMs are able to one-shot room temperature superconductors, or cures for cancer, we should be extremely worried about what they do with a virus. If they cannot[1], that suggests that we can focus less on "AI emails a DNA synthesis company the plans for a supervirus" as a threat model, and focus instead on other threat models, and ensuring they never get long term access to a lab[2].
One of the most frequently suggested ways AI could wipe out humanity is via an engineered virus.
I have no doubt that a sufficiently intelligent AI given sufficient time could socially-or-otherwise engineer his way to have a whole team of people working on developing super-viruses to wipe out humanity, and could also kill us in any one of a dozen other ways.
However much more important for very short term risks is whether a slightly more intelligent LLM, with spiky abilities, could do such a thing.
The argument against is that designing complex viruses requires iteration. It's insufficient just to trick a single lab into making a virus for you once, you need to have a lab where you can repeatedly iterate till you get it right. That's a far higher bar to clear, and requires a correspondingly more advanced AI, giving us more time to get alignment right.
Our earliest evidence for this will be AIs abilities in materials and biological sciences. Achievements here are in some ways similar to the Navier Stokes counterexample, in that they just require finding a single example that fits a bunch of complex criteria to achieve specific desired results. However, unlike NS, testing them requires significant interaction with the physical world and complex equipment.
If frontier LLMs are able to one-shot room temperature superconductors, or cures for cancer, we should be extremely worried about what they do with a virus. If they cannot[1], that suggests that we can focus less on "AI emails a DNA synthesis company the plans for a supervirus" as a threat model, and focus instead on other threat models, and ensuring they never get long term access to a lab[2].
Either because they're simply not capable of it, or it requires significant iteration.
For example - make it illegal to have an autonomously operated wet lab.