One open question in AI risk strategy is: Can we trust the world's elite decision-makers (hereafter "elites") to navigate the creation of human-level AI (and beyond) just fine, without the kinds of special efforts that e.g. Bostrom and Yudkowsky think are needed?
Some reasons for concern include:
Otherwise smart people say unreasonable things about AI safety.
Many people who believed AI was around the corner didn't take safety very seriously.
Elites have failed to navigate many important issues wisely (2008 financial crisis, climate change, Iraq War, etc.), for a variety of reasons.
AI may arrive rather suddenly, leaving little time for preparation.
But if you were trying to argue for hope, you might argue along these lines (presented for the sake of argument; I don't actually endorse this argument):
If AI is preceded by visible signals, elites are likely to take safety measures. Effective measures were taken to address asteroid risk. Large resources are devoted to mitigating climate change risks. Personal and tribal selfishness align with AI risk-reduction in a way they may not align on climate change. Availability of information is increasing over time.
AI is likely to be preceded by visible signals. Conceptual insights often take years of incremental tweaking. In vision, speech, games, compression, robotics, and other fields, performance curves are mostly smooth. "Human-level performance at X" benchmarks influence perceptions and should be more exhaustive and come more rapidly as AI approaches. Recursive self-improvement capabilities could be charted, and are likely to be AI-complete. If AI succeeds, it will likely succeed for reasons comprehensible by the AI researchers of the time.
Therefore, safety measures will likely be taken.
If safety measures are taken, then elites will navigate the creation of AI just fine. Corporate and government leaders can use simple heuristics (e.g. Nobel prizes) to access the upper end of expert opinion. AI designs with easily tailored tendency to act may be the easiest to build. The use of early AIs to solve AI safety problems creates an attractor for "safe, powerful AI." Arms races not insurmountable.
The basic structure of this 'argument for hope' is due to Carl Shulman, though he doesn't necessarily endorse the details. (Also, it's just a rough argument, and as stated is not deductively valid.)
Personally, I am not very comforted by this argument because:
Elites often fail to take effective action despite plenty of warning.
I think there's a >10% chance AI will not be preceded by visible signals.
I think the elites' safety measures will likely be insufficient.
Obviously, there's a lot more for me to spell out here, and some of it may be unclear. The reason I'm posting these thoughts in such a rough state is so that MIRI can get some help on our research into this question.
In particular, I'd like to know:
Which historical events are analogous to AI risk in some important ways? Possibilities include: nuclear weapons, climate change, recombinant DNA, nanotechnology, chloroflourocarbons, asteroids, cyberterrorism, Spanish flu, the 2008 financial crisis, and large wars.
What are some good resources (e.g. books) for investigating the relevance of these analogies to AI risk (for the purposes of illuminating elites' likely response to AI risk)?
What are some good studies on elites' decision-making abilities in general?
Has the increasing availability of information in the past century noticeably improved elite decision-making?
Apart from the charred papyrus fragments recovered in Herculaneum [buried by the Vesuvius eruption that buried Pompeii], there are no surviving contemporary manuscripts from the ancient Greek and Roman world. Everything that has reached us is a copy, most often very far removed in time, place, and culture from the original. And these copies represent only a small portion of the works even of the most celebrated writers of antiquity. Of Aeschylus’ eighty or ninety plays and the roughly one hundred twenty by Sophocles, only seven each have survived; Euripides and Aristophanes did slightly better: eighteen of ninety-two plays by the former have come down to us; eleven of forty-three by the latter.
These are the great success stories. Virtually the entire output of many other writers, famous in antiquity, has disappeared without a trace. Scientists, historians, mathematicians, philosophers, and statesmen have left behind some of their achievements—the invention of trigonometry, for example, or the calculation of position by reference to latitude and longitude, or the rational analysis of political power—but their books are gone. The indefatigable scholar Didymus of Alexandria earned the nickname Bronze-Ass (literally, “Brazen-Bowelled”) for having what it took to write more than 3,500 books; apart from a few fragments, all have vanished. At the end of the fifth century ce an ambitious literary editor known as Stobaeus compiled an anthology of prose and poetry by the ancient world’s best authors: out of 1,430 quotations, 1,115 are from works that are now lost.
It is almost impossible for jokes that are centuries old to retain any life. The fact that a few of the jokes of Shakespeare or Rabelais or Cervantes continue to make us smile is something of a miracle. Almost six hundred years old, Poggio’s Facetiae is by now largely interesting only as a symptom. These relics, like the remains of long-dead insects, tell us what once buzzed about in the air of the Vatican. Some of the jokes are professional complaints, of the sort secretaries must always have had: the boss routinely claims to
One open question in AI risk strategy is: Can we trust the world's elite decision-makers (hereafter "elites") to navigate the creation of human-level AI (and beyond) just fine, without the kinds of special efforts that e.g. Bostrom and Yudkowsky think are needed?
Some reasons for concern include:
But if you were trying to argue for hope, you might argue along these lines (presented for the sake of argument; I don't actually endorse this argument):
The basic structure of this 'argument for hope' is due to Carl Shulman, though he doesn't necessarily endorse the details. (Also, it's just a rough argument, and as stated is not deductively valid.)
Personally, I am not very comforted by this argument because:
Obviously, there's a lot more for me to spell out here, and some of it may be unclear. The reason I'm posting these thoughts in such a rough state is so that MIRI can get some help on our research into this question.
In particular, I'd like to know: