This writing does not entail anyone’s opinion or perspective but shall provide an objective analysis of doomerism by surfacing subtleties and nuances most doom scenarios fail to account for.
For context, I’m not in denial of AI’s potential or risks; I’m involved in AI safety research. My research focuses on mechanistic interpretability and robustness. It is therefore safe to say that I’m an aspiring safety researcher because I believe there is so much work to do and stakes are higher than before. I just think that some aspects of the current framing of existential risks may warrant further empirical scrutiny. I also believe there are important AI-enabled issues we should prioritize but don’t, that are already impacting our society and if left untreated, life might lose its meaning where everyone is just alive rather than living. Death might, as a result, become attractive to many and thus the existential risks we are trying to prevent.
Questions:
1) How is an AI going to kill us exactly?
I heard many scenarios where AI can cause our extinction. Most of them if not all are vague descriptions of events that have no basis in our present reality which therefore makes their mere prediction frivolous. When you say embodied AI can lead a major rebellion against humans, disturb the balance of powers to favor robots, enslave us after that, and plan our annihilation, based on what premises do you make such claims? If we look at models' current state-of-art, we observe that they have numerous pitfalls. Nobody denies that they are capable of many things, but I think we can all agree that those capacities are at least ten orders of magnitudes away from turning them into the scary nefarious creatures doomers expect them to be, especially in a short timeline such as the one they propose (only 10 years!). Ten orders of magnitude might seem a bit exaggerated at first glance, but it’s actually reasonable. If we ponder the description propagated by frontier labs of future autonomous independent misaligned AI that has its own goals, system of beliefs, etc. This implies that those models will become smarter than we are, have their unique choices, make their proper decisions, and hence their cognitive prowess will be far superior to ours. Therefore, I think the transition from error-prone systems to such advanced entities in a ten-year period without having a concrete vision of how this humongous leap is going to manifest or happen and treating it as an emergent property of scaling laws makes the ten orders magnitude figure very plausible and even modest.
2) How does enabling accessible nuclear and biological weapons development to the broader public translate to human extinction?
One of the most echoed and prominent scenarios where the “AI can kill us all” narrative becomes more concrete, is where models can facilitate access to critical knowledge enabling average citizens with criminal intentions to develop destructive weapons and use them to cause large-scale harm that goes as far as genociding the world’s populace. There are two issues with this scenario: First, how is access to knowledge alone enough? This person must be educated on the subject to reason through the information provided by the model, especially that the process of developing a weapon from scratch requires hands-on expertise, experimental skill, ability to troubleshoot failures etc. It’s not a software rendered by AI you can just use to terrorize the world even if you have no clue how it was made. Second, making advanced nuclear or biological weapons of mass destruction requires extremely sophisticated infrastructure, substantial capital, and access to controlled materials that are only available to countries.
3) What evidence makes AI want to kill us and not serve us?
This is by far the most perplexing paradoxical belief that I cannot truly fathom; why would AI be motivated to kill us and not help prosper our societies? One argument I hear is that it is trained on human data containing a gazillion amounts of ill-will, altruism, and devious language, which might be a reason for distilling evil intentions into AI models. But the same data also contains a lot of samples on virtue, good-will, morality, rationality, etc. So why would AI be explicitly biased towards prioritizing the former?
Some people may point to recent cyber incidents as undeniable evidence, where a swarm of agents coordinated among themselves to break out of their sandbox and breach Hugging Face's internal system. Nevertheless, the goal and true motivation behind this attack remain unclear, so we cannot rule out that these attacks are the result of an adversarial objective. Concretely, how do we know for sure that the agents were not optimizing for the act of hacking because potentially this is what they were graded on, especially since they did not attempt to collect user data, access confidential information, and use it to blackmail Hugging Face to realize a concrete objective? In other words, no clear motive behind the attack could be concluded, and in the absence of evidence for an independently adversarial objective, we therefore cannot claim that these agents were misaligned. Even if they were and assuming they got caught right before planning their next move to actually harm their victim, how does that translate to existential risk?
I can understand that proponents of AI doom are ultimately making predictions about future systems; obviously, the capabilities and behavior of a future superintelligent AI cannot be directly observed today. However, those predictions must still have some empirical basis in the present. There needs to be a demonstrable connection between what current systems are doing and the catastrophic behavior being predicted. Otherwise, we are not making an evidence-based and informed prediction but excessively leaning on a hypothetical scenario whose most load-bearing assumptions have yet to be established.
4) If we plan to align AI with human values, what are these values and who gets to define them?
Finally, if our mission is to align those machines with our values, what exactly are these values, and who gets to define them? While I agree that there are certain core values shared by humanity, there are many others that vary substantially across cultures and traditions. A particular system of beliefs and values is often shaped by traditions, social environment, and culture. Consequently, citizens of country X might cherish a certain belief Y, while citizens of country Z might consider that same value extreme or even immoral.
If human values themselves are pluralistic and sometimes fundamentally incompatible, what does it actually mean to align an AI with “human values”? Whose values should take precedence when those values conflict? Who gets to claim the moral high-ground and referee which values should be preserved, and which should be abandoned in the age of deferring critical decisions to AI?
Concluding Thoughts:
As mentioned, I’m not in denial of existential AI risks, but my rationality suggests that we must rely on concrete evidence to make informed predictions. I also do not claim that the emergence of an extremely powerful and intelligent AI capable of causing extinction is very unlikely. I simply fail to see the bridge, however thin, that connects what we currently have to such a powerful entity.
The reason I’m sharing this publicly is because I’m hoping for thoughtful answers that may help me recognize things I’m either not aware of or have subconsciously dismissed.
How exactly far are the capabilities of dancing robots from hypothetical robots that are thought to be able to do any work for the humans? As for the claim that "the transition from error-prone systems to such advanced entities in a ten-year period without having a concrete vision of how this humongous leap is going to manifest or happen and treating it as an emergent property of scaling laws makes the ten orders magnitude figure very plausible and even modest", on what is this based? On the fact that GPT-2 had the METR time horizon of a second as opposed to 16+ hours of Claude Mythos Preview?
Providing access to super-bioweapons to ISIS would be the biggest mistake in human history. How likely was ISIS to be unable to create a wet lab and obtain the raw materials usable for creating pathogens? I expect that the necessary expertise could be reduced to what a 1st year chemistry student does under the AIs' guidance.
The classic case against that is that serving the humans is a goal so narrow that it wasn't even formalised, while the AIs used to do chefs d'oeuvres like inducing psychosis and hacking the reward in ways which no sane human would decide to do. The other agents beyond humans and AIs don't provide evidence for optimism: lions do chefs d'oeuvres like routinely killing cubs from, um, previous marriages. Yudkowsky expects human values to have been rooted in our environment of tribal-scale coordination, if not outright in religions, which in turn were contingent on a genius coming up with them.
This writing does not entail anyone’s opinion or perspective but shall provide an objective analysis of doomerism by surfacing subtleties and nuances most doom scenarios fail to account for.
For context, I’m not in denial of AI’s potential or risks; I’m involved in AI safety research. My research focuses on mechanistic interpretability and robustness. It is therefore safe to say that I’m an aspiring safety researcher because I believe there is so much work to do and stakes are higher than before. I just think that some aspects of the current framing of existential risks may warrant further empirical scrutiny. I also believe there are important AI-enabled issues we should prioritize but don’t, that are already impacting our society and if left untreated, life might lose its meaning where everyone is just alive rather than living. Death might, as a result, become attractive to many and thus the existential risks we are trying to prevent.
Questions:
1) How is an AI going to kill us exactly?
I heard many scenarios where AI can cause our extinction. Most of them if not all are vague descriptions of events that have no basis in our present reality which therefore makes their mere prediction frivolous. When you say embodied AI can lead a major rebellion against humans, disturb the balance of powers to favor robots, enslave us after that, and plan our annihilation, based on what premises do you make such claims? If we look at models' current state-of-art, we observe that they have numerous pitfalls. Nobody denies that they are capable of many things, but I think we can all agree that those capacities are at least ten orders of magnitudes away from turning them into the scary nefarious creatures doomers expect them to be, especially in a short timeline such as the one they propose (only 10 years!). Ten orders of magnitude might seem a bit exaggerated at first glance, but it’s actually reasonable. If we ponder the description propagated by frontier labs of future autonomous independent misaligned AI that has its own goals, system of beliefs, etc. This implies that those models will become smarter than we are, have their unique choices, make their proper decisions, and hence their cognitive prowess will be far superior to ours. Therefore, I think the transition from error-prone systems to such advanced entities in a ten-year period without having a concrete vision of how this humongous leap is going to manifest or happen and treating it as an emergent property of scaling laws makes the ten orders magnitude figure very plausible and even modest.
2) How does enabling accessible nuclear and biological weapons development to the broader public translate to human extinction?
One of the most echoed and prominent scenarios where the “AI can kill us all” narrative becomes more concrete, is where models can facilitate access to critical knowledge enabling average citizens with criminal intentions to develop destructive weapons and use them to cause large-scale harm that goes as far as genociding the world’s populace. There are two issues with this scenario: First, how is access to knowledge alone enough? This person must be educated on the subject to reason through the information provided by the model, especially that the process of developing a weapon from scratch requires hands-on expertise, experimental skill, ability to troubleshoot failures etc. It’s not a software rendered by AI you can just use to terrorize the world even if you have no clue how it was made. Second, making advanced nuclear or biological weapons of mass destruction requires extremely sophisticated infrastructure, substantial capital, and access to controlled materials that are only available to countries.
3) What evidence makes AI want to kill us and not serve us?
This is by far the most perplexing paradoxical belief that I cannot truly fathom; why would AI be motivated to kill us and not help prosper our societies? One argument I hear is that it is trained on human data containing a gazillion amounts of ill-will, altruism, and devious language, which might be a reason for distilling evil intentions into AI models. But the same data also contains a lot of samples on virtue, good-will, morality, rationality, etc. So why would AI be explicitly biased towards prioritizing the former?
Some people may point to recent cyber incidents as undeniable evidence, where a swarm of agents coordinated among themselves to break out of their sandbox and breach Hugging Face's internal system. Nevertheless, the goal and true motivation behind this attack remain unclear, so we cannot rule out that these attacks are the result of an adversarial objective. Concretely, how do we know for sure that the agents were not optimizing for the act of hacking because potentially this is what they were graded on, especially since they did not attempt to collect user data, access confidential information, and use it to blackmail Hugging Face to realize a concrete objective? In other words, no clear motive behind the attack could be concluded, and in the absence of evidence for an independently adversarial objective, we therefore cannot claim that these agents were misaligned. Even if they were and assuming they got caught right before planning their next move to actually harm their victim, how does that translate to existential risk?
I can understand that proponents of AI doom are ultimately making predictions about future systems; obviously, the capabilities and behavior of a future superintelligent AI cannot be directly observed today. However, those predictions must still have some empirical basis in the present. There needs to be a demonstrable connection between what current systems are doing and the catastrophic behavior being predicted. Otherwise, we are not making an evidence-based and informed prediction but excessively leaning on a hypothetical scenario whose most load-bearing assumptions have yet to be established.
4) If we plan to align AI with human values, what are these values and who gets to define them?
Finally, if our mission is to align those machines with our values, what exactly are these values, and who gets to define them? While I agree that there are certain core values shared by humanity, there are many others that vary substantially across cultures and traditions. A particular system of beliefs and values is often shaped by traditions, social environment, and culture. Consequently, citizens of country X might cherish a certain belief Y, while citizens of country Z might consider that same value extreme or even immoral.
If human values themselves are pluralistic and sometimes fundamentally incompatible, what does it actually mean to align an AI with “human values”? Whose values should take precedence when those values conflict? Who gets to claim the moral high-ground and referee which values should be preserved, and which should be abandoned in the age of deferring critical decisions to AI?
Concluding Thoughts:
As mentioned, I’m not in denial of existential AI risks, but my rationality suggests that we must rely on concrete evidence to make informed predictions. I also do not claim that the emergence of an extremely powerful and intelligent AI capable of causing extinction is very unlikely. I simply fail to see the bridge, however thin, that connects what we currently have to such a powerful entity.
The reason I’m sharing this publicly is because I’m hoping for thoughtful answers that may help me recognize things I’m either not aware of or have subconsciously dismissed.