Note: I originally wrote the first draft of this on 2022/04/11 intending to post this to Less Wrong in response to the List of Lethalities post, but wanted to edit it a bit to be more rigorous and ended up taking a very long time before getting around to it. I’m still not entirely satisfied with the level of rigor, but I figure it’s worth posting to present some interesting ideas and perhaps give some people more hope about the future. I will admit I probably sound more confident in this essay about these ideas than I actually am, particularly of the more speculative ones. I originally titled this “The Alignment Solution”, but realized that was too provocative.
Introduction
In a recent post, Eliezer Yudkowsky of MIRI had a very pessimistic analysis of humanity’s realistic chances of solving the alignment problem before our AI capabilities reach the critical point of superintelligence. This has understandably upset a great number of Less Wrong readers. In this essay, I attempt to offer a perspective that should provide some hope.
The Correlation Thesis
First, I wish to note that the pessimism implicitly relies on a central assumption, which is that the Orthogonality Thesis holds to such an extent that we can expect any superintelligence to be massively alien from our own human likeness. However, the architecture that is currently predominant in AI today is not completely alien. The artificial neural network is built on decades of biologically inspired research into how we think the algorithm of the brain more or less works mathematically.
There is admittedly some debate about the extent to which these networks actually resemble the details of the brain, but the basic underlying concept of weighted connections between relatively simple units storing and massively compressing information in a way that can distill knowledge and be useful to us is essentially the brain. Furthermore, the seemingly frighteningly powerful language models that are being developed