Has anyone written an essay about how to fight against/correct for Trapped Priors? I would like to do something like that, but I want to make sure that I’m not reinventing the wheel here. Thank you!
I keep running into conceptual confusion around the term "alignment," particularly when reading older Less Wrong posts. Some people say "aligned AI" and mean "an AI that works for human flourishing," some people say that an AI "is aligned" if it reliably advances the intended objectives of some person or group (and doesn't have some secret set of goals / isn't scheming), and yet other people use "alignment" to mean something along the lines of "the ability of any system to reliably work towards some pre-defined goal." I usually have to work out which is being said on the spot, which is annoying given that the implications of each are very different.
Is there one commonly accepted definition? Is this confusion just a thing we've all accepted?
As Raemon put it,
An aligned AI is the one who is successfully pointed by humans to a goal. If mankind does solve alignment, then a power struggle over which goals the AI serves may have an effect on the world. Otherwise the AI pursues the goals which mankind never set, and the humans are wiped out or disempowered.
Gotcha. Is there a strong reason to assume that we'll succeed at creating AIs that can be pointed at a single target? I read this post and comment a while back and would love your thoughts.
The Hugging Face breach is probably not a clear warning shot. It might spur policymakers into action, but it seems like there are still a few mitigating factors preventing it from taking off—for example, some people who I respect aren't taking it very seriously yet due to their strong distrust of OpenAI and Sam Altman. And since no one was directly harmed, we're still left saying "What happens if capabilities increase further??" instead of "This is what happens if we don't intervene right now," which is obviously the stronger message.
In some ways, this is good: if we mobilize now, we'll probably do so even more if we get an indisputably clear warning shot (e.g. an AI commits some act of terrorism). On the other hand, I do hope we're not capitalizing on our goodwill early; there are some people out there who are determined to paint AI Safety people as perennial wolf-criers, and I worry that they'll say the same about us this time.
Thought in progress: epistemic humility is not a substitute for actual humility (or professed humility). You only get to cry wolf once, but you can probably warn about potential wolves several times—so long as you don't burn goodwill on an incorrect or overconfident prediction.
I think epistemic humility helps to increase trust and confidence in EA/Less Wrong-type spaces, but I think professed humility is far more helpful when it comes to public-facing AI comms, particularly as scenarios get more intense and specific (e.g. prefacing AI doom predictions with a decent amount of throat-clearing beforehand commensurate with the intensity and specificity of the forecast). For example, I think that AI 2027 might have been better received if the authors had spent less time trying to convince the readers of their credibility at the beginning and spent more time saying something along the lines of "we know this sounds crazy and are well aware of how sci-fi the scenario seems". (I'm not a huge fan of lampshading in fiction, but IRL, I think you do need to display self-awareness of outlandishness in order to be taken seriously, particularly if what you're predicting sounds insane to the average person.)
Of course, there are huge diminishing returns on this: the more throat-clearing you do, the less confident you seem. And throat-clearing should probably be saved for public-facing comms, because actual technical work seems to require people who are confident in their beliefs even when they are outlandish (as proven by the outlandish explosion of AI progress recently).
Still, I think that the AI safety community at large has a worse reputation than they deserve, and I think part of that is due to the appearance of overconfidence. This problem seems simple, tractable, and important.
Does the Fermi Paradox put an upper limit on the bounds of ASI capabilities?
Given that the Universe has had ~14 billion years to develop, it seems overwhelmingly likely that someone else out there has already maxxed out the tech tree and pushed AI as far as it can go. But we don't see any Von Neumann probes eating the Milky Way, nor do we see any evidence of interstellar travel within the Virgo Supercluster (~147 million light years)...let alone the rest of the observable universe!
From this observation, we can conclude one of two things. Either:
I concede that 2. is possible but I (perhaps naively) think 1. is far more likely. I would even go so far as to say that normal / sub-FTL interstellar travel is probably impossible on this view.
Of course, there are always the other standard possibilities e.g.
Either way, I think we shouldn't take the fact that we haven't been consumed or conquered by some alien civilization or ASI (yet) for granted.
I have a suspicion that p-zombie discourse is only going to get more relevant as LLMs get better. No one really argues that animals aren't conscious, even though they can't use words very well, but the release of GPT-3 caused a steady rise in people arguing that AIs are conscious. It's not clear to me that an LLM couldn't possibly be conscious, but it does seem that many people are taking LLM eloquence to imply that they are conscious, and I'm pretty sure we've been discussing this for years...