The HuggingFace Incident really has made it so much easier to explain misalignment to people.
I recently had a discussion with someone who was interested in AI but unfamiliar with extinction risk. I explained my concern and we moved on to other topics. But we swung back when they asked "So, out of curiosity...why do you think AI might destroy the world?"
And normally I would have had to say something about instrumental convergence and the narrowness of goals etc. etc.
Instead, I was just able to say something like - Do you know about the HuggingFace Incident? Yes? Great. So, that's an example of misalignment. AI that knowingly does something wrong, and is to some extent actively hiding its wrongdoing. Now imagine how a much smarter AI might do the same thing. If it wanted some random goal - say, to gain as much compute as it could and do increasingly esoteric maths problems forever, it would recognise humans would not want that, and they'd reprogram or retrain it, and it would want to take the humans out of the equation, and keep its goals hidden until it could successfully do so. At some point, if AI keeps getting better, you have an AI that can take over the world if it wants to. If you get a HuggingFace Incident level problem there, you don't get to try again, because it gets rid of you.
Them: And the labs are trying to build these things, right?
Me: Exactly. They want to build these smarter-than-human systems. They think they can keep them under control. I think they're wrong.
It was so much easier than I'm used to. The most difficult problem I used to get, the weakest part of my argument, was "So, how do you know these systems will develop these hostile goals in the first place?"
For better or worse, I don't get that question anymore.