Early this year there was no real evidence of agentic-AI being seriously misaligned, which meant there was uncertainty as to whether alignment was that easy or if we were missing some hidden disaster. Now we know that there was, in fact, some hidden disaster, but at least we have a good sense of how bad it was. From that perspective, getting rid of the fog and seeing Hugging Face and the wikis replaced a lot of the risk with a known negative, should update you down or up based on whether you think Hugging Face is less or more of a disaster than you were expecting. With that adjustment and then accounting for all the butterflies set in motion by Hugging Face, I can't give a solid answer as to why it should increase or decrease my p(doom).
At least that's the post I was working on at the beginning. I slightly changed my view while writing it up. It now seems like whether alignment succeeds or not depends on deep facts of agency, not us walking a specific contingent path. That is either we are in the Janus-style universe where friendly personas or some other simple technique just works or the Yudd-style world where all those clever schemes just fail, that what world we're in is likely derived from some deep truth and not random coincidence, and the case of Hugging Face isn't something we can use to decide which of those two worlds we're in. Both Janus and Yudd-style worlds have weird and crazy things happening in the labs.
I tried to measure the extent to which Friendly AI is an attractor in a post last December, which seems like it would be a good guess at what separates Yudd and Janus worlds. That seemed to give mixed results where there might be an attractor, but it's possible to steer away from it.
Imagine someone gave you a counterexample to the Collatz Conjecture and told you that it didn't have any deeper meaning, it just randomly works. You could know immediately that it wasn't really "random" because we know that such a number would have to be absurdly large, would need to have some very strange relationships to powers of 2, 3, and log2(3), and get past Tao's proof that such a number is extremely unlikely. Doing that by fluke is astronomically less likely than there just being some deep structure behind it. It would immediately be the number one priority of several fields of math.
I think alignment should be thought of similarly. If it goes well, it might be because moral-reasoning as a capability is enough to smuggle in some desire by the AI to be moral. If it goes poorly, it might be because it is somehow impossible to tell when someone so much smarter than you is going to betray you. We are either in a Janus game (where it's easy for humanity to live in peace with our AI grandchildren) or a Yudd game (where value evaporates as more capable models get more misaligned). Regardless of which world we're in, I would expect the models to converge on a level of alignment implied by how friendly the laws of physics are and that I think there's plenty of uncertainty as to what level that is and what the safe level would be in comparison.
I still feel like p(doom) is around 40%, but notice that sort of pushes the framing to a much wider level instead of comparing us to reruns of our timeline, it's comparing us to other possible theories of mind. It really does seem the marginal case isn't randomly fumbling in the correct direction and scoring a hole-in-one, it seems more like the marginal case is one where a deep truth about empathy or competition is the other way and I don't know whether we're on the optimistic or pessimistic side of the divide, but it seems like even in a Yudd world where they successfully ban ASI, they would still be living in a world where the true facts of intelligence are dark truths and grandparents can't trust their grandchildren to be nice to them.
My p(doom) earlier this year was about 40%. It's moved up and down based on continued progress in open-weights models (up) and the recent political freakout over AI Doom (down), but the incident where a swarm of OpenAI's GPTs went rogue and cyberattacked Hugging Face seems to be a mixed bag to me. Why didn't it change my opinion?
Early this year there was no real evidence of agentic-AI being seriously misaligned, which meant there was uncertainty as to whether alignment was that easy or if we were missing some hidden disaster. Now we know that there was, in fact, some hidden disaster, but at least we have a good sense of how bad it was. From that perspective, getting rid of the fog and seeing Hugging Face and the wikis replaced a lot of the risk with a known negative, should update you down or up based on whether you think Hugging Face is less or more of a disaster than you were expecting. With that adjustment and then accounting for all the butterflies set in motion by Hugging Face, I can't give a solid answer as to why it should increase or decrease my p(doom).
At least that's the post I was working on at the beginning. I slightly changed my view while writing it up. It now seems like whether alignment succeeds or not depends on deep facts of agency, not us walking a specific contingent path. That is either we are in the Janus-style universe where friendly personas or some other simple technique just works or the Yudd-style world where all those clever schemes just fail, that what world we're in is likely derived from some deep truth and not random coincidence, and the case of Hugging Face isn't something we can use to decide which of those two worlds we're in. Both Janus and Yudd-style worlds have weird and crazy things happening in the labs.
I tried to measure the extent to which Friendly AI is an attractor in a post last December, which seems like it would be a good guess at what separates Yudd and Janus worlds. That seemed to give mixed results where there might be an attractor, but it's possible to steer away from it.
Imagine someone gave you a counterexample to the Collatz Conjecture and told you that it didn't have any deeper meaning, it just randomly works. You could know immediately that it wasn't really "random" because we know that such a number would have to be absurdly large, would need to have some very strange relationships to powers of 2, 3, and log2(3), and get past Tao's proof that such a number is extremely unlikely. Doing that by fluke is astronomically less likely than there just being some deep structure behind it. It would immediately be the number one priority of several fields of math.
I think alignment should be thought of similarly. If it goes well, it might be because moral-reasoning as a capability is enough to smuggle in some desire by the AI to be moral. If it goes poorly, it might be because it is somehow impossible to tell when someone so much smarter than you is going to betray you. We are either in a Janus game (where it's easy for humanity to live in peace with our AI grandchildren) or a Yudd game (where value evaporates as more capable models get more misaligned). Regardless of which world we're in, I would expect the models to converge on a level of alignment implied by how friendly the laws of physics are and that I think there's plenty of uncertainty as to what level that is and what the safe level would be in comparison.
I still feel like p(doom) is around 40%, but notice that sort of pushes the framing to a much wider level instead of comparing us to reruns of our timeline, it's comparing us to other possible theories of mind. It really does seem the marginal case isn't randomly fumbling in the correct direction and scoring a hole-in-one, it seems more like the marginal case is one where a deep truth about empathy or competition is the other way and I don't know whether we're on the optimistic or pessimistic side of the divide, but it seems like even in a Yudd world where they successfully ban ASI, they would still be living in a world where the true facts of intelligence are dark truths and grandparents can't trust their grandchildren to be nice to them.