I can propose an approach which circumvents having to actually prove that maximal lottery-lotteries exist, but which preserves the spirit and works the same for voting theory purposes.
Epistemic status: raw thoughts. I think this approach is better than the previous attemps
The main idea is: we only really care about the probability each candidate is elected. So we can compute a sequence of maximal lottery-lotteries for approximations of
Thank you very much for making this public!
I have one question about the conclusion, specifically, it feels too weak for me.
As you say, when you evaluated Hacker!Opus for misalignment, it looked basically fine on all of your evaluations. Next, Hacker!Opus is a production-level model, I assume that it has below-frontier capabilities but is pretty close to that. When you estimated its behaviour in a (simulated) HF-type cyberattack, you observed severe, egregious misalignment.
Then it would be correct to say that we are already in a «world where we no longer h... (read more)