We did some nice theoretical and empirical analysis of this problem in https://arxiv.org/abs/2504.18530
iterate on the oversight design until Guards benefit more from increased capability than Houdinis
does seem to be our best option
Regarding the AI pause lever. What tools do we have, or could make, for telling what "Elo" the next generation would be? If we can't get good enough tools, how should we act under uncertainty?
Aside from simply guessing, we can evaluate new models in a controlled environment before releasing, and abort the release if the Elo gap is too high. This would only work if the Elo gap isn't so high that it lets the model escape the controlled environment or convincingly play stupid.
If we have to act under uncertainty, we should skew earlier for the pause. Still with the understanding that wasting the pause could be fatal too.
Let's say you're a club chess player with an Elo of 1500 (early intermediate). If you can beat Magnus Carlsen, rated 2800, at a game of one minute bullet chess, you win a billion dollars. Magnus only gets one minute on his clock, but you get one year on your clock.
Even given that advantage, you don't have a chance. However, there's a twist: During your year of time, you can challenge any other player to a chess game. They get one minute on their clock, you get as much time as you like. If you beat them, they can play Magnus for you instead, using the rest of your remaining time. Or, they can challenge another player, who can challenge another player... who can play Magnus in your stead. If at any point you or one of your proxies lose a game, you lose the challenge and go home empty-handed. We'll just pretend draws never happen.
The question: What is the optimal strategy for beating Magnus, and do you have enough time to have at least even odds of beating him?
Some assumptions:
You don't want to play Magnus directly, or you'll lose. Instead, you want to beat someone slightly better than you, who beats someone slightly better than them, and so on, until finally you beat Magnus. You rely on having a time advantage at each step.
How many steps should you take? Too few, and you'll get smoked in the first round. Too many, and you have to win too many games in a row. Each game is a chance to lose, and the more steps you take, the less of a time advantage your proxies have in each match.
You can choose the players you want to play and ensure they have an even distribution of Elos from your Elo of 1500 to Magnus' Elo of 2800. That means it's best to split your time advantage evenly across steps.
Let's say you chose to play 20 games, so you would play the first game, your first proxy would play the second game, ... and your 19th proxy would play Magnus. Each opponent would be 65 Elo higher than the last. Each game would get about 18 days, or 26,000x their opponent's one-minute clock. This time advantage puts your proxies effectively around 535 Elo ahead (net), winning around 96% of the time. That sounds high, but winning 20 games with probability 96% is only a 41% chance of winning every single one.
Not quite the 50/50 we wanted as a minimum. Plus, sitting through 20 stressful chess matches with a billion dollars on the line sounds like a drag. At the other extreme, you could play two games, with only a single proxy. Each player would be 650 Elo higher than the previous. Each game would be 6 months vs. one minute, giving a total advantage of 10 Elo in your favor in each game. Each game is basically a coin flip, and you end around a 27% chance of winning.
The sweet spot is somewhere in the middle. Seven games with six proxies. Each opponent is 185 Elo higher than the last. Each game is about 52 days vs. one minute, or 75,000x time advantage, which is worth +630 Elo. This means your players are net 445 Elo stronger, winning 93% of the time. To win seven games at 93% each is 59%. So you can have greater than even odds against Magnus after all!
If we don't already, we will soon have the first superhuman intelligences (these will probably be human-AI teams for now, and later pure AI). In order to ensure the godlike ASI we eventually build will be aligned to our values, we need to ensure each previous intelligence is also aligned. One failure anywhere on the ladder kills us.
Right now we are that lowly club player, hoping we can beat Magnus. Whether or not we can depends on whether the AI story lines up with the chess story or it doesn't. The chess story was chosen because I could actually estimate the relevant variables based on how chess is played, and it neatly communicates the idea. But, here are some ways the AI story might be different:
On the one hand, the Magnus Challenge is an inspiring story. You, a lowly 1500 Elo club player, have absolutely no chance of beating Magnus by yourself, even if he spends the first 6 moves swapping his king and queen just to troll you. But by defeating a successive chain of opponents and gaining their strength for yourself, you can beat him more likely than not.
A lot of people suggest the same plan for aligning AI. Will it work? Well, AI is already not perfectly aligned, so in a sense we've failed. But we don't need perfect alignment at each step in the chain. We need enough alignment that we survive until the next model, and the current model makes a sincere attempt to help us align the next model (or will get caught if it tries otherwise). Still, it's a high bar, and for the reasons mentioned above, the AI alignment chain might not be as easy as the chess victory chain.
The one timing lever we have, the AI pause, is something we have to carefully balance between blowing it at a point where it's not needed, vs. holding onto it so long we die before we can use it, when it could have helped. Ideally we pause before the biggest capabilities jump so we have more time when we are at the greatest risk of losing. Since capabilities jumps are currently not very high, I would personally lean toward "not yet". But get the infrastructure in place so we can do it quickly once it's time.
Whatever we do, we have to make sure we're not just aiming for reasonable certainty that the very next step in the chain won't cause disaster. We need much more certainty than that to ensure the entire chain is safe.