I think this misplaces the left side of the window where control is useful. In particular, there will be a somewhat large period where AIs are insufficiently intelligent to successfully execute a takeover, but still could act on misaligned, coherent, long-term goals. Insofar as control approaches can in fact catch or mitigate those actions, there are lots of scenarios where control applied in the earlier stages of RSI is counterfactual for ensuring that later, takeover-capable models are more aligned and/or more poorly positioned to escape control. For exa... (read more)
I think this misplaces the left side of the window where control is useful. In particular, there will be a somewhat large period where AIs are insufficiently intelligent to successfully execute a takeover, but still could act on misaligned, coherent, long-term goals. Insofar as control approaches can in fact catch or mitigate those actions, there are lots of scenarios where control applied in the earlier stages of RSI is counterfactual for ensuring that later, takeover-capable models are more aligned and/or more poorly positioned to escape control. For exa... (read more)