Summary * The primary ToC for control makes the case that control is compelling even if it does not scale to ASI. I think it is underdiscussed that this is also true for alignment (for all the same reasons). * Even though control does not need to scale to ASI,...
I am going to discuss five kinds of inner misalignment and two kinds of outer misalignment, which create a simple taxonomy of alignment failure modes. When I talk about a kind of misalignment here, I am talking about a reason for misalignment (like inner/outer misalignment), not a kind of misaligned...
I am going to argue that we will likely eventually get AIs that are strongly power-seeking, much more so than current SOTA LLMs.[1] TLDR 1. Right now SOTA LLMs are still largely in a simulator regime. This buffers against power-seeking. 2. Long-horizon RL or similar methods (applied to LLMs or...
Edit: we renamed the technique from Monitor Sensitive Training (MST) to Evaluation Conditioned Training (ECT)[1] TL;DR * We introduce Evaluation Conditioned Training (ECT), a new post-training technique where we augment training data with evaluation labels that describe how evaluation is going to be applied for each sample. We then change...