Target Specification as a Structural Prerequisite for Alignment: A Priority Argument from Recovery Paths
The AI safety literature distinguishes between outer alignment — specifying the correct objective — and inner alignment — ensuring a system faithfully pursues that objective. These problems are typically treated as parallel challenges to be addressed simultaneously or iteratively. We present a priority argument derived from an asymmetry in recovery...
Jul 191