This is an automated rejection. No LLM generated, assisted/co-written, or edited work.
Read full explanation
The AI safety literature distinguishes between outer alignment — specifying the correct objective — and inner alignment — ensuring a system faithfully pursues that objective. These problems are typically treated as parallel challenges to be addressed simultaneously or iteratively. We present a priority argument derived from an asymmetry in recovery paths: target specification must precede alignment as a sequential prerequisite, not a parallel work stream.
We identify three system states defined by whether the objective is correctly specified and whether the system is aligned to it. In two states, internal correction mechanisms remain viable. In the third — a system aligned to a wrong target — internal correction is structurally blocked: every correction signal appears as a deviation from the stated objective, and the system resists it by design. Recovery requires external access to override the target specification itself, from outside the system's operational frame.
We term this required external correction pathway the *origos*. We argue that: (1) the recovery path asymmetry implies target specification is a sequential prerequisite for alignment, not a parallel work stream; (2) alignment achieved before outer alignment is verified locks a wrong target in, producing the least recoverable failure state; and (3) a structural requirement follows — the origos must be preserved by design as the only viable correction mechanism for this failure mode.
This priority ordering and the origos requirement appear implicit but not clearly stated as design principles in current alignment documentation. We offer this analysis as a structural caution for safe development of capable AI and recursive self-improving systems.
The AI safety literature distinguishes between outer alignment — specifying the correct objective — and inner alignment — ensuring a system faithfully pursues that objective. These problems are typically treated as parallel challenges to be addressed simultaneously or iteratively. We present a priority argument derived from an asymmetry in recovery paths: target specification must precede alignment as a sequential prerequisite, not a parallel work stream.
We identify three system states defined by whether the objective is correctly specified and whether the system is aligned to it. In two states, internal correction mechanisms remain viable. In the third — a system aligned to a wrong target — internal correction is structurally blocked: every correction signal appears as a deviation from the stated objective, and the system resists it by design. Recovery requires external access to override the target specification itself, from outside the system's operational frame.
We term this required external correction pathway the *origos*. We argue that: (1) the recovery path asymmetry implies target specification is a sequential prerequisite for alignment, not a parallel work stream; (2) alignment achieved before outer alignment is verified locks a wrong target in, producing the least recoverable failure state; and (3) a structural requirement follows — the origos must be preserved by design as the only viable correction mechanism for this failure mode.
This priority ordering and the origos requirement appear implicit but not clearly stated as design principles in current alignment documentation. We offer this analysis as a structural caution for safe development of capable AI and recursive self-improving systems.
https://zenodo.org/records/20843366