x
Towards deployment-time misalignment continuation evals: lessons from recent AI agent incidents — LessWrong