x
Removing an Attacking AI Agent's Ability to Prove Attack Success: Interim Research Results — LessWrong