x
small targeted training as a mechanistic “amplifier” for discovering latent/learned deceptive circuitry — LessWrong