x
Attempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes — LessWrong