Attempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes
Epistemic status: empirical report from a 30-hour project sprint. Null result, reported honestly, with full code and data. TL;DR MacDiarmid et al. (2024) showed that a linear probe on model's internal activations can catch a sleeper agent about to defect despite knowing that directly asking the model fails completely. From...
Jul 204