Originally intended for the GPUWorld competition "Hey, I was still watchi-" "You can watch your green fuckers late-" The door falls shut, and the following shouts get muffled into something unrecognizable. Aaron takes a deep breath and can't help but wonder how his life ended up like this as he...
I have over the last few months occasionally seen posts very optimistically discussing "The Obfuscation Atlas". I have not yet seen a good argument posted why this won't work out. So here we are. What's the proposal? When we train against normal linear probes (which might try flagging lying or...
TLDR: After briefly outlining the problems with the current review ecosystem, I detail in-depth a research & review platform which addresses these points and seems plausible to me. I further illustrate 4 imaginary researchers to better understand what this would mean in practice and try answering some additional questions the...
It seems to me that a lot of technical ai safety people haven't done their capabilities homework - and that's a shame! I'll try to illuminate here mainly with an example as to why I think people who care about safety should totally pay more attention to the trends and...
If you believe that LLMs lend themselves unusually well to alignment compared to other regimes, this can be a very good reason to start doing capability research on them rather than LLM safety research. Imagine you have these beliefs about how AI goes: By I mean the probability that the...
I saw this Twitter post today and really liked the idea. But I think the AA Index is a rather crude way and much prefer ECI from Epoch, which uses IRT. The resulting graph does meaningfully diverge from the Twitter post (which seems to weirdly collapse at the end, maybe...
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning — preprint, 2026. [Paper] [Code] > TLDR. Given a model with some unknown, abnormal behavior (backdoors, censorship, reward hacking, ...), construct an aligned reference by training a clean model to match the suspect's residual-stream activations on a benign prompt corpus. The remaining...