Top postsTop post
TLDR: So there has been recent discourse on 𝕏, and recent news of major cyber attacks that were done with the help of AI. The missing frame here is the dual-use gap: as AI models become more capable, they create more upside for defenders and more downside for attackers. The...
TLDR: Pragmatic interp sounds great in the sense that you get to keep interp tools while actually moving safety metrics, but looking a bit closer it's kinda a trap. You pay interp's overhead but get judged against black-box baselines that don't, so the work that survives is whatever cleared that...