This is the first of a series of posts about automating AI Safety Research. TL;DR Automating large parts of AI safety research now seems possible. Alignment research faces three recurring challenges: measuring model capabilities and safety, developing techniques that make models safer, and establishing why those measurements and techniques should...
TL;DR There is some consensus that LLMs are bad at hard-to-verify tasks. The question is whether models are getting better at them over time. As a motivating example, I gave one research-reproduction task to 12 models spanning three years of progress, to illustrate how (1) what looked like an emergent...
Based on a 2-day hackathon brainstorm. Current status: 70% of the tooling is done, unsure of how to proceed. Not enough experience with multi-month sized projects to judge for feasibility. I'm looking for some feedback. Specifically I want feedback regarding my current implementation. The statement "SAEs could be useful for...
TL;DR * Small Language Models are getting better an at an accelerated pace, enabling the study of behaviors that just a few months ago were only observed in SOTA models. This, paired with the release of the suite of Sparse Autoencoders "Gemma Scope" by Google Deep Mind, makes this kind...