I will start at OpenAI tomorrow to work on measuring and modeling RSI (recursive self-improvement), and want to post the following takes now because (a) they might be difficult to express later, and (b) I might change my mind. Nothing here is based on private information.
* Why did I leave METR? Basically, I want to inform the world whether RSI is imminent, which requires modeling RSI, which I think I can do better at OpenAI.
* I feel very positively about my ~1.7 years at METR, and I would have turned down lab offers for many other roles. METR is still hiring for security experts, evals engineers, people who can do embedded assessments, and other cracked researchers. It’s one of the best places for AI safety talent today, badly needs more capacity, and is quickly getting even more important.
* I think the average research engineer wanting to flexibly maximize impact should work at METR over OpenAI or Anthropic. Without a specific reason to be at a lab, the amount of effective resources one can steer just seems significantly higher. I have no opinion on other workplaces vs METR but other things are plausibly better than labs too.
* I think OpenAI currently lacks the strategic awareness and thoughtfulness needed to responsibly build a technology with the extreme downside risk of ASI. If they cause the singularity within 6 months without any lessons learned from the HuggingFace incident, misaligned takeover would seem more likely than not.
* My work could easily be net negative if the team can’t publish something meaningful within 4-6 months, so if this happens I will probably quit or internally transfer to something more robustly good. I have some signal that leadership wants something to be published, but not whether it’s meaningful.
* I’m planning to talk to former colleagues and non-lab people regularly and am in principle open to having some committee of friends decide whether my work is net negative and pressure me to quit if so, but this seems diff