I wrote this for the SaferAI team. I'm publishing it with some small edits, in case it's useful to other orgs. A note on scope, this post is about sourcing candidates specifically. It doesn't cover how we define a role before opening it, how we assess candidates once they're in...
Current AI risk management relies on qualitative approaches, much like nuclear safety before 1975. We propose a shift to quantitative risk modeling, following the approach that transformed nuclear safety. We propose a methodology and demonstrate it by building nine probabilistic models of AI-enabled cyber attacks. This is a first attempt...
We (SaferAI) propose a risk management framework which we think should improve substantially upon existing Frontier Safety Frameworks if followed. It introduces and borrows a range of practice and concepts from other areas of risk management to introduce conceptual clarity and generalize some early intuitions that the field of AI...
Reading guidelines: If you are short on time, just read the section “The importance of quantitative risk tolerance & how to turn it into actionable signals” Tl;dr: We have recently published an AI risk management framework. This framework draws from both existing risk management approaches and AI risk management practices....
Abstract AI safety researchers often rely on LLM “judges” to qualitatively evaluate the output of separate LLMs. We try this for our own interpretability research, but find that our LLM judges are often deeply biased. For example, we use Llama2 to judge whether movie reviews are more “(A) positive” or...