This is a follow-up to two posts Geodesic released last week on our current research direction. The code for generating the figures can be found at this GitHub repository. In our previous post, we outlined Geodesic's focus on what we term the pre-RL alignment checkpoint of models -- the alignment-relevant...
The first three sections are written for a general TAIS reader who wants to understand what the state of Debate research is and some high-level takeaways of our work. A reader familiar with Debate may like to skip the setup and start with our presentation of An illustrative training run....
TL;DR Models might form detailed representations of the training task distribution and use this to sandbag at deployment time by exploiting even subtle distribution shifts. If successive rounds of training update the model’s representations of the task distribution rather than its underlying tendencies then we could be left playing whack-a-mole:...
My research this summer involved designing control settings to investigate black-box detection protocols. When designing the settings, we had to decide between trying to mitigate research sabotage (RSab) or Dangerous Capability Evaluation (DCE) sandbagging. It was not clear to us which to target, or whether a single setting could be...
Epistemic status: I’ve only spent 3 months working on sandbagging, and have had limited, mixed feedback from established researchers on the ideas in this post. However, I have had positive feedback from most junior researchers I asked, so I think it will be helpful to make these ideas concrete and...
Feedback request: Is the time right for an AI Safety stack exchange? Epistemic status: I think this is a really good idea, and most of the ~20 people I’ve asked in the community agree. I am uncertain of what exactly the best product would look like, and whether the community...