TLDR * Models often behave dishonestly without acquiring a coherent deceptive disposition. * We trained some mid-sized models on their own plausible but false reasoning. * True and false training usually produced nearly identical downstream effects. * Even statements contradicting latent knowledge transferred only weakly to unrelated dishonesty. * General...
Context: This is the first research output from Arcadia Alignment’s scalable oversight team, carried out in collaboration with external researchers and mentors (Simon and Jacob). We aim to do rigorous empirical work on debate - bridging the gap from theory to the alignment tasks we care about. Debate is a...
We are excited to announce that Resolution (fka Sequent) has a $160M grant from Coefficient Giving (cG) to put rigorous alignment research on a (closer to) even footing with the frontier labs. We will use it to accelerate progress towards higher-confidence alignment, or to find evidence and obstacles showing why...
The first three sections are written for a general TAIS reader who wants to understand what the state of Debate research is and some high-level takeaways of our work. A reader familiar with Debate may like to skip the setup and start with our presentation of An illustrative training run....
EDIT: We originally launched under the name Sequent. Read why we renamed to Resolution. Alignment is not on track Artificial superintelligence (ASI) may be developed in the next few years. It is unclear whether alignment is on track to be ready on the same timeframe. At a minimum, the empirical...
Summary This is a summary of a paper published by the alignment team at UK AISI. Read the full paper here. AI research agents may help solve ASI alignment, for example via the following plan: * Build agents that can do empirical alignment work (e.g.~writing code, running experiments, designing evaluations...
TLDR: * Behavior-only descriptions are useful, but insufficient for aligning advanced models with high assurance. * Two models can look equally aligned on ordinary prompts while being driven by very different underlying motivations; this difference may only show up in rare but crucial situations. * So persona research should aim...