Bringing More Expertise to Bear on Alignment
Preamble The preamble is less useful for the typical AlignmentForum/LessWrong reader, who may want to skip to Adversaria vs Basinland section. On 28th of October 2025, Geoffrey Irving, Chief Scientist of the UK AI Security Institute, gave a keynote talk (slides) at the Alignment Conference. The conference was organised by the UK AISI and FAR.AI as part of the Alignment Project, which aims to bring experts from relevant fields to make progress on the alignment problem. TLDR: * Adversaria vs Basinland. We might be in one of two worlds. One where alignment is adversarial (a security problem), one where it is navigational (a search for good basins of training behaviour). We don't know which world we are in, and how we train and deploy AIs may determine this. * We need new disciplines. The field is small, thinly resourced and approached from only a handful of angles. A few well-placed ideas from other disciplines could disproportionately shift what's achievable. * Even if this all fails, evidence of hardness is valuable. Moving past broad framing to details Alignment means ensuring that AI systems do what humans want. This is the broad framing. There is, of course, a lot of complexity and detail hidden in that simple description: Which (subset of) humans? On what topics or domains? What does "want" mean? It is also important to acknowledge that alignment isn't the only problem we need to solve to ensure AI goes well. We should aim to move from the above broad framing to talk about the details and how they interact, in depth. We may have different takes on the details and it is easy to get blocked by such disagreements, but the different areas of expertise also come with different sets of tools, which often work for many choices of details. We should plan for superintelligence Even if we are uncertain or may disagree about when or whether superintelligent systems may arrive, we broadly agree that it is wise to plan for it. While this is certainly the "hard case