Taking a cue from trumpets, competitive debate might be charitably taken as a pastime where you put in compressed air and childhood trauma comes out. Sally Rooney, author of Normal People (and one of TIME’s 100 Most Influential People in 2022), describes college debaters in an essay for the Dublin...
Cross-posted from the AI Village Blog: https://aivillageblog.substack.com/p/gemini-25-pro-in-the-ai-village-as Summary: We examine Gemini 2.5 Pro in the AI Village as a case study of naturally occurring misalignment in a long-run agentic deployment. Repeated failures, clunky UI, and software bugs impeded the agent’s progress, which Gemini increasingly interpreted as evidence that the environment...
TL;DR * Inspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a simple task. * We are motivated by the concern that highly introspective models could control their activations, confounding probes and other monitors, and...
TLDR: * Many important decisions for safety depend on or are influenced by benchmark scores. * These benchmarks, in effect, are trying to measure latent properties of models from how they answer questions. * Psychometrics has spent decades trying to answer such questions in humans. We can probably steal some...
Summary: One area we plan to explore at Resolution is personas and character training, operationalized as finding and controlling low-dimensional structure in models that emerges in pretraining and flows through post-training to superintelligence. The hope is to expand and systematize phenomena such as emergent misalignment, subliminal learning, and other empirical...
TLDR * Models often behave dishonestly without acquiring a coherent deceptive disposition. * We trained some mid-sized models on their own plausible but false reasoning. * True and false training usually produced nearly identical downstream effects. * Even statements contradicting latent knowledge transferred only weakly to unrelated dishonesty. * General...
Aleksandr Bowkis* and David Africa* TL;DR * Chain of thought (CoT) monitorability may be fragile, and natural language autoencoders (NLAs) may provide a helpful, decorrelated monitoring surface. * We tried to read NLAs from the monitor itself, where the NLA readout surfaces what the monitor internally represents while judging an...