This essay grew out of conversations with Danaja Rutar, Paul Colognese and Eric Michaud. It proposes an alternate hypothesis for how and why models might be becoming increasingly misaligned in training and eval environments while seemingly more aligned in real world use: the opposite of what we would expect if...
This post is my attempt at explaining my dissatisfaction with the current alignment narrative, and different paths that we might do well to explore more. In brief, I think control as the dominant narrative of what alignment looks like is misguided because it merely shifts the problem towards improving human...
Science as attunement, from Galileo to language models. Crossposted from my website. Written by me and edited in collaboration with Claude Fable (Anthropic). Many would agree that the scientific method is the best process we have discovered to understand the world. But what is the scientific method? The products of...
This started off as an entry for Dwarkesh's blog post contest, specifically an answer to his first question on why intuitions about slowdowns in reinforcement learning (RL) progress have either not come true or have had mixed success. His 1000 word limit turned out to be too little to make...
Over the past few years, the ways we think and process information as a society have undergone a marked shift, some features of which remain underdiscussed. As you read this sentence right now, millions of people are talking and thinking with the same entity - one of a very small...
I am a mathematician turned cognitive scientist/AI researcher. I wrote a short essay on the philosophy of mathematics that I think will be of interest to many people here. I combine platonist and formalist ideas and frame mathematics as an experimental activity, extremely similar to physics and other hard sciences...