Preface This is part of a series of essays on corrigibility and alignment attempting to formulate an understanding of what makes alignment seem intractable, and how a pragmatic approach to problem formulation, falsification, and decomposition might produce the groundwork for a formal solution someday. Note: I use some terms loosely...
OpenAI has been consistently making it hard to understand what each of their different models are for. Frustratingly, their approach—as indicated by their recent UI decisions in the webapp to hide all settings for model reasoning effort and even the actual model name/type itself under: first, a collapsible element, and...
In this note I’ll try and lay out briefly some thoughts about corrigibility and why it matters. First, some housekeeping. Definitions When I say “corrigibility”, I mean “correctable” as per the original sense of the term—etymologically derived from corrigible, of the same meaning in French, itself hailing from the Latin...
Introduction: The AI Haters In the early months of 2026, generative AI has now improved (at exceeding speed) to a point where many trademark critiques have become dated. The famous 'gotcha' that AI can never make normal-looking art because it always under or over estimates the number of fingers on...
So there’s been a lot of talk recently about AI slop taking over the internet. And I just wonder if it's going to end up changing the way we talk. And more importantly: should we (want to) change the way we talk? Many places (websites, schools, forums, although strangely enough,...