In a previous post, I argued that Bryan Caplan’s signaling theory isn’t a good explanation for why college graduates get higher-paying jobs. Instead, I claimed, understanding the role of higher education in the modern West requires sociological explanations. In this post I argue more specifically that getting an undergraduate degree...
I’ve tried several times to summarize the core question my research is trying to tackle (and, indeed, I often think of research progress as a process of asking increasingly good core questions).[1] This post gives the deepest version of that question I’ve found thus far: how should you relate to...
Note: I wrote this post in 2023, along with the rest of my meta-rationality sequence. I never got around to uploading it, but a recent comment from Vladimir Nesov inspired me come back to it and fill in the few remaining gaps. I still broadly stand by the ideas in...
This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment research” and “capabilities research" thereby lost most of its meaning.[1] In particular, I’ll chronicle the development of what I’ll call the “pragmatic alignment”...
This sequence is about the last decade in AI alignment. Over five posts, it recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progress in favor of iteratively improving existing systems...
Some context for this post: I’ve been working part-time as a consultant for the AI Futures Project over the last year. Most of the work I’ve done for them has involved critiquing and suggesting improvements for their AI 2040 scenario—some of which were addressed, and some of which weren’t. To...
In this post I’ll sketch out an informal model of intelligent agents as webs of beliefs (or belief webs for short). The belief webs framework pulls together ideas from active inference, agent foundations and machine learning. In doing so it aims to unify beliefs, goals and actions as three facets...