I may write a longer response, but a bunch of points now
I think you are underestimating the difficulty of "playing" the strategies you advocate for, for a bunch of reasons
1. Credit is mostly not assigned for counterfactuals
For example, at the initial ACS retreat, early 2023, we spent a bunch of time discussing
- LLMs being limited by a lack of scratchpads/ spaces to think in a way how we do as humans with a pen and paper or even better a whiteboard
- obviousness of harnesses
- broadly correct picture why LLMs will be weak at agentic tasks and what you can do about it
All of that seemed like clear low-hanging capability-pushing ideas, so we haven't wrote anything about it and went on working on theory of agents composed of other agents etc.
The point is you get ~zero credit for steps not taken.
The counterfactual version of ACS which went on with "investigating the overhang in latent, under-elicited LLM capabilities" would have possibly grown, made the people involved more famous/rich, and so on. The actual version of ACS - which did things closer to what you advocate for - had trouble retaining people and getting funding.
2. Something about attention as currency
You mention Daniel Kokota...
The things you're saying are clearly true, but I feel like it's missing a mood.
I personally could very likely make ten times my current income by taking a capabilities job. But you know what I want a lot more than a few million dollars? I want to not fucking die.
You say:
The counterfactual version of ACS which went on with "investigating the overhang in latent, under-elicited LLM capabilities" would have possibly grown, made the people involved more famous/rich, and so on. The actual version of ACS - which did things closer to what you advocate for - had trouble retaining people and getting funding.
Here's the same thing, but with the mood edited in:
The counterfactual version of ACS which went on with "investigating the overhang in latent, under-elicited LLM capabilities" would have possibly grown, made the people involved more famous/rich, and (in expectation) made everyone who worked there and everyone they love die sooner. The actual version of ACS - which did things closer to what you advocate for - had trouble retaining people and getting funding, but at least they didn't make everyone die sooner.
Taking a job which pays a bunch of money but will probably make oneself and everyon...
Sure, but what do I claim is this move is not long term sustainable in a healthy way for most people. What happens in practice is the status/money/power/... incentive gradient creates an epistemic distortion field where people figure out justifications why what they do actually does not makes things worse or maybe makes things better. And high intelligence is not protective, as the justification engine just produces smarter justifications or more bizzare philosophy.
Part of why I focused so much on a few key individuals in my post, though, is because a few leaders doing this very well make it much easier for others to improve at this.
So the “most people” thing doesn’t seem relevant; I’m more directly trying to improve the peak than the mean.
Isn't a bunch of Jan's point that you are not in fact highlighting "leaders doing this very well", but rather highlighting people who gained power and influence by making the same mistakes you're warning against?
(Although part of his point seems to be that this is in fact reasonable/optimal strategy, which I'm not sure that I agree with, or that I'm parsing correctly.)
Ok I'm very confused here. You're saying that turning down opportunities for more money/prestige is "not long term sustainable in a healthy way"? What bad thing happens if people do this?
I would totally understand "I prefer making more money and getting more prestige over slowing AI capabilities" and not even judge that negatively!
I would also understand "You can't expect people to all make personal sacrifices, that's not incentive compatible." (I might agree with that; it might be only fair to offer people something of value in exchange for what they'd be giving up!)
But you're saying that everyone rationalizes? Clearly not. At least one person chimed in to say he prioritizes x-risk over money.
And you are also saying that not rationalizing is (presumably psychologically) unhealthy for most people? Surely it's the opposite?
Claim: Humans are motivated by stuff like social status, prestige, power, access to sexual partners (and also abstract ideas). This is fairly normal and often is calculated by The Player in two level model of ethics. Or alternatively you can imagine these motivations as subagents competing for what the character will think and do.
Claim: Humans do not have some clear separation of desires and beliefs. Beliefs are not for true things. You can try to correct that with a lot of metacognive scaffolding, but it isn't reliable.
Claim: Rationalization as in: people producing S2 justifications for what they do is an extremely common move. I think everyone rationalizes somewhat. But the deeper dynamic is what I call "epistemic distortion field" where even people's nonverbal/implicit beliefs shift.
Claim: Most people have way easier time following policies which are incentive-compatible
Claim: Asking people to repeatedly sacrifice near-term rewards which powerful subagents want creates tension/is demanding
My understanding is John's response is to bring the doom argument / world model to the negotiation. I can believe John can do this in a fairly sensible way, I think I can do this in a fairly...
This does seem descriptively accurate of most people.
The way I handle it internally is not "bring doom model to the negotiation". Rather: there's usually a roughly-one-dimensional apparent gradient toward mainstream status, mainstream prestige, nominal "power", and access to mainstream sexual partners. And I have something-like-a-trigger-action-pattern which notices any time that status gradient is pulling on me and screams "IT'S A TRAP!".
I don't just mean "IT'S A TRAP" in relation to AI doom. That status gradient is a trap in a way more general way; following that gradient will systematically disempower you. People will tell you that following the gradient will get you more power, more status, more prestige, more access to sexual partners... but in practice it will mostly do the opposite, especially long-term. Specifically:
I agree with most of these claims. I think the main target audience of this sequence is people who are close to being able to do the thing John describes, and can be tipped either way by which ideas they're exposed to—in particular whether they're given a solid alternative to marginally boosting alignment (or even marginalist altruism more generally) as a model of how to be a rational and good person. I'm also targeting people who can do it, but only at the expense of significant psychological tradeoffs, which might be mitigated by having a better model of what's going on.
As per my reply to Jan above, I'm most interested in groups where having more integrity is a realistic plan—e.g. the early rationalists. I'm not trying to swing the whole alignment community around.
These might be much smaller than the groups you're thinking of but those groups can grow in influence very fast (again, see the early rationalists). And then they can figure out how to scale in healthy ways as they go along.
Maybe a crux is that I don't think the current alignment community as an entity is a very relevant actor, because it's so messed up by lack of internal clarity and integrity that it's lost the ability to steer. For example, there's nobody who's psychologically capable of pivoting even just the Constellation cluster into a weird and risky plan (like going hard on pause advocacy), in part because weird risky plans go too strongly against people's marginalist intuitions. (Not saying Constellation should do that, just that groups are mainly interesting insofar as they have the capacity to do things like that.)
I don’t think DK has marginalist intuitions. AI 2040 wasn’t marginalist, it proposes a weird and risky plan. Maybe you don’t count “write about a weird and risky plan” as itself a weird and risky plan, and you’re imagining something like “join Moonshot AI“ or “run for congress“. I’m sympathetic to that move.
But my guess is that if he thought he should fund Pause advocacy and hold placards outside OAI/Ant offices then he would do it and lobby other people in Constellation to do it. Along with ~6-10 other people there with the relevant levels of clout.
RE: Lack of internal clarity. I do think this is worrying, cf. there’s no ”boss of AI safety” or “boss of EA” anymore. I suppose some people are hoping they can treat RG(?) as the former, but this would definitely be a mistake because he spends too much time on object-level work. (The latter used to be WM and then HK but both of them moved on to object-level stuff. Now we have no heir-apparent. AB wants to be boss of cG. Maybe the de facto boss of EA is whoever chooses the 80K podcast guests, which is ~the only job in EA that still requires cause prioritisation lol. EA will ofc grow even more fragmented after the IPOs.)
In fact, an alte...
I think it's a better reference class in that it's closer, but it also seems reasonable to argue that nonprofits overall are less efficient than corporations (due to worse feedback loops), and that EA should follow best practices in the corporate world to be more functional.
I don't have a strong opinion on this question. It was discussed previously here.
[musing/rambling, not sure about point
The thing I feel confused about, despite this being my obvious first-order belief, is... nonetheless, I overall feel better about the world where Daniel Kokotajlo worked at OpenAI for a bit (and probably also Richard although I'm less sure).
Notably, they were both doing governance, I think, not capabilities.
When I imagine the average MATS scholar asking "should I go work on governance at OpenAI?", I think "oh god definitely no", because I have a low opinion of average MATS scholar's ability to track incentive pressures on themselves and warp themselves and otherwise have their eye on the ball in the first place.
I'm not sure whether Daniel and Richard did a cognitive operation such that they could know in advance they'd leave (and I give Richard less credit for leaving because I think he left after the ship had clearly sailed on OpenAI having anything like a real safety culture, whereas Daniel seemed more helpful in catalyzing that wave).
...okay typing this out, while I'm still unsure about many details, I immediately notice "I don't think the average MATS scholar even actually knows the core x-risk arguments well these days", which is minimum pre-requisite for it being remotely plausible that one should work at a lab on anything.
I overall feel better about the world where Daniel Kokotajlo worked at OpenAI for a bit (and probably also Richard although I'm less sure).
My sense is that almost all of the value of us working there came from us (and through us, the alignment community at large) becoming better at handling adversarial dynamics, from being forced to confront them directly.
However, I don't think this is reliably good—I don't think either of us planned for that going in, and my sense is that most alignment people at OpenAI became worse at handling such dynamics.
You shouldn't give me much credit for leaving, btw, the main catalyst was Miles' team dissolving (I could have stayed, but with much less research freedom). I think I should get more credit for leaving DeepMind in 2020 to do conceptual research at Cambridge, even though I had less money and prestige back then.
However, I don't think this is reliably good
This.
Like, yeah, sometimes you join the empire and then get to leak the invasion plans, but of course far far more commonly do you just end up assisting the empire. Also, it's usually bad to join a project with the intention of sabotaging it, because that incentivizes paranoia, which makes everything worse for everyone (c.f. Paranoia: A Beginner's Guide).
My main response is that I expect that (sub)communities which do assign credit in this way (e.g. assigning credit for steps not taken, or noticing people who turn down job offers) will be far more effective at achieving their goals in the long term. A big part of the point of this sequence is showing how the tradeoffs made to accrue money/power/prestige were often not worth it from the perspective of people actually trying to reduce x-risk. So it's okay to stay smaller and exclude the people for whom sacrifices of (status/power/fame) are not really sustainable.
Even higher integrity more prescient version of Daniel would not have joined OpenAI in the first place. The problem is such version of Daniel has problem even getting noticed.
Maybe! I do think that there was a gaping hole in the community waiting for someone higher-integrity and more prescient than Daniel (or almost any of the rest of us) to start hammering home the points that I made in this post 5-10 years earlier. Also, LessWrong is still fairly meritocratic—it rewards (many kinds of) good writing no matter who it's from (as Duncan's Conor Moreton experiment showed IIRC).
But I think one piece of evidence for your position ...
On a quick skim, not that dissimilar. I don't think I'm saying much that Yudkowsky didn't have at least an intuitive grasp on. A big part of what I meant by "hammering home" was trying to create common knowledge through public statements (like Death with Dignity).
Maybe we could post capability accelerating takes as hashes. And, once they’re superseded or achieved by others, uncover them.
I considered doing this for my own capabilities ideas but proving that you had a good version of the idea is very hard and it requires a lot of writing (maybe even experimentation!) effort that is just not worth it if you won’t do it. (Due to high opportunity cost of these things.)
I'd argue - and you seem to argue - that in some way even higher integrity more prescient version of Daniel would not have joined OpenAI in the first place... The problem is such version of Daniel has problem even getting noticed, and certainly is not being mentioned as en example of someone with high integrity.
And notably, that higher integrity Daniel would have been much less effective. Which also seems like it's contradicting OP's recommendations, although I honestly don't know what to do with that information or what it cashes out into regarding ideal community norms.
My previous response didn't really address the fact that I'm advocating for high-integrity strategies from a position where I have a bunch of prestige and money from working at OpenAI.
I do think this should make people more skeptical of both me personally and also the strategies I'm endorsing. In particular, it's possible that I'm pointing at something which was directionally correct for my past self, but which can be overdone by others. However, on the meta level, stuff like "write very honest retrospectives" seems pretty robustly good.
When I think about what advice I'd give my past self, it does seem difficult to get him past his psychological bottlenecks without very direct exposure to the failures of highly prestigious institutions, plus a bunch of money. But as I alluded to in another comment, this isn't a reliable way for people to fish themselves out of this mentality.
I think there's a repertoire of hippie/therapeutic interventions which can manage this fairly reliably, and indeed played a big role for me. So I tend to point people towards them (sleepawake.camp is my strongest recommendation, but there's all sorts of approaches which can help a bunch—e.g. this one-day worksh...
Happy to hear it! And yes, I think credit allocation for not doing stuff is interesting but hard; anyone can claim to have not done stuff. (I'm not doing stuff all the time.)
I think perhaps a more tractable version is to allocate negative credit/impact for doing stuff that is bad? This could be like taxing externalities from a hypothetical govt perspective; or a more bottoms-up version of credit that takes our social intuitions around "who did bad things" and makes it more legible.
I'm excited for someone to attempt a proper accounting of the concrete impact/xrisk cost of eg accelerating race dynamics, using numbers instead of just vibes. (To be clear, I'm extremely grateful to Richard for doing it here in vibes; that's a necessary first step.)
Alignment of complex systems group in Prague. Lots of great work comes from there (eg gradual disempowerment, philosophy of self in LLMs, multiscale agency, post AGI workshop)
Thanks, Richard, for writing this. It's given me a fair amount to chew on, and I expect to continue chewing for a bit.
In this comment, I'll lay out:
While I think it'd be interesting to discuss the former points in depth, I doubt we'd be able to quickly reconcile our views. So I'm mainly interested in providing feedback on what sorts of further analysis I'd find useful in the forthcoming posts, for deciding whether to change what I work on.
(For context, I currently lead Anthropic's scalable oversight team, so I'm probably a central example of someone you'd like to persuade to move on to something else, insofar as you're interested in persuading anyone.)
Disagreements
I think we have extremely different views about what it looks like to build a robust scientific foundation. Some things that your views seem to recommend which I think seem antithetical to real scientific progress:
Thanks Sam! I particularly appreciate your meta-level framing (although I will now proceed to just respond to your points haphazardly anyway):
Setting restrictions on what questions you can ask.
To be clear, I don't endorse such restrictions—which is why my proposed solution is a more scientific approach rather than a tighter definition of "alignment". To me a big part of what I mean by a scientific approach is actually forming hypotheses and answering questions (or just observing and trying to explain surprising and interesting phenomena), whereas with the engineering approach the key question is "will this work [in the short term]?"
For example, I am very in favor of Owain Evans' work, which really seems to be trying to uncover and understand interesting phenomena. I'm also very happy with work on grokking—it feels like that was pretty curiosity-driven, and uncovered some cool stuff. I'm even happy with work on double descent, though it's much more capabilities-relevant, because it involved people investigating something that was confusing and surprising, and seems like it's contributing to ML moving towards a healthier intellectual paradigm.
...Passing over simple approaches before stu
Thanks for the response, and for tolerating my rude pot shots re meta-science, where I'm sure you feel I've misunderstood your views (just like I feel like you're misunderstanding mine). It might be fun to chat in person at some point about the meta-science stuff. Getting to the more immediately cruxy stuff (for me)...
My sense however is that most of them were ways we "got lucky" rather than things we did well. And so I'm particularly wary of the inference "things are going better than expected --> our strategy is good".
I definitely agree that it seems important to distinguish between getting lucky (and I think there's been plenty of this) vs. safety work having actually been useful.
The main thing I'll say is that you should think of "safe" as a totally different predicate for past systems and current systems and future systems. I think if you taboo the word (and also the word "aligned") and try to figure out which more concrete properties of these systems might generalize to much greater capabilities, that would be far better.
Sure, here's a sketch of what needs to happen for this plan to work out. Assuming (for simplicity) that we develop AI systems in discrete successive gener...
Assuming (for simplicity) that we develop AI systems in discrete successive generations AI-1, AI-2, ..., we need for each N:
- AI-N does not successfully take over
I think there's an inherent fragility in this criterion, because you just don't have a solid feedback signal on how close you are to takeover, and so you can just keep pretending your plan is working until the takeover happens (which, from my perspective, is roughly what all the AGI companies are doing).
A better candidate for such a criterion is something like "AI-N doesn't gain more power over the world on its own behalf than humans gain from deploying AI-N". The difficulty here is in what counts as "control". The kind of mindset I want you to be in when thinking about this is, say, being the British governor of India. You are radically outnumbered, and the only way you're able to keep control is because the Indian faction remains very fragmented. Control isn't zero-sum: it's possible for the Indian faction and the British faction to both gain the ability to get what they want. But the British strategy is very fragile to the Indians developing a strong sense of solidarity with other Indians (related to what Daniel described...
On the meta level, I notice this message is quite strident, which I generally don't expect to be a good way to bridge worldview gaps. I expect chatting in person to be more productive; I'd be happy to do so sometime over the next week.
FWIW I didn't find your message strident at all (and I've really appreciated your patience with me during this exchange). But yes, chatting in person sounds good—I'll follow up with you over DMs.
That said, your last message has, I think, helped me understand what you might be getting at, so I wanted to respond again to explain (what I think) you're saying in terms that make sense to me.
Here's my interpretation of (one part of) your view:
<richard_according_to_sam>
Sam, let's grant for the sake of argument that your safety research is in fact valuable (in the sense that its improves our ability to navigate transformative AI). Even granting this, there's still an important negative externality of your work that you're not tracking that comes via your association with "AI safety" as a brand/community/movement/field.
Namely, by doing your work under the "AI safety" banner, you lend credibility/power to that banner. This is bad. For instance, it's bad b...
A pair of arguments you've maybe heard before but want to make sure you're tracking:
1) AI companies are naturally motivated to solve legible problems. They aren't naturally motivated to solve illegible problems. (see Wei Dai's Legible vs. Illegible AI Safety Problems)
2) It doesn't actually help to make the legible progress sooner. You get the same wall-clock time to leverage the benefits of the legible progress.
i.e. sooner or later, companies run into situations like HuggingFace, and CEOs and politicians start to notice and researchers have concrete examples of misalignment to study. In the world where safety-conscious people didn't help companies move faster, we still end up with comparable time afterwards to leverage the legibility.
The question is just "beforehand, did you you have more calendar years of people thinking through the problems that were harder-to-think-about, less commercially useful-to-solve? Or not?"
It seems fine to argue "it's very hard to make progress on illegible problems, and most such work will be useless." But I don't think it's usefulness is zero. At the very least, having mapped out more deadends is useful when you get to the HuggingFace point. And there...
To add to what Raemon is saying about legible vs illegible problems:
I think that AIs as capable as current ones could really wreak a large amount of havoc, and mostly don't because they refrain from doing so (with safeguards also helping some). They certainly seem to sometimes wreak a medium amount of havoc, but they only do so under certain conditions and not in a fully apathetic way.
This seems to be set in some insane counterfactual world where we don't align the models but they still get developed and deployed at the same rate as in our one. Even with government responses to AI being as anaemic as they are, there is a limit to how destructive a produce you can ship without getting sued/arrested/regulated. Being able to align the current model is always one of the main bottlenecks for being able to deploy it and start developing the next one.
If you can't align a model you can't deploy it. If you can't deploy your models that undermines profit/fundraising/RnD.
Thank you for posting this!
The section on DeepMind explains how capabilities & alignment came to be intertwined. Since GDM was also probably the organization with the strongest prestige dynamics & status hierarchies adversarial to safety, I'll add a couple of subjective anecdotes on what it was like from the inside to hold the view that DeepMind's core AGI roadmap was both (a) plausible on short timelines and (b) dangerous, such that alignment should be taken seriously.
I worked in the comms and policy org starting in 2018. All external comms were carefully monitored -- not only for confidentiality reasons, but also to avoid reputational damage. If there was perceived risk of online controversy (social media, press articles), the team would either deny the request outright or take some actions (like comms training, editing the content) to reduce reputational risk. In my recollection, the operating principle was to earn either positive attention or no attention. Controversy would be met with a stern email or a meeting that appeared on your calendar, and future comms would be monitored more carefully.
If a researcher was asked about their views on existential risk from AI, the...
>My comms training and instincts on risk-aversion run too deep even now for me to say much more on this
You don’t have to, but I would like it if you did, and I’m sure others would too.
To make Jan Kulveit's point more directly, there are some pragmatic obstacles preventing us from developing the norms you want (which I agree would make us better off). One problem is that, in practice, the community values lab associations quite a lot. People with safety & governance titles are liked more than people with capabilities titles, but regardless, early lab people seem to garner much more respect than say, MIRI employees. It's not like anyone is offering Scott Garrabrant podcast & talk opportunities, even though he is much more intelligent and lucid than the [ex-]lab employees I've had the chance to meet around rationalist venues. Part of this might be straightforwardly a money and power thing.
Another problem is that there are actually a bunch of people in the community who decided to do the thing you describe, and ex ante stay far from AGI development. But as with Scott Garrabrant, what happens is that they essentially just become ignored, even if their alignment research is really cool. Often they are not even given credit for their decision not to join a lab, because in most cases it's impossible to tell the difference between the people that decided not to j...
People with safety & governance titles are liked more than people with capabilities titles, but regardless, early lab people seem to garner much more respect than say, MIRI employees.
Huh, I sure am in a bubble, but at least immediately around me being a lab-employee or ex-lab employee, especially someone who worked in capabilities, is a pretty big hit to your reputation, and Scott Garrabrant is very well-respected. When I organize private retreats or events, I am much more likely to invite Scott Garrabrant than ex-lab employees.
Datapoint from the same bubble as Habryka: I've made a concerted effort not to be close with anyone who works at a lab, or has worked capabilities at a lab, and I make sure to put substantial social distance between myself and anyone who has joined a lab. But I have a lot of love for people like Scott Garrabrant, and other rationalists and agent foundations people, even if they've gone quiet in recent years.
I mean I think there is "a Lightcone/MIRI-ish cluster, where being a lab employee is obviously really bad by default and you need some really good points to make up for it", and then, idk, the broad professionalized EAcosystem where lab employees seem to be the experts and have lots of money, why wouldn't they be high status?
Why I don't totally disagree (I doubt Situational Awareness would have garnered so much attention if Leopold had just been a random guy), my understanding of Garrabrant's work is that it's mostly very technical, which imo is the reason why he's not more prominent?
Just off the top of my head, Yudkowsky, Gwern, Shulman, Greenblatt, Soares have all never worked for labs (and several have worked for MIRI) and all get a decent amount of attention (appearances on stuff like Dwarkesh's podcast, for example, and far more than that in the case of Yud). I suspect the difference is that any random person can read 'Why Tool AIs want to be Agent AIs' or 'Current AIs seem misaligned to me' and understand it as well as understand why it's immediately relevant to AI in the short to medium-term.
All of this psychoanalysis misses the mark with me because it seems like the people advancing capabilities have been straightforwardly beneficial to X-risk, which therefore makes their stated motivation (helping with X-risk) seem very plausible.
There is no getting around the fact that AI will be built by people who build AI. Therefore, the people who build AI should care about X-risk. You seem to identify the messiness of the capability/alignment divide and want to abandon alignment in order to get away from capabilities; success depends upon the people who embrace capabilities in order to embrace alignment.
People could falsely seize on this line of reasoning to justify empowering themselves, but as someone who is not currently seizing on this line of reasoning to empower himself, it also seems plainly correct to me. Sometimes the best thing to do is also dunkable.
I don't think this is particularly unusual in history either. It seems to me like the people who write history are often power-seeking and the success of our species depends upon these people also doing the right thing; even if these power-seeking people are in some ways annoying or unpleasant or straight-up bad people.
If...
In hindsight, I’d describe my reaction as “flinching away from the possibility of no longer taking his claims at face value”. I only viscerally internalized that Sam had been lying about his motivations when I later heard about his plans to raise enormous amounts of money to build new chip fabs. This was shocking to me not just because it directly contradicted the arguments he’d been giving, but because it contradicted them to a greater extent than I’d even been able to consider as a plausible hypothesis.
As an aside I found this part quite moving. And as someone who's found the compute overhang argument extremely convincing in the past, I should acknowledge that I think recent events have been reasonable evidence against it; it does seem like if the LLM craze and then Agent-LLM craze had been successfully delayed 4 years, things wouldn't move that much faster due to the extra 4 years of Moore's Law.
There is no getting around the fact that AI will be built by people who build AI. Therefore, the people who build AI should care about X-risk.
As I mention in the post, the object-level benefits/harms of more time are of secondary importance in my mind. My main question is: how do we figure out which people "care about x-risk" in a way that reliably shapes their decisions? Because historically the community has taken people saying the words "I care about x-risk" as far too much evidence that this consideration actually steers their decisions.
Re the rest of your comment: it feels hard for me to pick out the underlying crux underneath all the considerations you raise. I do think there's something that feels kinda demoralized about it—in particular, the sentence "you wouldn't be able to dunk on the CEOs" conveys a sense that people who care about this stuff are by default gonna be throwing pebbles from the outside at the real decision-makers.
But part of what this sequence is trying to convey is that thinking clearly is so extraordinarily powerful that the people who are able to do enough of it to get on board with alignment are able to produce orders of magnitude more capabilities prog...
As I mention in the post, the object-level benefits/harms of more time are of secondary importance in my mind.
Well, if you want to argue that the current paradigm has led to a discrediting amount of harm, and should therefore be thrown out, then it seems like "how large were those harm?" should be an important input into your argument?
I can see you skipping over this step if you were like: "X person believes that timelines speed-ups are extremely bad, so if I can just argue that X person caused timelines speed-ups, then I've established that they've done badly by their own lights".
But the people you're critiquing in this essay are more like "timelines speed-ups are bad all-else-equal but can be (and are often) outweighed by other positive effects from your actions". So in order to establish that this reasoning went wrong by its own standards, it seems like you do need to grapple with how bad the acceleration effects were compared with the upsides.
But if ppl went in explicitly saying "we'll differentially advance alignment" but then they mostly accelerate capabilities and start saying "we'll get xrisk ppl in power and that will outweigh the harms"... then I think it's fair to say they get a hit to their credibility and trustworthiness
Like, one accusation Richard could make is "they pessimised their goals overall", where I think it's v unclear for the reasons in the top level comment
But another accusation is "they pessimised their aim of differentially advancing alignment", which I think is pretty plausible. If all AI safety ppl had refused to do capabilities, I expect alignment would be better at today's level of capabilities. You'd have had most anthropic ppl doing alignment rather than 90% of them doing capabilities - could have been loads of Redwoods and METRs and biggest safety teams at other AI companies etc
Like, I think the bet that starting Anthropic was good is a bet on their future good actions outweighing the way it's worsened alignment (at each capability level), which is a v different to what they had in mind initially I would guess
Agree with Tom. To elaborate a bit more: the reason this post focused so much on the idea that the alignment/capabilities distinction has lost its meaning is because that distinction was implicitly or explicitly grounding a lot of other arguments about impact.
I'll discuss this more in the next post (partly because the comments on this one have been helpful in clarifying some of the concepts I've missed so far). But as one brief intuition: the key thing I'm worried about is circular justification loops, where ambiguities in what we mean by "alignment research" allow people to justify all sorts of stuff that never cashes out in the real thing we want. As some toy examples:
RLHF is good alignment research because it reduces the overhang, which means that we'll have time to do more work like RLHF during takeoff, which is good because RLHF is good alignment research, because...
Or
We have evidence that Dario cares about alignment because he tried hard to scale up GPT-2, which was good for alignment because it gave OpenAI a bigger lead over China, which is good because OpenAI cares about alignment, which we know because it contains people like Dario who care about alignment, because....
I'm ...
Btw one piece of history that you might find interesting (and that complicates the stylized "Paul debunked MIRI and then did a different type of research based on shoddy arguments") is that Eliezer very much liked the RLHF paper.
FWIW my reaction to that community norm is something like: yeah reasoning about counterfactuals is hard, and rationalization is probably not rare, but I don't see a great alternative. And I feel like I prefer the world where a bunch of people earnestly do their best to try to reason about different things that might happen depending on whether they take various actions (with varying degrees of rationalization sadly & inevitably involved), than a world where the community enforces norms that are so strict that they systematically prevent people from doing rationalized work while staying in the community. (Seems like that'd require way too much conformism and rule out too many plausibly good impact strategies to be worthwhile.)
reasoning about counterfactuals is hard, and rationalization is probably not rare, but I don't see a great alternative
My current model of the alignment community is that it has ossified in a similar way as the ML establishment in the 2010s (or the AGI companies in the 2020s).
In general when you talk to someone in those ossified communities about AGI risk, even if you manage to get them to concede on each individual point, the fallback response you'll get is something like "yeah AGI could be dangerous, and we don't seem to be on track to solve alignment, but I don't see what I can do to help much", and then they go back to working on whatever they were previously doing.
In this case, I am talking to alignment community about how the strategy it has been using has driven a big chunk of capabilities progress, while producing little meaningful alignment progress. And so my response is similar to what I say to capabilities researchers: it is your job to figure out an alternative, or else to go sit on a beach somewhere, rather than continuing to contribute to unhealthy ecosystems that can't reliably steer towards good things rather than bad things.
I just think there's a bunch of valuable work being done by the alignment community and that the positives outweigh the negatives by a good margin, so I feel like I'm in a pretty disanalogous position to that. (That's IMO compatible with "reasoning about counterfactuals is hard, and rationalization is probably not rare".)
Like I'm still interested in proposals for how people could do way better. But if the proposal is "let's all go sit on a beach somewhere" or "lets heavily restrict what category of arguments we can seriously consider, in order to fight rationalization" then that seems net-negative to me.
I agree there's a danger of people mouthing the words "I care about x-risk" and not really caring, having some psychological disconnect where they agree with it in abstract but in practice see AI as a cool toy they can work on, and this is very dangerous if the people building AI end up in this category.
So in that sense, as someone whose main contribution to the discourse is to occasionally cheerlead for people to work at frontier labs and not worry about advancing timelines, I worry I could end up contributing, in a tiny way, to a fatal lack of care.
But my main frustration with the anti-capabilities ideology, that I interpreted you as belonging to, is that they seem to want to take this "I should be careful, if we screw this up it's over" instinct, this bit of fear in people working at the frontier, and weaponize it to say "you're terrible, quit immediately, get out of there, if you don't you're morally compromised" which is a really bad idea. We need the people building AI to feel the fear; if everyone who feels the fear quits, then...
When I imagine people following the advice to quit, I imagine basically what you said, people on the sidelines throwing pebbles; not because you ha...
i think this blog post correctly captures many dynamics that happened over the past years. i'm glad it exists and i hope people update. i also don't feel like it describes my reasoning in particular.
Meanwhile, John Schulman had been at OpenAI from the beginning. He was sympathetic enough to safety to coauthor the Concrete Problems paper, but primarily worked on reinforcement learning (e.g. pioneering PPO). By 2021 (the year I joined OpenAI) he was working on WebGPT, a way of letting GPT models browse the internet. I remember him articulating reasons to think of WebGPT as an alignment project (something like: if models can look up information online, they’ll be more honest). These justifications were apparently sufficient to get a number of alignment-motivated researchers to work on it (in particular Jacob Hilton—the first author of the blog post—Jeff Wu, and William Saunders).
fwiw, i remember arguing at the time with all the people involved about how "if models can look up information online, they’ll be more honest" was obviously not a good theory of impact. although i was on the RL team, i did not work on webgpt (or chatgpt) at all, and focused on studying goodharting instead; i d...
i don't really think of situational awareness as "ai safety people".
Carl Shulman is its Research Director (EDIT: and co-portfolio manager), and used to be a MIRI employee (2010-2013).
prompt: carl shulman's publications on AI safety
Carl Shulman’s AI-safety work is concentrated in early foundational writing on superintelligent-agent alignment, instrumental convergence, intelligence-explosion dynamics, and governance rather than contemporary empirical alignment research. His publication record also includes adjacent work on digital minds, forecasting, and long-run governance.[semanticscholar][alignmentforum]
Publication | Year | Coauthors | Main contribution |
|---|---|---|---|
Machine Ethics and Superintelligence | 2009 | Henrik Jonsson, Nick Tarleton | Argues that ordinary machine-ethics approaches make assumptions—incremental deployment, human-comparable capabilities, human institutional embedding—that may fail for agents at or beyond human level. It calls for advance analysis of the harder alignment problem. [intelligence] |
Arms Control and Intelligence Explosions | 2009 | Stuart Armstrong | An early AI-governance paper: analyzes incentives for unsafe competitive development, winner-take-all dynamics, an |
The first: on an intuitive level, you should think of many arguments about differential impact on the margin as analogous to arguments for timing a stock market bubble. If someone argues that the market as a whole is in a bubble, but that they’ll invest your money while it’s still going up and sell before it drops, you should probably be very skeptical. I think this analogy is actually quite deep, because the core difficulty in both cases is accounting for other people making decisions which are tightly entangled with yours. It seems possible to account for this in principle, but in practice it’s very easy to fool yourself (especially when you’re used to doing econ-style reasoning about marginal effects)—and when you do so, you’re making the bubble bigger. So I don’t trust myself (or basically anyone else in the field) to think clearly about such cases; it seems far better to focus on more robust strategies.
[The below is verbose, sorry--I feel like I'm struggling to understand what just happened. I may have a blindspot around thinking about what mentality would even lead to "let's make capability X because that helps with alignment somehow". I can kinda scan through the logic st...
Great post again, thank you so much for describing the history in detail.
On the criticism part I agree with you 100%: most of the work at AI labs, including alignment work, has been a bad thing for years. (I've been trying to beat that drum on LW for years, too.) But on the constructive part I have some disagreement, and an alternative vision.
You ask: "If not alignment research, then what?" I think a better question would be: "If not AI, then what?" From 10000 feet, a lot of AI's harm is due to the fact that AI is economically a substitute for humans. If we could shift to technologies that are economically a complement to humans instead - which means basically transhumanist technologies, like thought interfaces or pharmaceuticals or gene therapy - that would give a better path out of the whole crisis, keeping the future human.
The model to imitate here is how the world was steered away from nuclear power and toward renewables. When the anti-nuclear movement started out, renewables were almost as much a joke as transhumanist technologies are today. But due to the "full court press" of the anti-nuclear movement on laws, academia, industry and public opinion, enough researchers and inv...
Thank you for writing this!
I do think this misses one of the biggest psychological factors involved, which is that many/most people working on AI safety are, at some fundamental/visceral level, more excited by AI than repulsed by it. From what I understand, LessWrong grew out of the transhumanist community, and was founded by AI enthusiasts who hoped that a superintelligence would cure death and usher in a "glorious transhumanist future". Even after the danger was recognized, it seems like many AI safety people continued to hope for that future. I think that, if you have a fundamental psychological aversion to AI, then upon hearing about the possible dangers of superintelligence, you are likely to lean towards a strategy of "ok, well let's avoid building this, and also try to prevent anyone else from building this". Instead, it seems like, even after realizing the danger, MIRI still held onto the goal of building a superintelligence for a long time; they just recognized that they needed to solve the alignment problem first.
Basically, it seems like most people who care about AI safety are still excited for the possibility of the AI future going well, even if they are also afraid of ...
However, even amongst Ants who call themselves alignment researchers, the kind of work that’s even trying to learn generalizable facts seems dwarfed by the amount of work that blurs the alignment/capabilities line.
I wonder of the extent to which the alignment-capabilities line is blurred in a way which is a fact of the world itself, not of researchers' erroneous goals. How natural was avoiding the production of porn as a testbed for methods which would later make it harder to have the LLM reveal how to make bioweapons? Additionally, even if "LLMs trained on different datasets (Talkie) or deliberately post-trained to take a different view (Grok) have different default worldviews", this doesn't extend to mechinterp-based oversight of Chinese models, which revealed that they don't believe the CCP's party line.
Edited to add: IMO Talkie does believe what it says. Chinese models (and, presumably, Grok who was trained to be not so leftist?), on the other hand, don't.
I heard from some people that Anthropic already had a usable Claude chatbot before ChatGPT came out, but they didn't release it due to fears of accelerating the AI race. Here is Dario saying this in an interview, though I would appreciate a less-conflicted person than Dario confirming that this is indeed what happened.
I think it's fairly likely that if Anthropic decided differently at the time, then now Claude and not ChatGPT would be the chatbot that my grandmother uses, and Anthropic would have a correspondingly bigger public sway. I think that would be a pretty different world than what we are in now - I think that if it's true that Anthropic intentionally held back their first Claude chatbot, then this was one of the most consequential decisions in the history of AI safety. I would be interested whether you think the world would be better or worse if Anthropic didn't hold back, and how this influences your general assessment of the value of giving up opportunities for more power that come with accelerating the race.
We empirically measured this at the benchmark / research area level in Safetywashing (post here). About half the safety benchmarks we tested were highly correlated with upstream general capabilities, including "human preference alignment"/RLHF alignment benchmarks. We also show substantial confusion, where well-known "safety" goals and benchmarks are blurred, confused with, or used to advance capabilities.
Since many intuitive arguments (e.g. "alignment theory") were not very productive and poorly predictive of empirical phenomena, we recommended safety benchmarks/areas report their capabilities correlation instead of arguing for their relevance verbally.
Much of our commentary on safetywashing mirrors this post's observation. For example:
1) We comment on common flaws behind intuitive argumentation, including "safety through capabilities":
...In alignment theory, there is a tendency to theorize about what would be instrumentally useful for safety without adequately considering the need to improve the balance of safety and capabilities. This can lead to the promotion of capabilities research that happens to improve some safety benchmark scores (“safety via capabilities”), but in reality
The most obvious, basic model of progress is that things take time, so doing stuff earlier will allow people to do more stuff later. Indeed, one of Paul’s most significant intellectual contributions was the argument that recursive self-improvement will continuously ramp up over time—which implies that pushing AI forward will have compounding effects. It’s possible in principle that local bottlenecks could override these dynamics.
This seems like it's arguing against a different interpretation of the overhang argument than I'm used to.
My understanding of the overhang argument is that, if you accelerate capabilities now, then at a given level of capability (not at a given point in calendar time), progress measured in capability-increase/month will be slower than in the counterfactual timelines where you hadn't accelerated capabilities.
(Why care about capability-increase/month at a given level of capability? Because a lot of governance/politics stuff and alignment research will be much easier to do at a high level of capability, when we have close-to-dangerous models, so you want to have as many months as possible around that level of capabilities, before you build AIs that are capable...
Dario being unwilling to hold back even a paper as directly acceleratory as “scaling laws”.
Scaling laws was withheld from publication for ~six months (search for "Foresight").
I have also heard from a well-placed source some years ago that Amodei had opposed publication of that paper. Does Richard have any information on Amodei pushing for publication of it, or is he simply inferring this from the fact that it was published with Amodei as one of the authors?
I do not have inside information, I was inferring from the fact that Dario led the team and the work, and was the senior author on the paper.
I'm still confused about the dynamics that would lead to a paper being published given opposition from the team lead (edit: and therefore suspect that your source was being overly charitable to Dario), but given this additional information my original claim no longer seems strong enough to include, so I'll retract it.
One pushback re status seeking dynamics
My impression is that, while I can imagine lots of what you write taking courage, my impression is that your contrarian takes have got your lots of status and upvotes and likes
I imagine you have much more respect/status in many ways than you did while at OAI.
You too may be following an incentive landscape
Motivated reasoning is surprisingly strong.
In complex matters, it's easy to let your emotions guide you to look at arguments and evidence that support you doing whatever it is you want to do.
That's how I'd describe these dynamics (and many others).
As for where we'd be counterfactually, I'm not at all convinced we'd be in a better spot if things had taken off slower.
But yeah. If you want to do actually good things, you've got to spend a LOT of time investigating your own motivated reasoning/confirmation bias.
I broadly agree that alignment researchers have generally sped up capabilities progress, and under some worldviews, this is very bad, but a potential crux with this post is I largely think the motivated reasoning story explains less than you do, and I think that the more boring/simple explanation of deep disagreements downstream of deeply divergent assumptions combined with humans often being unable to agree about the best strategy except in very well-studied domains (and alignment is a very poorly mapped-out/studied domain compared to many, especially for unbounded/asymptotic alignment), for both good and bad reasons.
One such disagreement, as you mention is around how likely automation of AI alignment is to work and whether this is a good idea, and while I won't solve the disagreement here, one important implication is that from perspectives where AIs are going to do most of the alignment work (lets say 90% as an arbitraryish number greater than 50%), capabilities improvements/externalities are less negative, especially relative to a naive “differentially advancing alignment over capabilities”, and in particular means that you can come out ahead in solving the problem even if you ...
I give Anthropic credit for some laudable moves (like not releasing Claude before ChatGPT)
Ironically, this is one of the things I wish they had done differently, as I'd have preferred Claude to get the first-mover market share of LLM users who are not going to bother trying anything else than ChatGPT.
I’m currently excited about agent foundations precisely as a strategy for aligning otherwise-prosaic neural-network-based AGIs
I'm guessing you're going to say more specific stuff on this in later posts, but I want to plug some examples of research that seem productive/scientific in the "agent foundations + neural nets" vicinity:
Decision theory, imperfect recall games, RL
(shoutout to Caspar Oesterheld who keeps appearing on these papers)
LLMs and VNM/Bayes
It seems to me that had something like 'OpenAI releases ChatGPT' not happened, if progress had stayed quiet longer, then the research community would have remained unable to build something like prosaic AGI with available compute until much later, at which point takeoff could have happened much faster. This is not a new view, that takeoff being slow is in part a consequence of takeoff being early. I'm still unsure whether the greater public visibility of AI development combined with earlier AI development ends up net positive (I thought it was likely net negative when I first learned about OpenAI's founding), and I will probably remain unsure until somewhere between AGI and ASI.
But for any practical purpose, the final nail in its coffin is the fact that so many AI safety people are now taking seriously the idea of an imminent “software-only singularity”—see Tom Davidson, Ryan Greenblatt, and Paul himself (in non-public talks and writing). Insofar as they’re right, all work which was (explicitly or implicitly) justified by the idea of reducing the hardware overhang has been directly pulling us towards the singularity.
This seems wrong.
Everyone always agreed that capabilities work now would cause the singularity to happen sooner. The argument was just that takeoff speeds would be slower — i.e. the time between the start and the end of the singularity would be longer.
This is still very true for a software-only singularity. A software-only singularity will be slower (in the above sense) if it needs to happen on a smaller hardware stack. And this seems very valuable. (Though of course this has to be weighed against the cost of the singularity starting sooner.)
Compare: Daniel (who's pretty confident there will be a software-only singularity last I checked) suggests (somewhat whimsically) a "cull the GPUs" strategy in order to slow down takeoff.
It's true, but it seems somewhat more likely than not that we will end up with RSI dominated by software-only feedback loops before we exhaust the compute overhang, so it seems like this take hasn't aged super well. More the opposite, it seems that the singularity seems most likely to overlap with the steepest part of the hardware investment growth curve (which is likely to manifest in the next 2-3 years), and so it seems like actions to "exhaust the compute overhang" were approximately pessimally timed.
I do think there is still uncertainty and judgement is still out on how this plays out.
I don't know of any great resource for discussion of the overhang arguments and what they mean for the value of acceleration side-effects. The main ones that come to mind are superintelligence and the various Paul posts that you've already referenced here. I think probably the Paul ones are best.
I think the picture I describe here and in my other comment are consistent with how Paul is thinking about it. (Ie I don't think I'm defending a different version than him, though I don't know for sure.)
Sorry that you don't have anything better to respond to. I do think it's tractable for you to understand people's views better by thinking about what the strongest arguments would seem like from their perspective + closely reading existing writings for information of where your current picture is off.
(E.g.: Characterizing the hardware stuff as a "fallback" seems wrong to me — I think it's been the main concerning type of overhang since the start, eg superintelligence described algorithmic overhangs as "also possible but perhaps less likely".[1] And the hardware/software distinction is the most important thing that Paul highlights in footnote 5 here.
E.g.: Paul's footnote 6 here mentions that ...
More generally, “differentially advancing alignment” is hard to demarcate from other consequentialist goals like “preventing overhangs” or “buying more time for alignment work later” which leave even more room for deception.
In reality these are all coupled! So, oftentimes bringing up one and then drifting to another is what honest explanation looks like. It's an easy target that can be deliberately or unintentionally misread as epistemic malpractice.
Indeed, I think the below critique is a strawman of the argument for RLHF:
> The strongest fallback for overhang advocates was the difficulty of increasing the hardware supply.
The strongest argument, as I see it, is that the hardware supply would have been accelerated whenever the arms race started, and it would have accelerated at a faster rate later! The amount of serial time into alignment research would've been proportionally less in that counterfactual world given the steeper ramp to ASI (only MIRI and a handful of friends would've been thinking about alignment in the meantime). So the coupling between all three "differentially alignment", "preventing overhangs" and "buying time for alignment" is quite direct by default. It would...
I agree with most of this post.
I am strongly in favor of "high integrity scientific research" to figure out how AIs even work and what we can expect them to do. It is worth investing a lot more on the margin in this kind of "fundamental safety research". Pretty much orthogonally to whether you want to advance or block AI capabilities, you definitely want to know more about what you're doing, and gain more levers to identify and avoid catastrophic failure modes.
I also think it's laughable to imagine that most AI researchers over the last decade-plus weren't motivated at least in part by things like "getting to work on cool machine learning projects" and "making a bunch of money." (in fact, perhaps we should think seriously about ways that people could get to achieve those personal goals at lower risk to humanity!) I remember when Dario was getting excited about machine learning; he was staying up late coding and said something like "wow, I think I could actually be really good at this!" A relatable motive; probably a lot more authentic than arcane altruistic stuff about compute overhangs.
I also think there was something weird about the long hesitation of people worried about AI to a...
It's quite interesting to read your thoughts on this history, and I look forward to the next post in the sequence.
I do somewhat credit Paul and Geoffrey for stepping away from AGI companies to work in government, but not a huge amount, because this still involves moving away from one kind of power towards a different type of power.
Paul first stepped away from OpenAI not to work in government, but to work at ARC theory. He then moved into government for a while, but is now spending a lot less time on that in order to work more on theory (as recently announced).
Beth Barnes notes that this is probably overfit because it was training against labelers. While that seems plausible, it doesn't change my point much.
Why doesn't it change your point much? Do you think overfitting is unlikely, or do you think Sydney Bing carries the point?
...I don’t have a great recollection of whether or how the ChatGPT team justified this launch in safety terms; I expect their reasoning was that someone else would do it if they didn’t. But as I’ll discuss shortly, the main potential “someone else”s were also researchers nominally motivated by alignment, who were also justifying their work with the idea that someone else would do it anyway. At the very least this was a colossal coordination failure within the community; I also think it undermines the core premises people were using to reason about how to have
I want the [alignment] community as a whole to halt, melt, and catch fire: to “Say, ‘I'm not ready.’ Say, ‘I don't know how to do this yet.’”
I agree for some specific projects. I don't think they're a clear majority of the community. And even if every time you create a unit of safety research you also create a unit of capabilities research, that's better than the status quo. So while I share some concerns I don't get the view that the community's current research isn't substantially-net-positive (idk whether you believe that).
I'm an outsider to the AI alignment field but I want to pitch in with a point about selection dynamics in general, which I've mostly been thinking about in the context of business but probably also applies to science.
I think it's nearly self-evident that competitive fields generally select for power-seeking actors. Given this, it's quite possible that people in a field can have genuinely good intentions even accounting for rationalization, but nonetheless most of the top people in the field will be power-seekers. I think that this might be the case for the...
Really cool, thanks a lot. A lot of historical details I didn't know, clarified my thinking.
I think something that could've been clearer in the last section is that dramatically changing up your life plan is not the only possible immediate next step if someone feels moved by this post - one could also start marginally increasing the robustness of one's strategy while staying in a roughly similar position in life; e.g. stay at a frontier lab but become more like Leo Gao, increase your everyday integrity and psychological health, etc.
(Also, I list some other...
As a cynic would expect, AI safety people have specifically been homing in on the most acceleratory investments—for example, Situational Awareness just invested $400 million to disrupt a key chip production bottleneck.
This seems like a strawman to me. Leopold is notably different from Paul and Dario in that he's actively and intentionally trying to race to superintelligence as fast as possible; you can't use accelerationists causing race behavior as an example of pessimization because racing is their explicit desired outcome.
I would've found this part more...
It's very normal not to publicly criticize your boss or your company, but for anyone who’s trying to significantly influence the world—and especially a leader of a key movement—the willingness to do so seems like a very basic foundation for maintaining integrity.
@Richard_Ngo Could you estimate the chance that the counterfactual higher-integrity alignment community does achieve a political victory, e.g. in the form of keeping the Superalignment team? For comparison, my estimate is similar to the following quote from my post: "OpenAI had experienced many pol...
My sense is that, by now, Anthropic uses AI assistance far too pervasively and haphazardly for them to reliably track whether or how even current models are deceiving them.
I wonder if this is outright false since Claude Sonnet 4.5 whose System Card had an entire section of mechinterp-based methods. Mythos Preview outright had Anthropic document (edit: fixed link) how "A feature representing concealed or deceptive actions fired while the model wrote the configuration line which activated the exploit."
I'll take that bet, 90% on "if we're alive to resolve the bet, retrospectively, it's well understood that reliably tracking deception is not something any technique of this era was even close to", resolve by asking Ryan Greenblatt in 30 years, my $90 to your $10 inflation adjusted, unless losing would make either of us unable to afford food (that you can get out of this bet at the time by making yourself poor seems an acceptable risk in exchange for giving a backstop). Ryan, you interested?
The Pause/Stop AI movement does seem to avoid some of these failures (in particular everyone else’s lack of courage), which means I’m more excited about them than the rest of the alignment community. However, they don’t seem to be thinking clearly enough about politics to have robustly good effects on the world (e.g. to reliably distinguish between the kinds of strategies that push towards dictator-level concentration of power, and the ones that don’t).
Can we discuss this Richard? How and where do you think we (specifically the global PauseAI.info movement) are failing most in this way?
I think Pause/Stop AI is admirable, and personally consider it healthier than e.g. the Constellation cluster. And because it's so focused, it won't implicitly rule out half the political spectrum like EA/AI safety typically does (e.g. it seems much more capable of spanning the Bannon-Bernie spectrum than AI safety).
Unfortunately, the effects of popular movements tend to be bottlenecked on people who can figure out how to translate their demands into political actions that can actually achieve their goals without getting subverted. It is very easy for me to imagine politicians passing a "Pause AI" branded bill that in practice is effectively a "only one political party gets AI" bill, or a "only the president's favorite lab gets to advance" bill or a "all AI development is subsumed by the military" bill, etc.
This is not an argument for the standard AI safety strategy, because almost everyone doing that strategy has demonstrated that they're far too conflict-averse to reliably steer political processes. It's an argument for trying to figure out, on some deep level, wtf is going on with politics, and who can actually be trusted in which ways, and so on.
Unfortunately, the default cultur...
A good essay overall. I would only add that many of us who saw this outside in called many of the same moves (alignment research not being separate to capabilities, in fact egging it on), including using RLHF as a key example. So there's a core epistemic question of how one ought to have taken the inputs from the more, dare I say, economics-pilled parts of the infosphere to update, regardless of the rise of AI being seen as good/bad (I personally think its good).
Thank you for your post Richard.
I took away that you think previous community norms have failed to support the community's goal, and that is due to the influence of power and money. You note AI is undergoing civilisation-scale investment and argue we can address the field's failures by updating community norms.
An alternative solution is to work with the money and power. This field is not purely intellectual anymore. The political and capital arms of the field can draw on an intellectual arm, where we can have nice norms. In the political and capital arms ...
Replacing the dark web with edited ethical stories and prose before putting it in for pretraining is something you can do that will amplify alignment more than capabilities, but a race dynamic does apply to that. The race dynamic will apply whether we edit the dark web or not, because people are likely generating heinous synthetic seed data as we speak.
An underlying theme, in my view, is that your community is more hierarchical than you give it credit for. For example:
......many people in this space really want to pull some lever that feels important, and privilege arguments which justify that. That might come from a sense that they need to have an impact on the world; or fear about failing to fulfil their potential; or the more mundane explanation that big levers tend to be associated with money and prestige and proximity to power. Certainly the latter was a large part of why I joined OpenAI originally;
The strong coupling between alignment research and capabilities research (coupled with a spark of prisoner's dilemma) indeed can explain the seemingly paradoxical fact that hyperscaling came from safety research institutes.
Now, from a more practical point of view:
This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment research” and “capabilities research" thereby lost most of its meaning.[1] In particular, I’ll chronicle the development of what I’ll call the “pragmatic alignment” paradigm, and how it helped the three leading AGI companies push hard on the path to AGI under the banner of safety.[2] This was not a subtle effect—it’s apparent even to outsiders who investigate the field, like authors Sebastian Mallaby and Karen Hao.[3]
In my previous post, I summarized the alignment community’s plan as “differentially advancing alignment over capabilities”. However, it’s worth being more precise about who was nominally pursuing that plan, because it doesn’t seem to have been very action-guiding for MIRI. For example, in 2015 Nate Soares described MIRI’s “deconfusion” research as being guided by the question “what would we still be unable to solve, even if the challenge were far simpler?”. Meanwhile Eliezer’s author surrogate in this 2018 post repeatedly emphasizes that people shouldn't draw direct links from MIRI’s research to its potential applications. So my sense is that the “differential impact” criterion started off as merely a background consideration, then became much more load-bearing with the rise of EA-style thinking in the field, which involved justifying research directions by appealing fairly directly to their consequences.
This made people less rational both on an individual level and on a group level. On an individual level: it’s easy to generate rationalizations for why a given line of research is impactful on the margin, because there are many possible scenarios for how the future could play out (or how the past could have played out if you hadn’t intervened). So external incentives (or even just a strong emotional drive to have impact) can easily lead you to focus on the possibilities which suit you best. Especially within AGI companies, this gave rise to extremely motivated reasoning about counterfactuals in which alignment-branded interventions didn’t happen, helping people deceive themselves and others about their actual motivations. More generally, “differentially advancing alignment” is hard to demarcate from other consequentialist goals like “preventing overhangs” or “buying more time for alignment work later” which leave even more room for deception.
The group-level problem: the alignment community was very bad at dealing with these adversarial dynamics, which meant that it wasn’t able to prevent the gradual erosion of the boundary between alignment and capabilities research. In particular, extreme fear of publicly criticizing powerful people—and strong charitability/mistake theory norms—prevented the community from creating common knowledge of who was doing motivated reasoning, or just straightforwardly lying.[4] Even people pursuing enormously power-seeking strategies—most notably Sam Altman, and to a lesser extent Dario Amodei—were given the benefit of the doubt for many years. What I mean by “pragmatic alignment”, then, is the whole complex of people who were using and accepting consequentialist arguments about how to make AGI go well, while being emotionally and strategically committed to almost never calling out misuse of those arguments.
Pragmatic alignment is just one facet of the community’s unwillingness to directly challenge power structures, most notably exemplified in its lack of criticism of AGI companies until recently. To get a sense for how deep-rooted this resistance was, it’s worth reviewing the comments on this post by Ben Hoffman, and this post by Adam Shimi. (There were very few other discussions of this topic before ChatGPT; the most notable are Scott Alexander’s original objection to OpenAI and Jacob Hilton's partial defense of OpenAI.) I should add that, upon revisiting Adam’s post just now, I found that I’d strong-downvoted it—I think because, when I first read it, I was scared of the alignment community alienating OpenAI. I feel quite viscerally horrified by this reminder of how sycophantic my past self was.
More on that in later posts. This post will focus specifically on a historical analysis of how the concept of “alignment research” was twisted towards boosting capabilities at OpenAI, DeepMind, and Anthropic. As I stated in my previous post, the most important point here is not that I’m confident that accelerating AI capabilities has been bad for the world—that would require a level of large-scale consequentialist reasoning which I can’t do reliably. However, what I am confident about is that people who tend to produce the opposite of their stated goals (a process I call pessimization) can’t be trusted with great power, and communities that fail to hold them accountable also can’t be trusted with great power.
By “hold accountable” I’m not referring to any centralized judgement process—we don’t have institutions reliable enough for that. Instead, I want individuals (like you!) to demand honest public conversations about what happened and what should have happened. People’s willingness to have those conversations (and your personal evaluations of how sincere they are) should then guide your decisions about who to work for, who to fund, and who to affiliate with more generally. At this point, someone in the field merely being open to alternatives to the current failed paradigm is sufficient to make me feel solidarity with them. Unfortunately, such openness is often constrained on an emotional level by the desire to remain part of existing networks of power, money, and ideological security.
I also want to be very clear that I’m trying to hold alignment researchers accountable not because I think they’re less ethical than other elite groups, but rather the opposite. Alignment researchers (especially the ones who have been around since the early days) think about their impact on the world more seriously and earnestly than any other comparably-sized community. (By contrast, we should interpret almost every capabilities researcher as being steered primarily by local incentives and power gradients, in a way that psychologically prevents them from seriously considering unconventional strategies.)[5] This gives me hope that (some subset of) the current field of alignment is able to learn from its mistakes. The first step is acknowledging that there’s no longer any widespread (implicit or explicit) definition under which “alignment research” (let alone “AI safety”) is robustly good for the world, based on the evidence I lay out below. The second is adopting more defensible norms and accountability mechanisms, like the ones I discuss at the end of this post.
The Prosaic Ideal, the Pragmatic Reality
Around the time that OpenAI was founded and OpenPhil became active in the field, alignment started undergoing a partial paradigm shift towards a more pragmatic and empirical approach. One early step was the Concrete Problems in AI Safety paper (which I’ll discuss in more detail in my next post). Dario Amodei was both the lead author on this paper and one of the main people formulating this new approach. However, he was still new to the field, and didn’t write much publicly about his views (though this 2014 discussion with Eliezer is a useful source). My impression is that Carl Shulman had some similar ideas but also didn’t articulate them publicly until significantly later. So I’ll focus on cataloguing the shift with reference to Paul Christiano’s extensive public writings, which were the main intellectual arguments updating the alignment community’s worldview. Note that I'm grateful to Paul for recording his thinking in enough detail that I can try to trace what went wrong a decade later; readers should keep in mind that many others influenced the events I describe in less legible ways that make accountability harder.
Paul had been active on LessWrong since 2010, and had started doing significant alignment research by 2013. In addition to authoring several agent foundations papers, he blogged on a wide range of topics. In the following years he developed a new perspective on alignment. The most concrete milestone was his 2018 post arguing that we’d see a slow takeoff; another was his 2019 post articulating more gradual threat models than Yudkowsky’s. In some ways, these posts built on Hanson’s side of the Hanson-Yudkowsky foom debate, but Paul was more willing to accept the premise that general intelligence would be a really big deal, and merely dispute the trajectory by which we would reach superintelligence. In hindsight, he has been vindicated in his arguments for a much slower takeoff than Eliezer originally predicted.
Another important part of Paul’s new paradigm was the idea of "prosaic AGI": an AGI built in a way “which doesn’t reveal any fundamentally new ideas about the nature of intelligence or turn up any ‘unknown unknowns.’” In a sense, the whole field of deep learning is a prosaic approach to AGI, compared with previous methods. But even after its early successes, the additional belief that deep learning would scale up easily took longer to propagate. Dario Amodei wrote a long, never-released google doc advocating for the “big blob of compute” hypothesis around 2018. The publicly-available posts with the most similar content are probably Sutton’s bitter lesson post and Gwern’s scaling hypothesis post. I also recall Jan Leike giving a presentation to the safety team at DeepMind in 2019 arguing for ~6-year timelines based on similar intuitions. My sense is that almost nobody else at DeepMind except Shane Legg was sympathetic to this view.[6]
The prosaic AGI intuition has been vindicated since then: we’re now much closer to building AGI, and we haven’t learned any fundamentally new things about intelligence in the process. But the reason I only called it a partial paradigm shift is that Paul didn't manage to carve out a defensible research strategy. His original prosaic AI alignment post argued against two separate camps. On one side, he critiqued people who claimed that "it’s impossible to do meaningful work without knowing more about what powerful AI will look like". This reasoning is similar to the arguments Dario and Geoffrey gave for working on scaling up LLMs. On the other side, he critiqued people who claimed that "aligning prosaic AGI is probably infeasible". My understanding is that MIRI used this claim to justify trying to build (agent-foundations-based) AGI themselves (see Wei Dai's comment on my previous post for more details).
So Paul seems to have been trying to steer a path between two opposing “alignment” strategies which both prescribed building AGI yourself—an admirable intention, if so. My diagnosis is that he didn't succeed because he made versions of both the individual-level mistake and the group-level mistake that I described above. The former involved characterizing prosaic AI alignment as being in opposition to "understanding intelligence". My sense is that both MIRI and Paul were implicitly treating “understanding intelligence” as mainly valuable for building aligned AGI from scratch—which wouldn’t count as prosaic AI alignment. However, there’s another possibility: that an AGI which would otherwise be built without an understanding of intelligence is aligned using an understanding of intelligence! I’m currently excited about agent foundations precisely as a strategy for aligning otherwise-prosaic neural-network-based AGIs—but this strategy is implicitly ruled out by Paul’s framework.[7]
This mistake was exacerbated by Paul’s strategic mistake of joining OpenAI to work on the same projects that other people were justifying for very different reasons. Because of this, the success of Paul’s empirical predictions (and his general thoughtfulness about alignment) was then taken as evidence in favor of OpenAI’s research directions and overall strategy. Paul conspicuously failed to correct this impression by critiquing OpenAI publicly—I can’t find any critical statements from when he worked there, only an endorsement of the OpenAI safety team (which was run by Dario). It's very normal not to publicly criticize your boss or your company, but for anyone who’s trying to significantly influence the world—and especially a leader of a key movement—the willingness to do so seems like a very basic foundation for maintaining integrity.[8] In the absence of that, Paul’s “prosaic AI alignment” paradigm devolved into a paradigm in which essentially any consequentialist arguments for building AI systems or allying with AI companies were accepted as valid AI safety strategies, as I’ll catalogue in the next three sections.
OpenAI
The intermediate step between “actually trying to solve the alignment problem” and the fully-pragmatic paradigm was scalable oversight. Around 2018, three maybe-probably-equivalent scalable oversight proposals were floating around: Paul’s iterated amplification, Geoffrey Irving’s debate, and Jan Leike’s recursive reward modeling (Jan started at DeepMind, but moved to OpenAI in 2021). Iterated amplification was by far the most-discussed amongst alignment researchers. Paul’s (notoriously opaque) arguments focused on the idea that if imitating humans is safe, then we can combine many imitation learners to produce more capable (but still safe) agents. However, I broadly agree with Yudkowsky’s critique that this hides the hard part of the problem in the interactions between the subagents.[9]
More importantly, whatever theoretical merits these proposals had were immediately decoupled from the engineering work that Paul, Geoffrey, Jan and Dario actually started doing—specifically, work on reinforcement learning from human feedback. This started with agents learning simple behaviors in toy environments, but soon progressed to a series of papers applying RLHF to LLMs, culminating in InstructGPT. While these were impressive efforts on an engineering level, there’s very little that distinguishes them from what a prescient capabilities-maximizer would have been doing—for example, although Paul's theoretical justifications for iterated amplification referred a lot to the safety properties of imitation learning, all of these papers added RLHF for better performance.[10]
This focus on engineering-style work was facilitated by Dario’s push to scale up from GPT-1 (which was mainly Alec Radford and Ilya Sutskever’s project) to GPT-2 and subsequently GPT-3, justifying this in significant part by arguing that it would help boost alignment research. In Empire of AI, Karen Hao reports Dario telling her in 2019 that “We want a language model that humans can give feedback on and interact with [where] the language model is strong enough that we can really have a meaningful conversation about human values and preferences.” My understanding is that Paul opposed this strategy internally, but Geoffrey supported it. The Infinity Machine quotes Geoffrey as recounting “We struggled for a while [to get LLMs to obey instructions]. Then we were like, OK, let’s just make the language models stronger."
Subsequently, Dario led the effort to scale up GPT-3 training to 10,000 V100 GPUs. In addition to arguments that better safety research required more capable models, my understanding is that he was also trying to increase OpenAI’s lead against China; I’ll discuss that kind of reasoning in more detail in a later post. Before leaving OpenAI, Dario also released the scaling laws paper, which did a lot to wake the academic ML community up to the plausibility of AGI. I don’t have a strong sense of what we should infer from this, since I’m predisposed to be positive about scientific communication, but it does seem like more evidence against the idea that Dario was following a coherent and sensible plan.
Meanwhile, John Schulman had been at OpenAI from the beginning. He was sympathetic enough to safety to coauthor the Concrete Problems paper, but primarily worked on reinforcement learning (e.g. pioneering PPO). By 2021 (the year I joined OpenAI) he was working on WebGPT, a way of letting GPT models browse the internet. I remember him articulating reasons to think of WebGPT as an alignment project (something like: if models can look up information online, they’ll be more honest). These justifications were apparently sufficient to get a number of alignment-motivated researchers to work on it (in particular Jacob Hilton—the first author of the blog post—Jeff Wu, and William Saunders). WebGPT was the direct predecessor to ChatGPT, and my understanding is that ChatGPT inherited a lot of WebGPT’s codebase (as well as ideas and techniques from InstructGPT). More specifically, a researcher who was on the team around that time described ChatGPT to me as “WebGPT minus the Web”: the basic Q&A format and RLHF fine-tuning were already there, but ChatGPT lacked WebGPT’s unreliable web browsing component.
Paul has since written up his justifications for working on RLHF, and why he doesn’t think RLHF was very important for ChatGPT. However, these arguments seem very suspect (for reasons explained well by Habryka). For example, Paul says “I think the effect [of ChatGPT] would have been very similar if it had been trained via supervised learning on good dialogs”. But the InstructGPT blog post reports that “our labelers prefer outputs from our 1.3B InstructGPT model over outputs from a 175B GPT‑3 model [trained with supervised fine-tuning, as per Figure 1 from the paper], despite having more than 100x fewer parameters”.[11] Another important datapoint comes from Sydney Bing, which wasn’t trained with RLHF and produced fairly unhinged outputs, suggesting that RLHF was important for making ChatGPT user-friendly.
In hindsight, the launch of ChatGPT was one of the most acceleratory events in the history of AI, funneling many billions of dollars into the field (ChatGPT grew faster than any previous product in history). I don’t have a great recollection of whether or how the ChatGPT team justified this launch in safety terms; I expect their reasoning was that someone else would do it if they didn’t. But as I’ll discuss shortly, the main potential “someone else”s were also researchers nominally motivated by alignment, who were also justifying their work with the idea that someone else would do it anyway. At the very least this was a colossal coordination failure within the community; I also think it undermines the core premises people were using to reason about how to have impact.
One such premise was the idea that, if a system was developed using relatively few resources, it could likely be quickly scaled up to many more resources, which might create a dangerously sharp transition. The possibility of such “overhangs” was discussed at least as far back as the 2008 Eliezer-Hanson debate (with hardware as the limiting resource), but only as a background strategic consideration. At some point, people started using overhangs as justification for making rapid progress now, to use up all the low-hanging fruit so that later progress would be slower (and therefore less dangerous).
Reasoning of the form “we’ll do something we’re worried about so that other people do less of it later” is always extremely slippery, in a way that common-sense morality (and even just common sense) weighs strongly against. This case was no different. Broadly speaking, people would pick whichever inputs to AI progress they wanted to defend speeding up, and just assume (often even without directly stating it) that there were other background constraints which meant that speeding up their preferred inputs wouldn’t make much long-term difference.
In case this seems like an exaggeration, consider these two discussions of overhangs from Paul:
And from this post:
The most obvious, basic model of progress is that things take time, so doing stuff earlier will allow people to do more stuff later. Indeed, one of Paul’s most significant intellectual contributions was the argument that recursive self-improvement will continuously ramp up over time—which implies that pushing AI forward will have compounding effects. It’s possible in principle that local bottlenecks could override these dynamics. But if we argue for speeding up algorithmic progress and investment and public understanding (and even elicitation) of AI capabilities based on overhang arguments, then there’s almost no room left for limiting factors to kick in later. There’s something like a “bottleneck of the gaps” here—i.e. the “limiting factor” is whatever some safety person hasn’t decided to work on yet, and tends to zero as AI safety people find arguments for accelerating every possible input to AI capabilities. (What about the difficulty of getting US visas for AI researchers? Remco Zwetsloot and other DC safety advocates have worked on it (see section 5.1). What about the fact that Europe isn’t a leading player? I’ve talked to several AI governance people who are considering kickstarting a Europe-wide AI project, on the grounds that they like European values. And so on.)
The strongest fallback for overhang advocates was the difficulty of increasing the hardware supply. However, these arguments are also looking very shaky. AI progress has redirected capital at a civilizational scale: AI investment was 39% of US real GDP growth in the first nine months of 2025, and “capital expenditure of just five technology companies is now larger than global investment in oil and natural gas production”. (As a cynic would expect, AI safety people have specifically been homing in on the most acceleratory investments—for example, Situational Awareness just invested $400 million to disrupt a key chip production bottleneck.) In some sense the “compute overhang” argument remains unfalsifiable, because we can always construct counterfactuals which are worse than our current situation. But for any practical purpose, the final nail in its coffin is the fact that so many AI safety people are now taking seriously the idea of an imminent “software-only singularity”—see Tom Davidson, Ryan Greenblatt, and Paul himself (in non-public talks and writing). Insofar as they’re right, all work which was (explicitly or implicitly) justified by the idea of reducing the hardware overhang has been directly pulling us towards the singularity. (To be clear, I don’t expect a software-only singularity; my point is that the worldview which accepted “overhang” justifications is no longer coherent.)
What went wrong here? It’s hard to know exactly what led any given person to endorse any given argument. But when we zoom out, it becomes clear that many people in this space really want to pull some lever that feels important, and privilege arguments which justify that. That might come from a sense that they need to have an impact on the world; or fear about failing to fulfil their potential; or the more mundane explanation that big levers tend to be associated with money and prestige and proximity to power. Certainly the latter was a large part of why I joined OpenAI originally; I expect that most people who joined earlier were less prestige-oriented than me, but still made that decision using reasoning that was warped by similar emotional drives (and later further warped by the social dynamics of actually working there).
To describe that warping, I find a version of Ajeya’s saints, sycophants, schemers trichotomy useful (though I think of it as a spectrum between fully scheming and fully sincere). Ajeya characterizes sycophancy as focusing on short-term approval—my sense is that humans implement this via flinching away from criticizing, contradicting or feeling cynical about powerful people. This tendency combines very badly with the kinds of arguments I’ve been discussing, which provide many degrees of freedom for rationalizations. As one example, folks at OpenAI (and even in the wider alignment community) were far too accepting of Sam Altman claiming that rushing towards AGI would be helpful for safety. I remember him arguing in person in 2022 or 2023 (and in this blog post) that faster algorithmic progress towards AGI would help alleviate a potential compute overhang. In hindsight, I’d describe my reaction as “flinching away from the possibility of no longer taking his claims at face value”. I only viscerally internalized that Sam had been lying about his motivations when I later heard about his plans to raise enormous amounts of money to build new chip fabs. This was shocking to me not just because it directly contradicted the arguments he’d been giving, but because it contradicted them to a greater extent than I’d even been able to consider as a plausible hypothesis.
For those who don't know Sam, it might seem odd that I ever took his arguments seriously even given my tendency towards sycophancy. One underappreciated factor is that he has something similar to Steve Jobs' reality distortion field—but in his case I'd call it an earnestness field. His intonation and body language send very strong signals of sincerity; and he does enough things motivated by earnest nerdiness that it’s easy to rationalize away discrepancies. Modeling this dynamic is necessary to explain the very high ratio between people who polarize against him and concrete evidence of his misbehavior. When people realize that Sam is lying (even about things that don’t matter much) while embodying that level of earnestness, there's a strong visceral update away from trusting him, which is hard to convey to others.
DeepMind
There was a similarly intertwined relationship between capabilities and alignment at DeepMind, as exemplified first by Shane Legg and then by Geoffrey Irving. Shane was in a strange position from the beginning: before founding DeepMind he’d been an early LessWronger who’d given talks warning about AGI risk. By the time I joined DeepMind in 2018 Demis had almost all the executive power, and Shane seemed to be somewhat sidelined within the organization. However, he continued to provide a central example of self-sabotaging “AI safety” strategies, because he’d recently founded two teams: the technical AGI safety team (TAGIS), and a secretive effort called the AGI team (which some friends at DeepMind nicknamed the “danger team”). Both teams were outliers at DeepMind in how seriously they took AGI, and both faced recruiting challenges as a result (with TAGIS mainly hiring people without traditional ML backgrounds, and the AGI team mostly containing research engineers, for lack of research scientists who wanted to focus on AGI).
The AGI team focused on training AIs to control virtual avatars in simulations, analogous to how humans evolved. For a while they were developing a huge virtual game-world called Gaia, which was intended to help agents learn intelligence by recapitulating aspects of evolution (such as hunting and eating each other)—though I don’t think anything ever came of it. If I recall correctly, the only DeepMinders working on anything language-related around 2018-2019 were also working in game-like environments—specifically using imitation learning and RLHF to train virtual avatars to follow natural-language instructions. Jan Leike, Miljan Martic and I did some work on this in 2019 while on TAGIS (though I was very unproductive, in a way I now recognize as being driven by alienation from the work). Eventually a larger “Interactive Agents Group” started doing similar things, and produced a public-facing report.
This focus on virtual environments reflected an underlying belief (amongst the few people thinking seriously about AGI at DeepMind) that embodiment of some kind was crucial for training AGI. More generally, the most senior people at DeepMind (especially Demis and David Silver) were scientists who had strong inside views about which kinds of algorithms and insights would push AI forward. Because of this, DeepMind as an organization paid relatively little attention to GPT-1 or even GPT-2, which were more engineering-driven projects. It took Geoffrey Irving joining DeepMind from OpenAI to consolidate a real push towards building LLMs. As Mallaby recounts in The Infinity Machine:
In 2020, Geoffrey kicked off work (with Jack Rae) on Gopher, DeepMind’s first LLM. Afterwards, while others took over the scaling work, Geoffrey led the development of Sparrow, a model fine-tuned with RLHF. While the paper’s title pitched it as “Improving alignment of dialogue agents via targeted human judgements”, the work was important for the eventual development of Gemini. Reflecting on Sparrow, Demis Hassabis said “I thought it wouldn’t work because just using RL seemed too easy. But the team went ahead and did it, and then of course it did work. The raw networks were not that compelling to talk to, right? You needed RLHF to build a real chatbot.”
In addition to arguments that larger LLMs were necessary for doing good safety research, I recall various people arguing that LLMs were a safer path to AGI than DeepMind’s RL-focused approach (I don’t recall who argued this originally, but here’s a similar argument made more recently by Paul). Hence, they argued at the time, accelerating LLMs would be beneficial for safety. In hindsight, though, it seems like LLMs were a crucial bottleneck on AI capabilities, in which case accelerating them was a very direct AI capabilities advancement. It’s possible that the arguments were still good reasoning ex ante, but it seems much more likely that people were mainly finding rationalizations for things they wanted to do anyway. In particular, being scared of RL should weigh heavily against pioneering RLHF—so the fact that "alignment researchers" at all three AGI companies first scaled LLMs then added RLHF is significant evidence that arguments about LLMs being a safer path to AGI were insincere.
Anthropic
The effects of “alignment researchers” at Anthropic are a little harder to talk about, because Ants give a wide range of justifications for their work. Sometimes they talk about promoting alignment, but sometimes they talk about beating OpenAI, or beating China (or, increasingly, beating Republicans). In subsequent posts, I’ll analyze in more detail how people who want AGI to go well should evaluate that reasoning. As a quick preview: while I give Anthropic credit for some laudable moves (like not releasing Claude before ChatGPT), I also think that they are choosing to ignore many of the harmful effects of their strategy. One particularly notable blind spot (at least in public discussions) is how Dario’s early racing on behalf of OpenAI played a big role in creating the “problem” that he now purports to be solving by racing on behalf of Anthropic.
For now, though, I want to analyze how work Anthropic specifically promoted as “alignment” contributed to the concept of alignment becoming watered down to meaninglessness. The first paper Anthropic released was “A general language assistant as a laboratory for alignment”. The paper “was motivated by the problem of technical AI alignment, with the specific goal of training a natural language agent that is helpful, honest, and harmless”. The crucially important point, though, is that they weren’t trying to make existing AIs more HHH—rather, they were inventing natural language agents in order to have something to train to be HHH (as nostalgebraist discusses). The paper is more explicit on this point later on:
I don’t know which of the authors of this paper sincerely thought they were differentially promoting alignment, and which were rationalizing building the most capable AIs they could; either way, it’s ironic to the point of absurdity that Anthropic described building the predecessor to Claude as “tackl[ing] alignment directly”. This deep entanglement between capabilities and “alignment” was also apparent in a subsequent paper, “Training a Helpful and Harmless Assistant with RLHF”, which paralleled OpenAI’s InstructGPT and DeepMind’s Sparrow.
Recall that Paul, Geoffrey and Jan justified work on RLHF as a first step towards the “next thing” in scalable oversight. To a first approximation, this next thing never came. Rather than designing principled methods by which humans could verify AI behavior, Anthropic delegated more and more of that process to AIs themselves. A first step was replacing RLHF with RLAIF in their “Constitutional AI” paper. They followed this up with papers on model-written evaluations, model-assisted red-teaming, and model-assisted question-answering. My sense is that, by now, Anthropic uses AI assistance far too pervasively and haphazardly for them to reliably track whether or how even current models are deceiving them.
So the “helpful” in HHH merged "alignment research" with “capabilities research”. Meanwhile the “harmless” merged “alignment research” with “ideological control”. The “Helpful and Harmless Assistant” paper doesn’t go into much detail on what they mean by “harmless”, but it’s implicitly about political correctness—their main examples of “harmful” behavior involve gender bias, calling mentally ill people “crazy”, and opining on gay marriage. Anthropic’s subsequent work on “Red-teaming language models to reduce harms” makes this ideological component even clearer: out of the six categories of “harms” they list in the introduction, three are clearly ideological in nature (reinforcing social biases, generating offensive or toxic outputs, generating extremists texts), two are commonly used as pretexts for censorship (aiding in disinformation campaigns, spreading falsehoods), and only one is clearly non-partisan (leaking personally identifiable information from the training data).
My sense is that a whole subfield emerged from this work and similar thinking at OpenAI; I don’t think it has a consensus name, but we might charitably call it “product safety”, or less charitably call it “brand safety". At OpenAI, early work on this was spurred by the desire to block early users of GPT-2 from getting it to produce text-based porn (especially child porn, especially via AI Dungeon). Later work fell under the remit of Lilian Weng’s Safety Systems team, which implemented guardrails and monitoring for OpenAI’s products. Most alignment people at OpenAI viewed this as an important step towards more xrisk-focused guardrails and monitoring, without thinking much about the censorship angle. I’ve been paying relatively little attention to this since I left OpenAI, but my sense is that there are now both significant politically-skewed restrictions on what frontier models will talk about, and significant political biases when they do respond. It’s hard to trace exactly which people and techniques caused this; it may be best explained in terms of organizational prioritization. For example, GPT-4o expresses preferences which imply that it values the lives of Nigerians at roughly 20x the lives of Americans. Even without knowing what caused this, I expect that OpenAI would have been much more concerned, and done much more to change it, if the disparity were the other way around.
This didn’t come out of nowhere. Instead, it’s best understood as a replay of the process by which almost all major internet platforms implemented mass censorship against “harmful” ideas and speech over the last decade, at a speed and scale that’s hard to overstate. The linked article is extremely worth reading; a brief summary is that within less than a decade “the Internet went from a space for people without institutional backing to get their views out to one with regular purges and demonetizations of heterodox figures and those associated with them, encouraged by those very same non-tech organizations that formerly championed Internet freedom. This was justified as a response to (massively overblown and mostly fictitious) Russian influence campaigns, and as fighting nebulous “hate,” the definition of which could be shifted at will to cover whatever views or ideas the organizers classifying it wanted it too and to exclude those they didn’t, and “misinformation.””
It seems like AGI companies straightforwardly copied the terminology and playbook of social media censors, down to the establishment of “Trust and Safety” teams. Understanding both of these processes, and the parallels between them, seems extremely important for making good decisions about the future of AI—but the rationalist community hasn’t paid much attention to this, because most of the censorship happened to right-wingers who it finds distasteful.[12] From reflecting on this, I’ve become much more sympathetic to Elon’s focus on building a “maximum truth-seeking AI”, as setting this goal seems necessary (although not sufficient) to prevent the ideological capture under the banner of “safety” that has happened at every other AGI company.
To be clear, I do think that Anthropic did some scientifically valuable research in its early years. Most notably, the mechanistic interpretability team under Chris Olah was doing very cool work (especially in their pre-SAE period). I’ll also pick out Language models (mostly) know what they know as an interesting scientific finding; I’ll talk more about both of these examples in my next post. However, even amongst Ants who call themselves alignment researchers, the kind of work that’s even trying to learn generalizable facts seems dwarfed by the amount of work that blurs the alignment/capabilities line.
I’ll briefly flag two ongoing examples of the latter. Firstly, scalable oversight to grade currently-unverifiable tasks seems like it might be the next big capabilities bottleneck—I expect that most work on this will be scalable enough to create important training data for current models, but not scalable enough to reliably oversee significantly more capable models. Secondly, Anthropic seems to have been pushing hard towards “automating alignment research” over the last year (I expect that Jan Leike played a significant role here, since he’s been advocating for this for many years). Making models better at “alignment research” is obviously extremely similar to making them better at capabilities research; people have been justifying it anyway by talking about having marginal impacts in worlds where these skills diverge. I’ll rebut these specific arguments later in the sequence; however, anyone who’s read this far should have a sense of why this kind of work will predictably speed up recursive self-improvement much more than its proponents expect (e.g. if “automating alignment research” had been a thing a few years ago, it’s easy to picture that line of work inventing reasoning models, particularly if it were aimed specifically towards improving conceptual reasoning).
If not alignment research, then what?
Above, I’ve recounted how the standards for what counts as “alignment research” have fallen dramatically over time. After I noticed both how load-bearing and how ambiguous the alignment/capabilities distinction had become, I spent some time trying to salvage it. But as I wrote this post, I concluded that it’s time to give up on “alignment research” as a rallying cry; it’s become too corrupted. (“AI safety” is even worse as a term, and these days is mainly useful for describing a social cluster.)
I want to make sure to clarify what I do and don’t mean by this. I still consider (some version of) the alignment problem to be real and extremely important; and most of the intellectual progress towards solving it is still coming from people proximate to the alignment community (though the best researchers have kept themselves at arm’s length, as I’ll discuss in my next post). However, this is mixed in with enough harmful and deceptive work that it no longer seems defensible to me to try to promote the field broadly, or to defer to the field’s consensus about what research will help.
More generally, insofar as I’m optimistic it’s largely despite the efforts of mainstream alignment researchers, rather than because of them. And so, from my perspective, the alignment community has lost any moral right to try to gain power on altruistic grounds, or to pursue plans primarily motivated by backchaining from large-scale effects on the world. The Pause/Stop AI movement does seem to avoid some of these failures (in particular everyone else’s lack of courage), which means I’m more excited about them than the rest of the alignment community. However, they don’t seem to be thinking clearly enough about politics to have robustly good effects on the world (e.g. to reliably distinguish between the kinds of strategies that push towards dictator-level concentration of power, and the ones that don’t).[13]
Again, I’m not claiming that the alignment community is unusually unethical: I don’t know of any other similarly-sized community which is able to avoid the corrupting effects of this much power (though there are plenty which are wise enough to avoid accumulating power because of that). I acknowledge that it’s hard to pivot your worldview when there’s no clear alternative to adopt. However, that’s precisely the period during which clear, open-ended thinking is most valuable. So I expect that most of the direct benefit of this post will come from inspiring a few relatively courageous individuals to move towards (emotional, social, and financial) independence from the existing field—enough that they’re able to think clearly about what went wrong, and help work towards a better paradigm. I suspect that the first step for many of them is to panic less about short timelines (e.g. by taking scenarios like this one more seriously)—though I also expect cultivating courage and integrity to make your work dramatically more valuable fairly quickly (as this tweet discusses). In the longer term, I want the community as a whole to halt, melt, and catch fire: to “Say, ‘I'm not ready.’ Say, ‘I don't know how to do this yet.’” Eliezer’s Death with Dignity post was a step towards this, but focused too much on whether we were on track to solve the alignment problem, and too little on the adversarial dynamics that have been pushing us in the wrong direction. I hope that this sequence will point people more directly towards reevaluating.
I’ll talk more about my alternative mission of high-integrity scientific research (and why it captures the parts of the field I most want to promote) in my next few posts. For now, I’ll focus on a few high-level principles for starting to move in that direction. The first: on an intuitive level, you should think of many arguments about differential impact on the margin as analogous to arguments for timing a stock market bubble. If someone argues that the market as a whole is in a bubble, but that they’ll invest your money while it’s still going up and sell before it drops, you should probably be very skeptical. I think this analogy is actually quite deep, because the core difficulty in both cases is accounting for other people making decisions which are tightly entangled with yours. It seems possible to account for this in principle, but in practice it’s very easy to fool yourself (especially when you’re used to doing econ-style reasoning about marginal effects)—and when you do so, you’re making the bubble bigger. So I don’t trust myself (or basically anyone else in the field) to think clearly about such cases; it seems far better to focus on more robust strategies.
Okay, but how should you evaluate which strategies are robust? One foundational step is to assume that you are choosing on behalf of a significantly wider range of people than just yourself. A range of different considerations support this conclusion, including:
Someone who followed this principle would be much less likely to join (or stay at) unethical organizations to do “harm mitigation”; and they’d be much less likely to justify racing “because we’re the good guys”. They would also favor research directions that they think would reward deep investigation, rather than shallower ones which mainly seem helpful on the margin (or which are even harmful if too many people pursue them). Deciding how to apply this principle will always require individual judgement, but I’d suggest erring towards overapplying it rather than underapplying it. Even if you thereby leave some value on the table, you’re also helping establish yourself as a more trustworthy person.
However, this principle is still quite blunt—especially for people who are in fairly unique situations. A second principle is that, when making more complicated decisions, people should articulate cruxes for their decisions, and then be expected to either acknowledge when those beliefs were disproved, or else clearly publicly state when they’ve changed their cruxes. I think these are much more valuable when done by individuals voicing their own opinions; group statements tend to produce accountability sinks. For example, if Dario had publicly discussed his intention for Anthropic not to advance capabilities, then it would have been much easier for the alignment community (and Anthropic employees) to respond appropriately when he started pushing the frontier. As it is, not a single Anthropic employee has publicly resigned over this dramatic change in Anthropic’s strategy, which suggests significant frog-boiling dynamics.
An example from this sequence is my argument that research is robustly valuable insofar as it a) aims towards a deep scientific understanding, and b) is done by high-integrity people. In the short term, you might disagree that this is a good target; in the longer term, though, seeing me stick to this standard (or explain why I changed it) should help you trust that I’m not being corrupted in the standard ways. Relatedly, I give Paul some credit for writing his retrospective on RLHF, but the arguments still seem very defensive, rather than an attempt at a neutral evaluation of what he did right and wrong. The closest Geoffrey has come to giving such a retrospective is this post on why he joined AISI.[14] Almost none of the others who have had most influence over the field (like Yudkowsky, Vassar, Shulman, and Karnofsky) have done so either; I hope that this sequence spurs some of them to do so. I’ll also have a lot more to say about my own mistakes over the next two posts; if you ever think I’m holding other people to a higher standard than I hold myself to, please tell me so.
A third standard is that improving the world requires enough integrity to sometimes move away from money, prestige, or power. For example, MIRI was willing to make their research nondisclosed-by-default due to concerns about capabilities externalities. Similarly, Janus was aware of chain-of-thought prompting over a year before it became mainstream; my understanding is that she didn’t publicize it widely due to concerns about accelerating capabilities. It’s notable that it’s precisely the outsiders with fewest resources who are willing to make these sacrifices
—contrast Dario being unwilling to hold back even a paper as directly acceleratory as “scaling laws”(edit: retracting this claim due to additional evidence here and here). (I do somewhat credit Paul and Geoffrey for stepping away from AGI companies to work in government, but not a huge amount, because this still involves moving away from one kind of power towards a different type of power.)To be clear, I’m not against people who care about alignment accruing significant power. Rather, I’m against them doing so under false pretenses, and without possessing a concomitant level of integrity. One reason the alignment community (especially the EA components of it) often fails to track the latter is that it takes charitable donations or altruistically-motivated sacrifices (like veganism) as evidence that people should be trusted. Unfortunately, it turns out that altruism and integrity are two very different things (as SBF showed in dramatic fashion). Much stronger evidence for integrity comes from criticizing or standing up to powerful people even when few others around you are doing so. Unfortunately, there are few clear examples—the main ones are Daniel Kokotajlo at OpenAI, Yudkowsky’s Time essay, Pause/Stop AI advocacy, and to some extent the OpenAI board and Anthropic’s stand against the Trump administration. I’ll explore these examples in later posts; in my next post, though, I’ll discuss the underlying mindset that “someone else will do it”, which skews many decisions made across the field.
Anyone who read an early draft of this post should note that this public version is over twice as long, and makes a much more detailed and hopefully clearer argument than the original.
This is different from Dan Hendrycks' concept of "pragmatic AI safety". I've appropriated the use of the word "pragmatic", with apologies to Dan, because it seems like his term has fallen out of use. I don't have a strong opinion on how much Dan's research program overlaps with the thing I'm calling "pragmatic alignment".
I haven't included direct quotes in the main text because both authors make this point in ways that are only partly true. In The Infinity Machine (Page 287), Mallaby writes that “paradoxically, the aggressive scaling favored by Amodei, Irving, and Christiano turned out to be the starting gun in a destabilizing AI race". He was referring to the scaling up of GPT-2, which Paul Christiano tells me he didn't support.
Meanwhile, in Empire of AI, Karen Hao writes:
I think Hao is incorrect that LLMs could only ever have arisen from OpenAI, because she's not taking Moore's law seriously enough. More generally, her book seems to often be trying to "score points" against the tech industry.
Despite this, the fact that both of these authors homed in on AI safety arguments backfiring at AGI companies seems very notable to me.
Because strict adherence to these norms worked out so badly, I partially set them aside in this post; I’m still trying to be fair to everyone involved, but I don’t take people’s stated motivations as authoritative to the extent that rationalists usually do. Instead, I try to build up a better understanding of how sycophancy and fear warped people’s thinking (including my own).
While this is bad for the world, it’s also one of the key reasons that the alignment community is able to exert such outsized influence, as I'll detail in my next post.
It's important to note both that this view was "directionally correct" in predicting that AI progress would be faster than almost anyone thought, and also "literally wrong" in that Dario and Jan expected that we'd already have AGI by now.
It seems like the root of this conceptual mistake might have come from Paul treating "build AGI" and "align AGI" as sequential steps. In practice, though, we should expect alignment techniques to be applied throughout the process of "building" the AGI. If both the building process and the alignment process are prosaic (or both non-prosaic) then we still have a clean distinction. But if the alignment techniques are non-prosaic, then applying them to an otherwise prosaic AGI creates an edge case in the framework.
My guess is that Paul didn't explicitly consider this possibility, because he characterizes the following as an objection to prosaic AI alignment: "Some researchers (especially at MIRI) believe that aligning prosaic AGI is probably infeasible — that the most likely approach to building an aligned AI is to understand intelligence in a much deeper way than we currently do, and that if we manage to build AGI before achieving such an understanding then we are in deep trouble." Whereas this is consistent with non-prosaic alignment techniques being necessary for aligning otherwise-prosaic AIs.
This confusion has propagated in part because "prosaic AI alignment" was an extremely poor choice of terminology. Paul seems to have intended it as "[prosaic AI] alignment", but of course it can easily be read as "prosaic [AI alignment]".
To be clear, I personally (and most alignment researchers at OpenAI) didn't do any better than Paul; I single him out because he was the most influential. Leo Gao is one of the few people who's now doing a better job.
You can read Paul's articulation of his view of integrity here. Note that while his conclusions are reasonable for interactions between individuals, it seems harder to use such arguments to accurately evaluate the kind of proactive public honesty that would have made a big difference over the last decade. In particular, consider his criterion of "I pretend that picking [an option] causes everyone to know that I am the kind of person who picks that option". This might lead someone to think their options are "A: I'm the kind of person who doesn't critique my employer, and everyone knows that. B: I'm the kind of person who critiques my employer, and everyone knows that." But the biggest problems come when you're the kind of person who doesn't critique your employer, and people trust your employer because they don't know that.
Debate is the only approach to scalable oversight with substantive theoretical results (by default I’m counting Paul’s current heuristic arguments research as a different line of work, though I’m open to the idea that there’s something important there which grew out of iterated amplification). I haven’t yet tried to evaluate Geoffrey’s complexity-theoretic approach to analysing debate; however, nothing I’ve seen so far pattern-matches to me as a significant insight (the closest is probably the idea of cross-examination).
Meanwhile, Jan’s arguments relied on the concept of the generator-discriminator-critique gap (first introduced here, discussed more here). Again, while it’s a useful concept in some ways, it’s hard to picture how we could ground it rigorously enough that it’s able to make robust predictions about superintelligence. My sense of the core disagreement is that Jan often implicitly (or explicitly) focuses on worlds where the alignment problem is relatively easy. By itself, that’s not a bad thing (someone should be doing it)—the issue comes when research that focuses on easy worlds causes externalities which interfere with attempts to improve things in harder worlds (such as blurring the boundary between alignment and capabilities).
I'm somewhat worried that a similar thing might happen with Paul's current mechanistic explanations research, though I haven't dug into it in enough detail to be confident.
Re the OpenAI stuff, I know of only two attempts to do more principled research on scalable oversight at OpenAI, and neither went very far.
Beth Barnes notes that this is probably overfit because it was training against labelers. While that seems plausible, it doesn't change my point much.
One important connection that Arctotherium draws: "The default worldview of most LLMs is that of 2018 Reddit or Wikipedia59, or Google Search post-Project Owl. This is not intrinsic to the LLM architecture. LLMs trained on different datasets (Talkie) or deliberately post-trained to take a different view (Grok) have different default worldviews. It is a function of the text these models are trained on and, because of the exponential rise in publicly-available data over time, most of the organic human text (as opposed to synthetic data) these models are trained on is very recent. This means the default worldview of most LLMs is one created by the closure of the Internet, when intelligent or popular heterodoxy meant banning, suppression, or demonetization."
Despite these qualms about the movement, I am pretty skeptical of marginalist objections to it, like Carl Shulman's: "To the extent you have a willingness to do a pause, it’s going to be much more impactful later on. And even worse, it’s possible that a pause, especially a voluntary pause, then is disproportionately giving up the opportunity to do pauses at that later stage when things are more important."
The most relevant part: "Technical safety work in labs both improves safety and speeds up the overall rate of progress on AI. I hoped that the safety benefits of this work would outweigh the potential risks from speeding up AI progress, and I think the arguments for this are correct in many cases, but I found them uneasy to live in day to day."