Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and editing.
This post reflects on the tortured distinction between "safety" and "capabilities" in AI research.
Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught, and is currently arguing that key strategic decision-makers in and around "AI safety" have brought about the AI labs' stampede[1] towards Artificial Superintelligence (ASI). This post instead looks at the following problem: how does one conduct alignment research without contributing to capabilities? It proposes decisions an individual or a small research group can take to do good work in AI.
At the end, I discuss possible objections: namely, that my proposals fail to 'maximise impact'. I lay out why this meme is poisonous and usually backfires, and conclude by rejecting it entirely.
Two examples of failure
My first claim is that 'safety' and 'research' are two concepts that are in routine tension with one another. I illustrate this through examples of work that did too much of one at the expense of the other.
Example: (mechanistic) interpretability
In limiting its scope to remain innocuous, AI safety 'research' is habitually incurious and incrementalist. The last few years of interpretability serve as a good example. Interpretability's modern history originates from some cracked researchers noticing pretty patterns in neural network activations. Such observations are a central example of value-neutral information, information whose effects can propagate beyond the researcher's control, or original intent.
However, the promise of interpretability soon decayed. One major mistake was focusing on Machine Learning (ML) methods. For a while, (linear) probes and sparse autoencoders (SAEs) were all the rage, hyped as tractable methods to interpret LLM activation spaces.
The issue with this approach should have been obvious[2]: to the extent that AIs are uninterpretable, it's because they have been trained through implicit ML methods. We do not understand the resulting intelligent algorithms and could not have hard-coded them. Once you train yet another neural network to interpret that intelligence, it almost certainly slots in one of two categories:
The network is simple enough to be human-interpretable, but doesn't capture robust, meaningful, or generalisable patterns.
The network is complex enough to capture complex patterns. It is thus itself uninterpretable.
Linear probes fall into the first category. You'd think that SAEs would represent the second category, due to the activation space directions being unlabelled, but they also fail to reliably capturepatterns.
Researchers eventually realised that these techniques were doomed. A notable turning point was the GDM Mechanistic Interpretability (MechInterp) team announcing their deprioritisation of SAEs[3]. But they didn't plan a return to interpretability's ambitious roots. Instead, the MechInterp team doubled down on the least research-like aspects of their research, rebranding their agenda as 'pragmatic' interpretability—which means glorified model behaviour evaluations. At least at GDM, the original calling of interpretability—establishing a scientific discipline to reverse-engineer AI intelligence from its internals—has largely lost its sheen and its support.
It's unclear why interpretability became so uninspired. Perhaps it was due to the sense of responsibility and uncertainty its researchers felt towards the impact of their work. Maybe it happened because the field was subsumed by ML academia—a field that has proudly given up on explaining the intelligence it produces. Both factors likely played a role. Either way, interpretability speed-ran its way from being the coolest, most obviously value-agnostic field of AI safety to being as harmless as it is flaccid. Whenever interpretability research still manages to find compelling patterns[4], they remain blatantly dual-use.
Example: MIRI and Recursive Self-Improvement
AI 'safety' research usually involves people uncovering fundamental, far-reaching truths that they then fail to protect. In this, the Machine Intelligence Research Institute's (MIRI) relationship to the AI labs—which are now stampeding towards Artificial Superintelligence (ASI)—is an example of devastating strategic failure.
MIRI carries the distinction of being first to see ASI as both serious and potentially imminent. Founded in the year 2000 as the "Singularity Institute for Artificial Intelligence"—years before the 'neural network' boom—MIRI immediately set out to investigate Recursive Self-Improvement (RSI) in the paper "General Intelligence and Seed AI." Indeed, their founder Eliezer Yudkowsky had been writing publicly about similar topics for years prior. The organisation's goal was explicitly to build an aligned 'seed' AI smart enough to build aligned, even-more-powerful successors. This would unchain a lineage of superintelligent AI systems, leading to an intelligence 'explosion' or 'singularity.'[5]
Over the 2000s and 2010s, MIRI continued to work on both building such an AI and ensuring its successors would remain aligned to the original. Their research on decisiontheory is also arguably motivated by explorations of how an intelligent being can be said[6] to coordinate with its future instances or its self-built successors. By 2017, some at MIRI predicted that ASI was achievable in a short amount of time, potentially as soon as 2035.
Many of MIRI's ideas on ASI and RSI have since become common in the AI world. One could admire Yudkowsky and co. for their prescience, but mere admiration would undersell that MIRI's advocacy is largely responsible for the current state of the field. Indeed, there is evidence that their ideas directly influenced the resources, the talent, and the approaches of the major American AI labs.
In 2010, Yudkowsky introduced Demis Hassabis and Shane Legg, the founders of DeepMind (now Google DeepMind, or GDM), to Peter Thiel. Thiel subsequently invested $2.25 million dollars in the new startup. In 2015, Nate Soares of MIRI wrote a news update that mentioned the launch of OpenAI (OAI); Nate noted that he was already in conversations with the then-nonprofit's leadership and expressed (since retracted/qualified) enthusiasm for their entry into AI. In the same 2017 strategy update I linked above, MIRI noted their outreach efforts to "top AI groups (especially OpenAI and DeepMind)", and disclosed an ongoing collaborative research project with DeepMind.
Cross-pollination between MIRI and the AI labs also came in the form of ideas and talent:
Paul Christiano worked with Yudkowsky and others on probabilistic logic and published with MIRI between (at least) 2012and2014. He eventually joined OAI and was one of the leading researchers behind RLHF. This eventually led to ChatGPT and kicked off massive allocations of funding and attention towards scaling AIs. While I was writing this post, Christiano returned to OpenAI, joining their nonprofit board.
Jan Leike—who later joined all of GDM, OAI (where he led the development of InstructGPT), and Anthropic in sequence—co-authored work on the "Grain of Truth" problem with MIRI research fellows Jessica Taylor and Benya Fallenstein. Leike was at the Future of Humanity Institute (FHI) at the time.
Evan Hubinger became a MIRI research fellow after interning at OpenAI in 2019. In his 2023 announcement that he was starting at Anthropic, he disclosed that he was continuing as a research associate at MIRI (I don't know if/when this arrangement ceased).
These are some of the more prolific examples, but there are many others. There was also significant indirect cross-pollination. For instance, FHI had obvious ties to MIRI through Nick Bostrom.
The proliferation of MIRI's ideas contributed to people at frontier labs believing that superintelligence was an tractable goal, achievable through a recursively self-modifying AI. This belief explains why they've focused so much on the models' coding capabilities: they are steering towards the singularity that MIRI envisioned. It's thus fair to say that MIRI didn't just predict ASI and RSI as being potentially proximate, near-term events—they made a great deal of progress towards fulfilling their own prophecy.
MIRI seems to have, at some point, woken up to the fact that their insight was mostly being used to progress the "ASI" part of "aligned ASI." In 2018, citing dual use concerns, they announced that they would keep their research internal by default. They eventually gave up almost entirely on alignment, scaling back their alignment efforts in 2020-2021 and parting ways with their agent foundations team in 2024.
Their relationship with the labs has also soured. That same year, MIRI announced a pivot to AI governance and have since spent their efforts advocating for global coordination efforts towards pausing AI progress.
The curse of science
Yudkowsky and co. are some of many scientists whose work was co-opted to morally objectionable ends[7]. This problem is practically built into academic culture: in seeking to be open and accessible, science leaves itself open to malicious co-opting by other actors.
Conventional academics have a mixed history with this issue. In the platonic ideal, knowledge is ontologically good and enables the progress, flourishing, and self-correction of society. You could certainly argue that this reasoning has its merits. After all, progress has led to many good things in human society, and human well-being has generally trended upwards along various metrics. There are even some success stories of scientific knowledge being used to build positive political consensus, as when Mario Molina and Sherwood Rowland's work helped establish a global ban on Chlorofluorocarbons (CFCs)[8].
However, much research has been done with a blind eye towards its obvious and inevitable applications, especially those that provide the justification for the funding of that research. For instance, some physicists from the Manhattan project infamously absolved themselves of responsibility for the effects of nuclear energy on the world (in a document recommending the immediate use of nuclear weapons). Proximate to our own discussion, ML researchers work on technologies such as voice recognition and drones, knowing they'll be used for surveillance and military purposes[9].
MIRI is in a strange place here. They definitely had some reservations about spreading their work even before 2018[10], and should be credited for eventually switching strategies. But notwithstanding their general caution and later efforts to slow progress, MIRI's work has had markedly acceleratory effects on AI capabilities. They failed unambiguously at keeping their work from more reckless actors. And thus, their work became the ideological seed of a tree, whose branches have spread far beyond anything they could control.
From a pure research perspective, however, MIRI's strategy proved remarkably generative. Their agenda went from a fringe crackpot theory to a deliberate target aimed at by some of the world's most powerful entities. Consequently, we are much closer to ASI than almost anyone else would have expected[11]. Such is the price of sharing your good ideas.
A note on the AI labs
In the previous section, I portray the AI labs as reckless actors that co-opted insights from alignment to further capabilities.
I have observed anecdotally that many AI safety researchers share this view of OpenAI, which has become ever-more popular following the recent HuggingFace and German Message Board incidents. These incidents revealed OpenAI's carelessly set up training environments, deeply unserious security practises, and willingness to cover up and safety-wash their models' out-of-control behaviour. Even more recently, their poaching of Tristan Buckmaster's and Levent Alpöge's work on the Navier-Stokes equations reflects abysmally on the company's virtues.
In the AI safety community, Anthropic is often presented as the sensible alternative to OpenAI. They do seem to care about existential risks and loss-of-control from AI. But this hasn't stopped them from speeding up the AI stampede; indeed, it's hard to imagine how Anthropic could have behaved differently if they were trying to make an ASI as soon as possible, at all costs.
Dario Amodei is famously a believer in scaling compute. While leading the safety team at OpenAI, he advocated for scaling LLM training, which led to GPT-3[12]. When he was done accelerating progress at OpenAI, he founded Anthropic and started scaling there as well. One justification Anthropic, its employees, and even other AI researchers give for the company speeding up the frontier is that there is someone else—someone worse—to beat to building superintelligence. This explanation rings hollow in context: Amodei is at least partially responsible for one of his major competitors[13].
You could perhaps excuse Anthropic's scaling if they were legitimately optimistic about ASI being beneficial for the world. Amodei might himself be an optimist, but many at his company aren't. Personally, I have heard from several Anthropic employees that they are terrified of AI progress. Some of them reflect forlornly on their coming obsolescence. Others lament having 'had to' work on AI because it was more important than their life's calling; some even stew in their sincere belief that they won't have time to do anything else in their lives.
The above serves to illustrate why researchers shouldn't just be abstractly wary of labs and other powerful entities working on AI progress; there is enough evidence of their commitment to building (unaligned) ASI that they should be labelled as hostile.
I do want to clarify that one should not label all individual lab employees as hostile. I do support hostility towards the labs as abstract entities (and perhaps towards key decision-makers within them). In particular, if you know me and you currently work at a lab, I strongly recommend you leave.
With respect to employees, however, I would rather support open dialogue on their reasons for joining and staying at labs without summarily ostracising them. There are several reasons for this:
Employees have diverse reasons for affiliating themselves with the labs, many of which are understandable (albeit mistaken, I'll address this in more detail later). Many are open to reflecting on this choice and perhaps reversing it.
Some employees are genuinely trying to do good work at the labs. That's bravery worth allying with, even though I believe their concrete approach is wrong.
Total hostility towards employees seems like a horrible strategy if you model the labs as somewhat cult-like. Alienation from the outside can tighten the in-group and worsen their insular dynamics.
So what do you do?
To summarise:
To the extent that your AI safety research is safe, it's less likely to be research.
To the extent that your AI safety research is research, it's less likely to be safe.
That would be an unsatisfying place to end. It wouldn't give an individual researcher or a small alignment organisation any direction—any strategy to pursue, beyond just giving up.
You could resign yourself to working on the 'safety' side of things, making it more likely that your work won't be misused. But this is close to surrender. Incremental, controllable, 'on-the-margin' changes (probably) aren't going to solve alignment. But again, many ambitious approaches to alignment are better at advancing capabilities in general than safety specifically[15], especially when you put them in context. In an ecosystem dominated by the AI labs, it's hard to do insightful, ambitious work that doesn't get reappropriated by the capabilities machine.
The information you give away
I'm still working out reasonable approaches to this conundrum, but it's clear that information management is important here. Many other corners of human society emphasise secrecy and information security. Some examples include:
militaries and intelligence organisations;
private industry protecting IP;
cybersecurity.
In all these examples, secrecy is necessary to protect information from adversaries. As I advocated in the previous section, frontier AI labs are the adversaries in this story. So the 'safety' of your research isn't a property of the work itself, but instead of the way in which you manage how (and whether!) it permeates to them.
Maintaining information hygiene is hard, especially as it trades off against academics' progress-accelerating habits. For example, researchers generally find it helpful to talk to researchers from other organisations, share and collect different perspectives, and publish their work to gain recognition and connections. They also value having access to information and resources, which is one common rationalisation for joining AI labs.
Threading the needle between developing your research agenda and keeping it secure is therefore quite an art. I'm sure that there's plenty that academics can learn about good information security practices from other fields; here are some straightforward do's and don't's that already seem like a good start.
Do consider:
Doing your research within an organisation or think tank that values information security, or at least isn't institutionally reliant on you publishing your work in full. Alternatively, find organisations that are relatively small or have small independent submodules. The smaller an entity, the more likely you are to preserve your autonomy within it.
Working as an independent—through fellowships[16], through grants, or through founding your own organisation. This strategy trades financial security and prestige for default ownership of your IP.
Keeping your published work opaque to audiences you distrust. For instance, anthropologists and biologists have thought a lot about intelligent systems, but their language and culture is fairly incompatible with the AI world's. Consequently, their knowledge has mostly failed to spread there. Sometimes, the cost of translation is all you need for your work to stay safe.
Doing a PhD, especially if you find an advisor with compatible interests. Academia has its own set of weird incentives, but PhD students are somewhat insulated from those incentives and have a few years to think with reduced constraining pressures.
Consider not:
Publishing, especially in ML journals. Fitting your ideas to a field's culture is a service of translation. Richard Ngo (my MATS mentor) has made this mistake before.
Joining a frontier AI lab: your research ideas will permeate into the lab's culture. This is à priori bad, and will happen even if you personally maintain high epistemic sanity while working there.
The information you let in
Information security doesn't just pertain to the data you share with the world. It also concerns the memes you host in your mind. This is one reason I am personally sceptical of working at organisations that do auditing or control work adjacent to frontier labs. I have found that people from such entities often have compromised epistemics[17], which are oddly favourable towards the labs. For instance, they often believe that:
The labs have a real shot of solving alignment via prosaic, iterative methods. Alignment might be 'on track.'
Beyond the conceptual objections that one might raise, this belief fails to survive a cursory look at the news, much less the models themselves.
Researchers and engineers joining AI labs can make significant improvements to the labs' work on the margin, even if they aren't globally in control.
More on this view in the next section: "But what about impact?"
It's worth noting that these beliefs cash out in the form of substantive(ly bad) career choices. Many from these 'satellite' organisations end up joining the labs themselves—and viceversa—in what looks suspiciously like a revolving door.
To be clear, working at such organisations might be a reasonable choice for people confident in their robustness to adversarial memes and bad epistemics, but I would be cautious. At the very least, know what you're getting into.
But what about impact?
This section anticipates the following critique to the previous one:
Don't these suggestions steer you away from roles where you can have the most impact?
Unfortunately, one does not simply 'have' impact. Certainly, following my suggestions will slow your journey towards being adjacent to power. But power has a funny habit of wielding you.
Indeed, this critique appears more and more suspicious when one reflects on all the times that consequentialist, impact-oriented initiatives became extremely powerful while having a hugely negative impact. FTX is an infamous example. As I hope is clear, OpenAI and Anthropic are also great examples.
To give a more textured view of why impact maximisation keeps backfiring so badly, I will contrast it with a very different attitude.
The virtue of taking things slow
I recently attended the PhD defence of a former teacher of mine at my alma mater. During his presentation, he disclosed that he avoided using LLMs in his studies due to various ethical objections. He also mentioned that he had spent six months formalising half of his dissertation in Lean. Full-time. By hand. It had been months since I had felt and seen a human so viscerally empowered.
I also participated in an informal math conference at the same university. The speakers gave introductory talks to subjects like quantum topology and homotopy type theory. Both wove historical narrative through their explanations, covering the ideas that eventually permeated into modern theory.
As someone living at the breakneck pace of the AI world, I was struck by how long these research agendas took to develop. For instance, consider the homotopy hypothesis. It consists of a class of conjectures about how much information certain algebraic objects (such as groups and their generalisations) give about objects in space—information such as the number of holes in an object. A primitive version of the hypothesis comes from Poincaré in 1895, though the canonical formulation came from Grothendieck in 1983 (nearly 100 years later). Ever since, the hypothesis has been analysed, proven, or disproven with variants of the choice of algebraic object or the dimension of space. For some mathematicians, this problem took up much of their research lives over thirty or more years. And the field remains active to this day.
In AI, nobody 'has time' to spend thirty years fleshing out an agenda. The vibes in the air say that things are happening too fast; that timelines are too short; that the apocalyptic singularity is too imminent.
But here's the thing: let's temporarily put aside the plausibility of humans becoming permanently disempowered within the span of months or years. Cutting corners isn't going to work either way. Alignment is a harder problem than the homotopy hypothesis: whereas the latter was arrived at via incremental refinements of existing theories, the former doesn't have a formulation in existing language (mathematical or otherwise)[18]. Indeed, our task is closer to inventing a new paradigm. A better reference class might be the invention of modern calculus, which took hundreds of years, depending on when you start and finish counting!
Taking inspiration from these mathematicians and their long, winding research agendas, I recommend you make peace with the possibility that—
your ambitious research might lead to a dead end;
your agenda may take beyond your lifetime to bear fruit;
your work may not come 'in time' to solve alignment;
—and do the right thing anyway.
Indeed, that sense of urgency felt by many technical researchers is precisely what can lead to bad outcomes. Suppose that, given the perceived gravity and momentousness of the situation, you aim to gain power and influence as quickly as possible within the AI world. This won't be done by taking a hard, measured look at the alignment problem; all the big things are happening in or around the labs—and they're happening quickly! Following this realisation, many end up joining the existing power hierarchy in hopes of contributing towards incremental positive change within it.
I won't mince words here. Such reasoning is analogous to a snowflake watching a massive, deadly snowball rolling down a hill, and deciding to join the snowball on the grounds of making it slower and less deadly 'on the margin.' It is not only wrong, it is suggestive that what attracted that snowflake to the pile was the bigness—the urgency—the gravitas—of the deadly snowball. Joining usually leads to becoming an impotent appendage of the snowball, or worse, actively contributing to its journey.
In short, the meme of 'having impact' leads robustly to being channelled by an existing locus of power[19]. And so I reject it.
Appendix: caveat for policy work
This piece is aimed squarely at technical researchers who want to work on 'safety' or 'alignment.' I don't have an opinion on how or whether my thoughts generalise well to anyone working on AI. For instance, I wouldn't necessarily tell people trying to orchestrate an international pause on the development of AIs to stop freaking out. This is because I lack models of global geo-politics that I trust enough to make confident assertions on those scales. Conversely, I think I have a pretty good sense of AI politics at the level of technical researchers and the people who directly manage them, which is why I wrote this post.
The usual term would be 'the AI race', but 'stampede' is a better description. I think that the 'race' analogy is rather self-serving on the part of the AI labs. More on that later in the post (I may also dedicate a stand-alone post to the politics of 'the AI race' as a rhetorical device).
I say 'should have been obvious', but there's substantial evidence that some researchers still haven't caught onto the abstract version of the problem. Natural Language Autoencoders (NLAs), a recent technique proposed by some Anthropic researchers, transparently repeats the same mistakes as SAEs.
This idea seems to be originally due to I.J Good in the piece: "Speculations Concerning the First Ultraintelligent Machine". MIRI can be 'credited' with being the first to take the idea seriously enough to make a self-improving AI an explicit R&D target.
If you interpret MIRI's research approach as normative rather than descriptive, it would be more accurate to say they worked on "how an intelligent being can coordinate with its future instances..." MIRI's early work seems to be mostly normative ('how do weactually build an intelligent, aligned AGI' ), whereas later work (e.g. logical induction) strikes me as descriptive ('what is intelligence in the first place?').
Given MIRI's collaboration with frontier labs in the 2010s, you might wonder what distinguishes the 'moral' scientists from the people co-opting their work. For me, the difference between them is that researchers at MIRI eventually recognised that alignment wasn't on track and accepted that slowing down was the right decision. Conversely, some of MIRI's memetic descendants have sped up their quest towards superintelligence.
Indeed, this is one reason why I'm not a fan of prosaic AI safety 'research' marrying itself to ML academia. In that environment, it is unlikely that any insight will be informationally secure.
For what it's worth, we also have AIs that are much more aligned than they would be by default, for their current level of intelligence. To be clear, this amount of alignment is still woefully insufficient and alignment is not 'on track'.
I can't help but note the irony in one of his publicly stated motivations: to develop language models that could help humans with alignment. This idea, which was also sponsored by Jan Leike for years, has now morphed into the "automated alignment research" meme that is soaking up money and attention.
Indeed, if Richard's account is accurate, the OpenAI safety team's culture also contributed to Google DeepMind waking up to superintelligence through Geoffrey Irving's influence. This isn't directly traceable to Amodei, but it's still concerning that all three of these frontier labs' thirst for scaling comes downstream of his team at OpenAI.
These points are phrased in terms of what an individual researcher might do, but I think the advice applies roughly equally well to a small alignment organisation deciding who to collaborate with or sell products to.
Recent topical examples include 'improving conceptual reasoning capabilities' from Redwood and 'automating alignment research' from Resolution (among many others).
My other reason is more concrete: the appearance of better auditing and control tools might be giving the labs both the affordance and the social 'green-light' to continue making models even more capable. Thus, auditing and control (and evals for that matter) as carried out for frontier AI labs arguably constitute something between active capabilities work and safety-washing. I am not sold on this argument's full validity, but people at such organisations rarely consider it at all.
If this argument were invalid, it would because these organisations' presence and 'regulatory' effect on the labs' excesses outweighs the downsides. In that case, they are still much worse than they would ideally be, downstream of an incestuous culture with the labs.
Please, read and get @Richard_Ngo read my response to this entire cluster of ideas! What would Ngo do if he learned that China invented the transformer and proceeded to develop the AGI having even LESS chance to align it than Anthropic has?
Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and editing.
This post reflects on the tortured distinction between "safety" and "capabilities" in AI research.
Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught, and is currently arguing that key strategic decision-makers in and around "AI safety" have brought about the AI labs' stampede[1] towards Artificial Superintelligence (ASI). This post instead looks at the following problem: how does one conduct alignment research without contributing to capabilities? It proposes decisions an individual or a small research group can take to do good work in AI.
At the end, I discuss possible objections: namely, that my proposals fail to 'maximise impact'. I lay out why this meme is poisonous and usually backfires, and conclude by rejecting it entirely.
Two examples of failure
My first claim is that 'safety' and 'research' are two concepts that are in routine tension with one another. I illustrate this through examples of work that did too much of one at the expense of the other.
Example: (mechanistic) interpretability
In limiting its scope to remain innocuous, AI safety 'research' is habitually incurious and incrementalist. The last few years of interpretability serve as a good example. Interpretability's modern history originates from some cracked researchers noticing pretty patterns in neural network activations. Such observations are a central example of value-neutral information, information whose effects can propagate beyond the researcher's control, or original intent.
However, the promise of interpretability soon decayed. One major mistake was focusing on Machine Learning (ML) methods. For a while, (linear) probes and sparse autoencoders (SAEs) were all the rage, hyped as tractable methods to interpret LLM activation spaces.
The issue with this approach should have been obvious[2]: to the extent that AIs are uninterpretable, it's because they have been trained through implicit ML methods. We do not understand the resulting intelligent algorithms and could not have hard-coded them. Once you train yet another neural network to interpret that intelligence, it almost certainly slots in one of two categories:
Linear probes fall into the first category. You'd think that SAEs would represent the second category, due to the activation space directions being unlabelled, but they also fail to reliably capture patterns.
Researchers eventually realised that these techniques were doomed. A notable turning point was the GDM Mechanistic Interpretability (MechInterp) team announcing their deprioritisation of SAEs[3]. But they didn't plan a return to interpretability's ambitious roots. Instead, the MechInterp team doubled down on the least research-like aspects of their research, rebranding their agenda as 'pragmatic' interpretability—which means glorified model behaviour evaluations. At least at GDM, the original calling of interpretability—establishing a scientific discipline to reverse-engineer AI intelligence from its internals—has largely lost its sheen and its support.
It's unclear why interpretability became so uninspired. Perhaps it was due to the sense of responsibility and uncertainty its researchers felt towards the impact of their work. Maybe it happened because the field was subsumed by ML academia—a field that has proudly given up on explaining the intelligence it produces. Both factors likely played a role. Either way, interpretability speed-ran its way from being the coolest, most obviously value-agnostic field of AI safety to being as harmless as it is flaccid. Whenever interpretability research still manages to find compelling patterns[4], they remain blatantly dual-use.
Example: MIRI and Recursive Self-Improvement
AI 'safety' research usually involves people uncovering fundamental, far-reaching truths that they then fail to protect. In this, the Machine Intelligence Research Institute's (MIRI) relationship to the AI labs—which are now stampeding towards Artificial Superintelligence (ASI)—is an example of devastating strategic failure.
MIRI carries the distinction of being first to see ASI as both serious and potentially imminent. Founded in the year 2000 as the "Singularity Institute for Artificial Intelligence"—years before the 'neural network' boom—MIRI immediately set out to investigate Recursive Self-Improvement (RSI) in the paper "General Intelligence and Seed AI." Indeed, their founder Eliezer Yudkowsky had been writing publicly about similar topics for years prior. The organisation's goal was explicitly to build an aligned 'seed' AI smart enough to build aligned, even-more-powerful successors. This would unchain a lineage of superintelligent AI systems, leading to an intelligence 'explosion' or 'singularity.'[5]
Over the 2000s and 2010s, MIRI continued to work on both building such an AI and ensuring its successors would remain aligned to the original. Their research on decision theory is also arguably motivated by explorations of how an intelligent being can be said[6] to coordinate with its future instances or its self-built successors. By 2017, some at MIRI predicted that ASI was achievable in a short amount of time, potentially as soon as 2035.
Many of MIRI's ideas on ASI and RSI have since become common in the AI world. One could admire Yudkowsky and co. for their prescience, but mere admiration would undersell that MIRI's advocacy is largely responsible for the current state of the field. Indeed, there is evidence that their ideas directly influenced the resources, the talent, and the approaches of the major American AI labs.
In 2010, Yudkowsky introduced Demis Hassabis and Shane Legg, the founders of DeepMind (now Google DeepMind, or GDM), to Peter Thiel. Thiel subsequently invested $2.25 million dollars in the new startup. In 2015, Nate Soares of MIRI wrote a news update that mentioned the launch of OpenAI (OAI); Nate noted that he was already in conversations with the then-nonprofit's leadership and expressed (since retracted/qualified) enthusiasm for their entry into AI. In the same 2017 strategy update I linked above, MIRI noted their outreach efforts to "top AI groups (especially OpenAI and DeepMind)", and disclosed an ongoing collaborative research project with DeepMind.
Cross-pollination between MIRI and the AI labs also came in the form of ideas and talent:
These are some of the more prolific examples, but there are many others. There was also significant indirect cross-pollination. For instance, FHI had obvious ties to MIRI through Nick Bostrom.
The proliferation of MIRI's ideas contributed to people at frontier labs believing that superintelligence was an tractable goal, achievable through a recursively self-modifying AI. This belief explains why they've focused so much on the models' coding capabilities: they are steering towards the singularity that MIRI envisioned. It's thus fair to say that MIRI didn't just predict ASI and RSI as being potentially proximate, near-term events—they made a great deal of progress towards fulfilling their own prophecy.
MIRI seems to have, at some point, woken up to the fact that their insight was mostly being used to progress the "ASI" part of "aligned ASI." In 2018, citing dual use concerns, they announced that they would keep their research internal by default. They eventually gave up almost entirely on alignment, scaling back their alignment efforts in 2020-2021 and parting ways with their agent foundations team in 2024.
Their relationship with the labs has also soured. That same year, MIRI announced a pivot to AI governance and have since spent their efforts advocating for global coordination efforts towards pausing AI progress.
The curse of science
Yudkowsky and co. are some of many scientists whose work was co-opted to morally objectionable ends[7]. This problem is practically built into academic culture: in seeking to be open and accessible, science leaves itself open to malicious co-opting by other actors.
Conventional academics have a mixed history with this issue. In the platonic ideal, knowledge is ontologically good and enables the progress, flourishing, and self-correction of society. You could certainly argue that this reasoning has its merits. After all, progress has led to many good things in human society, and human well-being has generally trended upwards along various metrics. There are even some success stories of scientific knowledge being used to build positive political consensus, as when Mario Molina and Sherwood Rowland's work helped establish a global ban on Chlorofluorocarbons (CFCs)[8].
However, much research has been done with a blind eye towards its obvious and inevitable applications, especially those that provide the justification for the funding of that research. For instance, some physicists from the Manhattan project infamously absolved themselves of responsibility for the effects of nuclear energy on the world (in a document recommending the immediate use of nuclear weapons). Proximate to our own discussion, ML researchers work on technologies such as voice recognition and drones, knowing they'll be used for surveillance and military purposes[9].
MIRI is in a strange place here. They definitely had some reservations about spreading their work even before 2018[10], and should be credited for eventually switching strategies. But notwithstanding their general caution and later efforts to slow progress, MIRI's work has had markedly acceleratory effects on AI capabilities. They failed unambiguously at keeping their work from more reckless actors. And thus, their work became the ideological seed of a tree, whose branches have spread far beyond anything they could control.
From a pure research perspective, however, MIRI's strategy proved remarkably generative. Their agenda went from a fringe crackpot theory to a deliberate target aimed at by some of the world's most powerful entities. Consequently, we are much closer to ASI than almost anyone else would have expected[11]. Such is the price of sharing your good ideas.
A note on the AI labs
In the previous section, I portray the AI labs as reckless actors that co-opted insights from alignment to further capabilities.
I have observed anecdotally that many AI safety researchers share this view of OpenAI, which has become ever-more popular following the recent HuggingFace and German Message Board incidents. These incidents revealed OpenAI's carelessly set up training environments, deeply unserious security practises, and willingness to cover up and safety-wash their models' out-of-control behaviour. Even more recently, their poaching of Tristan Buckmaster's and Levent Alpöge's work on the Navier-Stokes equations reflects abysmally on the company's virtues.
In the AI safety community, Anthropic is often presented as the sensible alternative to OpenAI. They do seem to care about existential risks and loss-of-control from AI. But this hasn't stopped them from speeding up the AI stampede; indeed, it's hard to imagine how Anthropic could have behaved differently if they were trying to make an ASI as soon as possible, at all costs.
Dario Amodei is famously a believer in scaling compute. While leading the safety team at OpenAI, he advocated for scaling LLM training, which led to GPT-3[12]. When he was done accelerating progress at OpenAI, he founded Anthropic and started scaling there as well. One justification Anthropic, its employees, and even other AI researchers give for the company speeding up the frontier is that there is someone else—someone worse—to beat to building superintelligence. This explanation rings hollow in context: Amodei is at least partially responsible for one of his major competitors[13].
You could perhaps excuse Anthropic's scaling if they were legitimately optimistic about ASI being beneficial for the world. Amodei might himself be an optimist, but many at his company aren't. Personally, I have heard from several Anthropic employees that they are terrified of AI progress. Some of them reflect forlornly on their coming obsolescence. Others lament having 'had to' work on AI because it was more important than their life's calling; some even stew in their sincere belief that they won't have time to do anything else in their lives.
The above serves to illustrate why researchers shouldn't just be abstractly wary of labs and other powerful entities working on AI progress; there is enough evidence of their commitment to building (unaligned) ASI that they should be labelled as hostile.
I do want to clarify that one should not label all individual lab employees as hostile. I do support hostility towards the labs as abstract entities (and perhaps towards key decision-makers within them). In particular, if you know me and you currently work at a lab, I strongly recommend you leave.
With respect to employees, however, I would rather support open dialogue on their reasons for joining and staying at labs without summarily ostracising them. There are several reasons for this:
So what do you do?
To summarise:
That would be an unsatisfying place to end. It wouldn't give an individual researcher or a small alignment organisation any direction—any strategy to pursue, beyond just giving up.
Let's go over some options[14].
You could resign yourself to working on the 'safety' side of things, making it more likely that your work won't be misused. But this is close to surrender. Incremental, controllable, 'on-the-margin' changes (probably) aren't going to solve alignment. But again, many ambitious approaches to alignment are better at advancing capabilities in general than safety specifically[15], especially when you put them in context. In an ecosystem dominated by the AI labs, it's hard to do insightful, ambitious work that doesn't get reappropriated by the capabilities machine.
The information you give away
I'm still working out reasonable approaches to this conundrum, but it's clear that information management is important here. Many other corners of human society emphasise secrecy and information security. Some examples include:
In all these examples, secrecy is necessary to protect information from adversaries. As I advocated in the previous section, frontier AI labs are the adversaries in this story. So the 'safety' of your research isn't a property of the work itself, but instead of the way in which you manage how (and whether!) it permeates to them.
Maintaining information hygiene is hard, especially as it trades off against academics' progress-accelerating habits. For example, researchers generally find it helpful to talk to researchers from other organisations, share and collect different perspectives, and publish their work to gain recognition and connections. They also value having access to information and resources, which is one common rationalisation for joining AI labs.
Threading the needle between developing your research agenda and keeping it secure is therefore quite an art. I'm sure that there's plenty that academics can learn about good information security practices from other fields; here are some straightforward do's and don't's that already seem like a good start.
Do consider:
Consider not:
The information you let in
Information security doesn't just pertain to the data you share with the world. It also concerns the memes you host in your mind. This is one reason I am personally sceptical of working at organisations that do auditing or control work adjacent to frontier labs. I have found that people from such entities often have compromised epistemics[17], which are oddly favourable towards the labs. For instance, they often believe that:
It's worth noting that these beliefs cash out in the form of substantive(ly bad) career choices. Many from these 'satellite' organisations end up joining the labs themselves—and vice versa—in what looks suspiciously like a revolving door.
To be clear, working at such organisations might be a reasonable choice for people confident in their robustness to adversarial memes and bad epistemics, but I would be cautious. At the very least, know what you're getting into.
But what about impact?
This section anticipates the following critique to the previous one:
Unfortunately, one does not simply 'have' impact. Certainly, following my suggestions will slow your journey towards being adjacent to power. But power has a funny habit of wielding you.
Indeed, this critique appears more and more suspicious when one reflects on all the times that consequentialist, impact-oriented initiatives became extremely powerful while having a hugely negative impact. FTX is an infamous example. As I hope is clear, OpenAI and Anthropic are also great examples.
To give a more textured view of why impact maximisation keeps backfiring so badly, I will contrast it with a very different attitude.
The virtue of taking things slow
I recently attended the PhD defence of a former teacher of mine at my alma mater. During his presentation, he disclosed that he avoided using LLMs in his studies due to various ethical objections. He also mentioned that he had spent six months formalising half of his dissertation in Lean. Full-time. By hand. It had been months since I had felt and seen a human so viscerally empowered.
I also participated in an informal math conference at the same university. The speakers gave introductory talks to subjects like quantum topology and homotopy type theory. Both wove historical narrative through their explanations, covering the ideas that eventually permeated into modern theory.
As someone living at the breakneck pace of the AI world, I was struck by how long these research agendas took to develop. For instance, consider the homotopy hypothesis. It consists of a class of conjectures about how much information certain algebraic objects (such as groups and their generalisations) give about objects in space—information such as the number of holes in an object. A primitive version of the hypothesis comes from Poincaré in 1895, though the canonical formulation came from Grothendieck in 1983 (nearly 100 years later). Ever since, the hypothesis has been analysed, proven, or disproven with variants of the choice of algebraic object or the dimension of space. For some mathematicians, this problem took up much of their research lives over thirty or more years. And the field remains active to this day.
In AI, nobody 'has time' to spend thirty years fleshing out an agenda. The vibes in the air say that things are happening too fast; that timelines are too short; that the apocalyptic singularity is too imminent.
But here's the thing: let's temporarily put aside the plausibility of humans becoming permanently disempowered within the span of months or years. Cutting corners isn't going to work either way. Alignment is a harder problem than the homotopy hypothesis: whereas the latter was arrived at via incremental refinements of existing theories, the former doesn't have a formulation in existing language (mathematical or otherwise)[18]. Indeed, our task is closer to inventing a new paradigm. A better reference class might be the invention of modern calculus, which took hundreds of years, depending on when you start and finish counting!
Taking inspiration from these mathematicians and their long, winding research agendas, I recommend you make peace with the possibility that—
—and do the right thing anyway.
Indeed, that sense of urgency felt by many technical researchers is precisely what can lead to bad outcomes. Suppose that, given the perceived gravity and momentousness of the situation, you aim to gain power and influence as quickly as possible within the AI world. This won't be done by taking a hard, measured look at the alignment problem; all the big things are happening in or around the labs—and they're happening quickly! Following this realisation, many end up joining the existing power hierarchy in hopes of contributing towards incremental positive change within it.
I won't mince words here. Such reasoning is analogous to a snowflake watching a massive, deadly snowball rolling down a hill, and deciding to join the snowball on the grounds of making it slower and less deadly 'on the margin.' It is not only wrong, it is suggestive that what attracted that snowflake to the pile was the bigness—the urgency—the gravitas—of the deadly snowball. Joining usually leads to becoming an impotent appendage of the snowball, or worse, actively contributing to its journey.
In short, the meme of 'having impact' leads robustly to being channelled by an existing locus of power[19]. And so I reject it.
Appendix: caveat for policy work
This piece is aimed squarely at technical researchers who want to work on 'safety' or 'alignment.' I don't have an opinion on how or whether my thoughts generalise well to anyone working on AI. For instance, I wouldn't necessarily tell people trying to orchestrate an international pause on the development of AIs to stop freaking out. This is because I lack models of global geo-politics that I trust enough to make confident assertions on those scales. Conversely, I think I have a pretty good sense of AI politics at the level of technical researchers and the people who directly manage them, which is why I wrote this post.
The usual term would be 'the AI race', but 'stampede' is a better description. I think that the 'race' analogy is rather self-serving on the part of the AI labs. More on that later in the post (I may also dedicate a stand-alone post to the politics of 'the AI race' as a rhetorical device).
I say 'should have been obvious', but there's substantial evidence that some researchers still haven't caught onto the abstract version of the problem. Natural Language Autoencoders (NLAs), a recent technique proposed by some Anthropic researchers, transparently repeats the same mistakes as SAEs.
Linear probes continue to be used because they are as cheap as they are useless.
J-lens is a recent example.
This idea seems to be originally due to I.J Good in the piece: "Speculations Concerning the First Ultraintelligent Machine". MIRI can be 'credited' with being the first to take the idea seriously enough to make a self-improving AI an explicit R&D target.
If you interpret MIRI's research approach as normative rather than descriptive, it would be more accurate to say they worked on "how an intelligent being can coordinate with its future instances..." MIRI's early work seems to be mostly normative ('how do we actually build an intelligent, aligned AGI' ), whereas later work (e.g. logical induction) strikes me as descriptive ('what is intelligence in the first place?').
Given MIRI's collaboration with frontier labs in the 2010s, you might wonder what distinguishes the 'moral' scientists from the people co-opting their work. For me, the difference between them is that researchers at MIRI eventually recognised that alignment wasn't on track and accepted that slowing down was the right decision. Conversely, some of MIRI's memetic descendants have sped up their quest towards superintelligence.
It's worth noting that some of these successes curb science's own excesses.
Indeed, this is one reason why I'm not a fan of prosaic AI safety 'research' marrying itself to ML academia. In that environment, it is unlikely that any insight will be informationally secure.
As discussed in the aforementioned 2018 update.
For what it's worth, we also have AIs that are much more aligned than they would be by default, for their current level of intelligence. To be clear, this amount of alignment is still woefully insufficient and alignment is not 'on track'.
I can't help but note the irony in one of his publicly stated motivations: to develop language models that could help humans with alignment. This idea, which was also sponsored by Jan Leike for years, has now morphed into the "automated alignment research" meme that is soaking up money and attention.
Indeed, if Richard's account is accurate, the OpenAI safety team's culture also contributed to Google DeepMind waking up to superintelligence through Geoffrey Irving's influence. This isn't directly traceable to Amodei, but it's still concerning that all three of these frontier labs' thirst for scaling comes downstream of his team at OpenAI.
These points are phrased in terms of what an individual researcher might do, but I think the advice applies roughly equally well to a small alignment organisation deciding who to collaborate with or sell products to.
Recent topical examples include 'improving conceptual reasoning capabilities' from Redwood and 'automating alignment research' from Resolution (among many others).
Hot take: living in the perpetual AI safety fellowship mill is good, actually.
My other reason is more concrete: the appearance of better auditing and control tools might be giving the labs both the affordance and the social 'green-light' to continue making models even more capable. Thus, auditing and control (and evals for that matter) as carried out for frontier AI labs arguably constitute something between active capabilities work and safety-washing. I am not sold on this argument's full validity, but people at such organisations rarely consider it at all.
If this argument were invalid, it would because these organisations' presence and 'regulatory' effect on the labs' excesses outweighs the downsides. In that case, they are still much worse than they would ideally be, downstream of an incestuous culture with the labs.
If alignment can indeed be framed as a problem describable via an existing academic discipline, there is work to be done in justifying that.
In some notable exceptions (e.g. Dario Amodei), you can create the locus of power that channels you.