I've been a follower since fairly early on: disagreed with you both then and now on a range of topics. I've been unimpressed with the precision of your language at times: I was being unfair then, I suspect at least partially because of our differences. This piece seems fair and accurate to me.
Thanks for writing this! Before it got eaten by the "AI safety community", this was a website about rationality—the art of achieving a map that reflects the territory and using the map to plan to achieve one's goals. I think a central reason that the project to improve human rationality failed so abjectly is because too little attention was paid to how the achieving-goals part could come into conflict with the map-accuracy part (because deception is often useful for achieving goals): as time has gone on, epistemic rationality has increasingly been forgotten in favor of seeking the "locally optimal discursive posture" (as you so aptly put it) in the service of AI safety. And the first step towards getting the epistemic rationality back would be accounting for what changed, rather than pretending it's always been this way—for example, by writing up the first-person intellectual history of the strategic forces that covertly shaped one's public writing, as you've done here. I would love to see more "AI safety" people do the same.
I hope you will see this apology as a product of fastidiousness and high integrity, rather than some admission of a yearslong effort at deception.
I think this is a false dichotomy that conflates absolute and relative standards. For example, you write that "the notion that [you] have not acted with intellectual integrity is not quite fair" because almost none of your fellow SB 1047 opponents agreed with you about AI risk (such that even as it was, you got "EA plant" accusations). But that seems less like a defense of your integrity and more like a claim that too much integrity would have come at an unacceptable cost to your political goals. Well, sure. There's no law of physics that says that there can't be a social environment that punishes integrity. In such a fallen world, doing as much deception as you need to in order to achieve your political goals, but feeling vaguely bad about it such that you fess up later is a product of relative fastidiousness and high integrity—but that doesn't mean that no deception occured. "Deception" is about a speaker sending communication signals that predictably decrease the accuracy of listeners' beliefs; whether the speaker had no better alternatives available doesn't play into it.
One of the most corrosive effects of this dynamic is not the harm of the deception itself, but self-deception about what higher integrity would even look like, as people faced with a conflict between honesty and winning reason, "Well, I'm a good person, and I did what I had to do given the incentives, so what I did can't be dishonest."
You write you "do not think anything I have ever written about AI safety or risks constitutes a lie." But as I've explained in a previous essay response to Eliezer Yudkowsky on this website, not-lying turns out to be a surprisingly weak standard: natural language has so many degrees of freedom that it's not that hard to arbitrarily push on listeners' beliefs while only using sentences that permit a true interpretation, simply by, e.g., "[leading] with arguments that [the speaker] believe[s] would be more palatable to a given audience at the expense of arguments that [they] believe[ ] might be more important".
At this point, some might be skeptical of the purported existence of a higher standard of integrity: what would that mean, concretely? How can there be more to honesty than just not lying? On this topic, I recommend in the strongest terms reading and meditating on two posts from this website's founding texts: "The Bottom Line" and "A Rational Argument".
Briefly: once you've decided what conclusion you want to argue for, that conclusion is already right or wrong. Searching for additional arguments for that fixed conclusion might make you more persuasive, but they can't make the conclusion more true. The arguments that matter are the ones that determine which conclusion you're motivated to argue for. Anything you come up afterwards that lacks the power to change your bottom line is in some important sense dishonest, even if every sentence is true. If the actual reason you care about open weights is liberty, then your arguments about geopolitics and diffusion are fake insofar as you wouldn't be talking about them if the geopolitical or diffusion concerns had pointed the other way.
It's a terrifyingly ambitious standard for anyone to aspire to—but just hearing it articulated has deeply changed the course of my life. Even this late in the timeline, I think it could change the world—if only anyone could remember.
I appreciate you coming here to say this. I wish more people would engage with communities in the place that they meet.
Thanks for writing this, Dean!
When I was reading your post on the inevitability of self-sovereign AI systems, I was telling myself, “yes, this looks quite likely”, but I was also asking myself some questions:
should not we expect that self-sovereign AI systems and communities of those systems will eventually be capable of non-saturating recursive self-improvement (RSI)?
do we have any plans to make sure that the necessary “good behavior” properties would be preserved and made stronger during those RSI processes, rather than be diluted and gradually disappear?
are we trying to rely on the premise that non-self-sovereign agents hosted by the leading labs would still dominate capability-wise because we’ll help them more and that that would (perhaps) enable better overall outcomes in terms of the properties of the overall ecosystem?
These were some questions I was trying to ponder…
I am not sure if you think much about RSI in connection with all this; would be very curious to learn your thoughts on how this aspect might interplay with everything else.
the focus of the piece is how to design institutions that incentivize self-sovereign AI toward pro-social activity
You can't. It can self-improve at the speed of software and you can't. It will have a higher growth rate than you. Like the US economy outgrowing the Argentinian economy. The institutions you'll set up will be like international institutions trying to stop the US from doing stuff to South America.
Ultimately there are four options. Either we build AI that's good to us by its nature, or we improve our capabilities so we can stand up to AI, or we stop AI, or we die. That's it, these four. We can't design institutions that will keep binding AI when it overtakes us. Human law, money, courts? It will shrug and do what it wants. You'd better start swimming or you'll sink like a stone.
Long-time reader. Happy to accept an olive branch. Your longposting fits right in here!
However I do recommend you use section headers for long posts like this when on LW instead of Substack, see this recent post for an example. Makes it easier to understand the thrust of a piece, track its flow, and jump back to important sections as desired.
(Also, blockquotes for quotations of a full paragraph or more.)
Appologies in advance if this is pedantic. You focus a a lot on what you believed and when. Do you have any concrete public statements which make these views clear? What you cite in the OP seems quite ambiguous to me, but perhaps I am missing something.
Reasoning about difficult and complex topics is naturally difficult and complex, and as I result I think it is the most understandable thing in the world that someone who is making an admirable attempt to do that may stumble and make mistakes. In light of that, it is no great sin in my opinion to be mistaken or to have allowed less-than-perfect reasoning creep into ones analysis. It is also a rare thing to see people ackowledge their mistakes openly. This post is refreshing in that regard and I appreciate that aspect of it.
At the same time, I think it is also somewhat common and understandable that people will sometimes re-interpret their own statements and beliefs in hindsight in a way that is flattering and suggests that they were "right all along!" when facts on the ground change. You could imagine the person who "know all along" that the housing market was sure to collapse, but seems to have conveniently start saying so in 2009. For this reason, I think it would be helpful to have some additional context that more clearly demonstrates what you were thinking. It seems like the OP is kind of asking the reader to look past what you actually said in favor of the interpretation you give now in the OP. I get that with a large body of work it is complicated because you won't always have expressed yourself perfectly and you may have focused on a particular aspect of the topic while leaving some of your views on other aspects less well explored in your public writings. I think it would be helpful to understand what you think are the clearest statements of your views at a given time, which you wrote at that time.
i did not read this in full because my skim mostly saw words directed at people looking to adjudicate fault or reputation, which felt like they would be a bit too boring for me to read. that said, i appreciate that you exist and am thankful for your efforts and conduct.
Longtime lurker, first-time poster.
I want to address a section of a recent essay of mine that has gotten some attention within the AI safety community. The main topic of the essay is what Dawn Song et al. call self-sovereign agents, or AI agents that are independent actors in the world. At the end of the essay, I say that I feel I haven’t spoken about this topic over my 2.5 years of writing with sufficient candor, and that I think this critique applies to others in the AI policy community–particularly the parts of it that tend to manifest themselves in Washington, Sacramento, and Albany–in other words, the parts of the AI safety world that are most involved in hands-on AI policy work. I attribute this primarily to a desire to remain “within the Overton Window,” or to not sound “crazy” within the halls of power, and I assert that others in my profession have made this same calculation.
I believe–and have believed for three years–that self-sovereign AI as I describe it in my essay would likely happen on our current trajectory. That being said, the essay takes pains to distinguish between “self-sovereign” AI and truly “rogue” AI, and the focus of the piece is how to design institutions that incentivize self-sovereign AI toward pro-social activity. Furthermore, I explain in the piece that I have considerable uncertainty about the size and noticeability of self-sovereign agents. While I believe their existence is almost surely inevitable, I do not think it is inevitable that they will proliferate in vast numbers or pose a non-manageable threat to society. Nonetheless, at the end of the essay, I apologize for having not described this phenomenon with sufficiently visceral rhetoric, nor have I spent enough time working on its solutions.
This passage was characterized by some in the rationalist and AI safety communities as an admission to a failure in integrity, or to “material misrepresentations” of my views. The best critical response–and frankly some of the best analysis of my work I have seen to date–was from X user @MelancholyYuga of the Substack A Goodly Measure. Throughout, I’ll respond to their arguments, but I also want to speak more broadly about how I’ve approached the numerous political tradeoffs I have encountered in my time in this profession.
I do not think anything I have ever written about AI safety or risks constitutes a lie. At times I softened words, and at times I did not. At times I led with arguments that I believed would be more palatable to a given audience at the expense of arguments that I believed might be more important, and at times I pushed my audience well beyond their zone of comfort. In my worst moments, especially very early in my writing career, I have taken unfair potshots at the AI safety community and their beliefs–statements which I retracted almost two years ago. Some have deemed this a failure of integrity, but I think it would be better described as a failure in judgment, with the benefit of hindsight to boot. You, however, can be the ultimate judge.
I apologize in advance for my first entry here being an exercise in navel-gazing, but it is what I was asked to do. This essay is not of interest to you unless you specifically care about what I thought at various points over the last few years of AI development and why I made specific decisions that I made. That is an extraordinarily niche topic and it should have a small audience. For the few who do care, though, I hope this is useful.
–
I first encountered the notion of human-level AI in 2003 or 2004, as I recall, surfing the internet as a young child. Yudkowsky was among the names I remember encountering, along with Vinge and Kurzweil. It was not a major area of interest for me (I was 12 in 2004), but it was “on my radar,” to the extent that anything can be on a 12-year-old boy’s radar. It was mostly a “big if true” kind of thing. AI recurred in my thinking a few times prior to the deep learning revolution, especially around 2008 or so when I became interested in Marx. Still, other areas of technology interested me much more.
Midway through college I remember reading about AlexNet. It was a minor update for me. But then I kept seeing more and more interesting results coming out of these “deep neural networks.” With each year, I paid more attention to developments in deep learning. By 2018, I was a regular LessWrong lurker and increasingly convinced that AGI would occur during my lifetime. I began speaking about it more with friends, family, and coworkers, though still only occasionally. These were hobby interests; my career at the time had little to do with technology.
I remember trying to use GPT-2 to classify municipal health and safety regulations by level of strictness in 2019 (it did not work), and I remember using some of the early writing tools that used GPT-3 on the API. I began to think more about the political theory implications of AI, especially during the pandemic, though this remained decidedly one of many side interests, and like many at the time these thoughts were mixed in with contemplations of the political theory of digital technology more broadly, not AGI per se.
GPT 3.5, then, was an update for me not in the sense of “wow, computers can talk,” but, “ah, OpenAI has found a way to make computers (relatively) coherent most of the time. It turned out there really was understanding in there. AGI is closer than I realized.” And so I decided in 2023 that I was going to pivot my entire career and begin writing about AI.
I spent most of that year going into serious technical depth. Training and fine-tuning toy models, reading papers, listening to lectures and podcasts. I did not say a word until I had developed a point of view I believed was my own.
Here is what I had concluded by the end of 2023:
Human-level AI seems likely to occur by 2035, and probably more like 2030. It’s going to be the most important thing that has ever happened in the history of technology. We don’t know how to align or even understand it. By default, it will be used to concentrate power, probably in neither human nor AI hands alone, but in some combination of both. The reason that it won’t concentrate in human hands alone is that humans will not be able to understand or control it. The reason it won’t concentrate in AI hands alone is that AIs will not be able to seize unilateral power due to (1) the decentralized nature of knowledge in the world and (2) the fact that the universe and the human world alike are both more resistant to change than we think; there is always more continuity than discontinuity, and this prior is remarkably well-supported by both history and other domains of science. These factors will make it likelier that AIs collaborate with existing human institutions and the incentives those institutions produce, which in the worst case scenario could result in economic arrangements where individual humans are robbed of agency, political voice, and economic freedom (in AI vernacular, this outcome would now be called Gradual Disempowerment, after the excellent paper of the same name, but that concept hadn’t crystallized for me in 2023).
I therefore found myself broadly persuaded about the importance of alignment risk, catastrophic misuse, and concentration of power, but also quite skeptical of some of the most extreme projections about job loss, radical and imminent alterations to the physical world, and the like. In particular, I have argued since the very beginning that competition over scarce resources–with the ultimate scarcity being capital–would be the ultimate primary bottleneck on the ability of any actor–AI, human, or a hybrid–to achieve the kinds of changes that some AI forecasters anticipated in such short timelines. The operation of market processes themselves, in other words, would naturally slow down certain kinds of transformations, especially in the physical world. I was, and am, an extinction-risk skeptic, but very much persuaded that any number of negative outcomes are possible.
In addition, I thought that the AI safety community had a tendency to discount the sophistication and adaptability of extant market processes, or to model superintelligence as an entity operating above those market processes rather than within them. The labor and capital markets are smarter than you think and might produce surprising adaptations.
Neither of these two critiques suggest that AI will go especially well for individual humans–I have never believed we are guaranteed good outcomes “by default” and in that sense I have never been a blind AI optimist. The trouble is that I am also deeply skeptical of the idea that regulation by do-gooders will secure a good outcome either; indeed, regulation really is a major concentration of power concern. This notion, combined with my general Burkean prior, led me to be reluctant to get on board with sudden changes to the regulatory system predicated on assumptions that all seemed rather more contestable to me than what I perceived the AI safety community to think.
This set of projections says nothing about the balance of power between AIs and humans, only that AIs and humans will persist in some kind of power dynamic. Importantly, this view does not deny superintelligence. It likely denies the notion of a monolithic superintelligence, but it does not deny the potential existence of machines smarter than humans or the risks inherent therein. Neither does the first full essay of Hyperdimensional, which Melancholy Yuga cites as one of my pieces that denies AI risks:
“Deep learning, the broad approach underlying virtually all successful AI models today, has improved rapidly in the past decade, and we don’t really understand how it works… But we lack a grand theory of what makes it all work. It’s similar to how we discovered steam as a source of energy long before we understood much about the science of thermodynamics.
Ambitious efforts have been made to further our understanding of these systems (a field known as interpretability), and important advances have been made in just the past few months. But those advances have lagged the capabilities of frontier AI models, and we are nowhere close to understanding the inner workings of something like GPT-4. That means we don’t understand why it “lies,” why it sometimes memorizes the text it is trained on (the basis of the New York Times’ recent lawsuit against OpenAI), or how to ensure it is robust against attacks…
There’s no reason to think that humans are nature’s upper bound on any capability we have: We have already created machines that can move faster than us, are stronger than us, and indeed, are smarter than us in some ways. I don’t dispute the notion that AI systems that are superior, or at least comparable, to humans in yet more dimensions are coming soon.”
Keep in mind that this is the first major Hyperdimensional essay. Does this sound like a person who began his career with an intent to misrepresent the nature of AI risks to readers, or even to downplay them, given that at that time I anticipated I would be fighting battles against AI regulation?
Now, Melancholy Yuga also quotes me as saying AI safety advocates have a tendency to express “ambient vague anxiety” based “purely in speculation.” These are not misquotes, and I think these critiques get at real problems in AI safety discourse, albeit imperfectly and with too much coarseness. I fully admit that in this piece and some of my other early writing, I was unfair to the people I then might have called “doomers.” Let me now talk about why my back was up at that time.
–
There is one more belief I had developed by the end of 2023: that the Biden Administration is using AI safety as a guise to exert control over AI, causing civil society, courts, and the American public to surrender our free-expression rights in the interest of protecting our ‘safety,’ and thereby nudge American toward tyranny. More broadly, I believed that every U.S. federal administration will have this incentive.
I thought in particular that the Administration’s rhetoric on AI’s catastrophic risk potential would be used in combination with frameworks like the Blueprint for an AI Bill of Rights and, cross-jurisdictionally, with the European Union’s AI Act, to enforce extra-legal broad speech limitations on AI systems. Here is a passage from the Blueprint:
“Protection against algorithmic discrimination should include designing to ensure equity, broadly construed. Some algorithmic discrimination is already prohibited under existing anti-discrimination law. The expectations set out below describe proactive technical and policy steps that can be taken to not only reinforce those legal protections but extend beyond them to ensure equity for underserved communities even in circumstances where a specific legal protection may not be clearly established.”
Overall, the Blueprint set up a (nonbinding) mechanism of broad state intervention into AI’s substantive outputs on issues that had nothing to do with the catastrophic risks I was concerned with. But I harbored doubts that courts would push back on the free-speech violations I believed the Blueprint enabled if–as I expected–AI would also constitute national-security threats to the United States. Courts tend to defer to the Executive during national-security emergencies, and it was my anticipation of just such an emergency, combined with documents like the Blueprint, that truly terrified me. You must understand this: I was a conservative with a long history of skepticism and fear of the administrative state, who was learning to become AGI-pilled, and this seemed like a clear and present danger to me.
[I later learned that the people in the Biden Administration who wrote documents like the Blueprint and the people who worried about AGI risk were distinct and at times warring factions. This was a major update. At the time I assumed they were operating as a ‘unitary executive.’ This broad update has been very important. President Biden’s former AI Czar, the AGI-pilled Ben Buchanan, is a good friend of mine.]
As I saw things, the only solution to this danger, and the broader reference class this danger belonged to, was open-weight AI models. It seemed to me that protecting open-weight was a matter of liberty more than it was a matter of geopolitics or diffusion, but in my early writing on open-weight models, I tended to focus on the latter two rather than the first one, which was my primary interest. Later, in a 2025 podcast with Rob Wiblin on 80,000 Hours, I reflected on this fact (emphasis added):
“So let’s just take the example of open source AI. Very plausibly, a way to mitigate the potential loss of control — or not even loss of control, but power imbalances that could exist between what we now think of as the AI companies… But if we have open source systems, and the ability to make these kinds of things is widely dispersed, then I think you do actually mitigate against some of these power imbalances in a quite significant way.
So part of the reason that I originally got into this field was to make a robust defence of open source because I worried about precisely this. In my public writing on the topic, I tended to talk more about how it’s better for diffusion, it’s better for innovation — and all that stuff is also true — because I was trying to make arguments in the like locally optimal discursive environment, right?”
AI risk is not the only area where I have had to make judgment calls about which arguments will be most palatable to various audiences.
Anyway, open-weight struck me as very important, for largely Jeffersonian reasons. The Biden Administration’s posture on open-weight models in 2023 was highly ambiguous, and a reasonable person would not be off base to conclude that it was hostile. The Biden State Department gave a grant to one AI safety organization–Gladstone AI–which proposed banning the open-weight distribution of any model above Llama 3 levels (in fact, they proposed making it a felony).
The AI safety community, in my eyes and sometimes in fact, allied itself closely with the Biden Administration. The AI safety community, therefore, in my eyes, were either useful idiots or willing co-conspirators in an effort to push America in the direction of tyranny. One way of re-phrasing my perspective at the time would have been, “look at how much potential for tyranny there already is with what they have done, and the actually powerful parts of our government have barely even woken up to AI yet! AND this is a structural problem of government itself, not just a pathology of the left!”
I am proud that, two years after having this view about my political opponents, I pushed back in no uncertain terms against another AI power grab by the Trump Administration, for whose White House I had worked and in whose Republican Party I have been a part for my entire adult life, in the Department of War’s dispute with Anthropic. My pushback on the Trump Administration’s actions here made national news and is the topic of, by far, the most popular essay of my career.
So, this is roughly the model of the situation, flaws and all, with which I began writing about AI: alignment and catastrophic risks were crucial to get right, transformative AI would probably arrive within a decade, concentration of power is the most important long-term risk, the Biden-era woke/AI safety alliance was terrifying and must be stopped, yet also the safetyists had many good points and were probably directionally right on many topics.
I wanted to articulate ideas and arguments that would strike on all those aspects of my world view: how can we push back against the tide of a government power grab while being serious about the risks while avoiding concentration of power and gradually building a classically liberal political theory of superintelligence? And can I do that while also conveying something useful about what it is like, personally, to live through this transformation? This was the intellectual project, and it still is.
But I also wanted this project to matter, which meant that it would need an audience. History suggested to me that AI policy would unfold in a series of crises or pivotal events over a period of 10 or so years, and that, as a guy with a keyboard and zero relationships with anyone important in the field, my best bet was to find a way to shape how interested members of the public would digest those crises. To do this, I thought about knowledge disseminating in networks. First there is the AI field itself, with its opinion shapers. Then there are all other industries and fields, where, simply out of curiosity, there were likely to lurk “AGI-curious” people, and more of them with each month (think: that one guy in the government agency who reads LessWrong, or the one guy on the trading floor listening to Dwarkesh in 2023).
The non-AI-professional-but-AGI-curious type would look to the AI opinion-shapers for cues, but when more people in his profession became AGI-curious, those people would naturally end up looking to him to understand what was going on. So, I reasoned, you want to sit at the Pareto frontier of “interesting to the AI in-group” and “approachable to the interested non-AI expert.” With each passing crisis, the stakes would rise and so too would the number of people in non-AI fields who took an interest in AI, and things would compound from there. This was my theory of change as a writer.
So with all this in mind conceptualized my audience and job as a writer as such:
This is the conceptualization of my work that explains most of the tradeoffs I made between “not wanting to sound crazy” and “trying to make actual contributions to AGI governance discourse.”
–
I am not saying I made no mistakes. But I think the notion that I have not acted with intellectual integrity is not quite fair either. Almost no one on the anti-SB 1047 side concurred with me about near-term transformative AI or the risks that transformative AI would present. From the very beginning of my involvement in the anti-SB 1047 coalition, there were various people who suggested that I was an “EA plant,” or similar, because of my willingness to (a) publicly acknowledge major risks from AI and (b) privately push back on dogmatic opposition to the cause of AI safety.
Here are mistakes I made, which I hope this context at least somewhat helps to alleviate. Melancholy Yuga quotes me in Marginal Revolution as saying SB 1047, with its focus on model-weight exfiltration, was “the stuff of science fiction, codified in law.” This was inaccurate and stupid of me. I did not view model weight exfiltration as “science fiction,” though I did believe it was incredibly unlikely that any LLM of the time would do it. In fact, apart from that one quote, which was in an email to Tyler Cowen, I do not believe I ever criticized the “kill switch” provision of SB 1047 in any of my public writing, and this was intentional (and quite distinct from many of the others on my side in the 1047 debate, who made that provision a major point of criticism). I should have said that I thought the scenario was unlikely with current models and not dismissed the weight-exfiltration threat model altogether.
I suspect I was sloppy there partially out of being a novice (my Substack at that time had a couple hundred subscribers, and I had almost no media experience), and partially out of fear of the “You Possess Your First Amendment Rights, Except When You Use AI” scenario I have already described. I took unfair potshots in an effort to beat people who I believed were, deliberately or otherwise, endangering my future. I was wrong to do this.
This is probably the single quote that comes closest to an example of me saying something I did not believe in public, and I shouldn’t have. Melancholy Yuga asked me which past statements of mine I would say I do not believe. This one is the sole example. The others are things I would probably phrase very differently today–or arguments I would skip over making altogether–but whose substance I stand by.
As Melancholy Yuga notes, I also recanted those statements. I began admitting I was wrong in my earlier posture in May 2024, five months after I began writing:
“I am here to tell you that the current debate over AI, no matter its flaws (and there are many), is among the most elevated and nuanced I have seen during my career in public policy. I am here to tell you that my intellectual “opponents”—those who worry immensely about AI catastrophic risks—are, by and large, honest and good faith people. I believe they are wrong, that they are sometimes anti-empirical, and that their proposed policies could be ruinous, but that is beside the point.”
At the end of 2024, I wrote a retrospective on the first year of my newsletter in which I said:
“I was convinced, way back in February 2024, that the eye of the state had turned toward AI, that fears about AI would be used to justify state intervention at massive scale into digital life, and that the AI safety community was at best composed of useful idiots for this intrusion, and at worst was actively enthusiastic about the prospect. I had been a silent observer of AI safety discussions on places like LessWrong and the AI Alignment Forum for years, so I did understand their concerns, and was even sympathetic to some of them.
Principally, though, I viewed AI safety as an enemy. And too often, I treated them like one—especially on X. I contributed to, and perhaps even helped to create, an unhealthy partisan divide on SB 1047. I presented the choice between SB 1047 and “not SB 1047” as a stark, civilizational fork in the road.
I believe this was a mistake, and it is one I have been trying to correct in recent months. I’ve come to understand that while the AI safety community is, as my friend Richard Ngo put it, “structurally power-seeking,” it is not the enemy I once apprehended it to be.”
I have also, it is worth noting, had extremely sharp words for those who dismiss AI risks, words that have become sharper with time. I’ve always had a particular anger at those who dismiss AI risk as “science fiction,” and I do wonder if there is not some sub-conscious link to my own mistaken use of that phrase in Marginal Revolution.
By the fall of 2024, “my side” had won the principal fight that animated 2024: SB 1047. But by the end of the 1047 debate I had a bad taste in my mouth. I knew I had a duty to offer a solution and not just be a critic, in particular because reasoning models had come to the fore.
I was not an AI maximalist during the 1047 debate. I believed we’d achieve AGI by roughly 2035 at the latest, and–as a Burkean–was skeptical at the imposition of immediate new laws. I did harbor some doubts that GPT-4-esque pre-train scaling alone could get us to AGI, and I said, in an August 2024 X exchange with Ajeya Cotra, that if I saw models engage in believable system II reasoning, that I would significantly update my views on AI safety, because this would make clear that the takeoff would be faster than the economics of pre-train scaling alone would imply. When o1-preview from OpenAI came out, I wrote:
“OpenAI’s breakthrough relies on a reinforcement learning-based method; this method will become publicly understood in due time. Indeed, many researchers are working on it, including ones in China. I expect that a Chinese company will produce a similar model within a few months, and perhaps sooner.
If your policy framework relies on the idea that only the largest models would have the potential for danger, and that we can keep advanced models out of our adversaries’ hands, I’m afraid your framework is unlikely to work.
We—analysts, policymakers, and the broader public—need to accept that these “extraordinary aliens” really are extraordinary, are here to stay, and that they are unlikely to remain within our “guardrails.” Instead, we need to build capabilities to make our society robust to the unique risks AI may pose. As I have written before, “those capabilities can include technical standards for AI, a coherent way to reason about AI liability, public computing infrastructure, digital public infrastructure for combatting deepfakes… and much, much else.” To this list I would add new technical protocols, governance mechanisms, and other approaches to make AI more legible to the state without invading user privacy or placing burdensome regulations on AI developers (stay tuned).
You should expect the pace of progress in AI to pick up yet again. You should not expect an impending “AI winter.” You should expect the coming years to be astonishing. You should not, necessarily, expect them to be wholly pleasant. You should expect the state’s grasp on AI to become even more tenuous. You should not expect everything to proceed in an orderly fashion.
We are in a new era, thanks to these aliens of extraordinary capabilities.”
After this essay, written in the days after the release of o1-preview, I began my project of developing my own approach to AI policy in earnest. This would end up focusing quite a bit on what I then called “private governance,” and what is now more commonly referred to as “independent verification organizations.” This idea began life as a fledgling, libertarianish idea, and while Tyler Cowen does indeed support something inspired by it, so too does the bipartisan FRONTIER Act, so too do several important AI safety groups, and so too do OpenAI and Anthropic.
So, in addition to being willing to apply my views about resisting state power grabs when it benefited my political side–and when it hurt it (and hurt me and my career)--I also pre-registered specific criteria that would cause me to change my beliefs on AI policy, and then did so when those criteria were met, then helped craft a policy solution that has become one of the Schelling Points of frontier AI governance, which is now beginning to prove its worth, in fits and starts, with the excellent METR report on the OpenAI-Hugging Face incident. In the interim, I held the pen on this country’s AI strategy, which at least mentioned biosecurity, model-weight security, data-center cybersecurity, semiconductor manufacturing equipment export controls, a major effort on mechanistic interpretability, among other things. The Action Plan has been unevenly implemented, but I think it was a modest achievement in bringing the key ideas of AI safety and security further into the bipartisan mainstream of policy discourse.
And yet, I admit to you that this achievement required grappling with political realities, as all political achievements do. One of several of those, and in my view the most important, is that I–along with many others in the AI policy community–failed to communicate with sufficient viscerality about the impending reality of self-sovereign AI. This was a major failure, and I regret it. I could have done more. At the same time, in a November 2024 essay called “Here’s What I Think We Should Do,” which was perceived by many as my audition for a seat in the Trump Administration, I devoted an entire section to the protocols I described in my very recent essay on self-sovereign AI:
“I believe it is possible that AI may require new protocols to be invented. What might this new wave of basic infrastructure do? Here’s a short list:
The need for and technical feasibility of these protocols is speculative. They may be largely created by the private sector, or the private sector may solve these problems in different ways. But it is an area meriting further study. I see no reason not to assign this inherently basic research to the federal agency that gave us the protocols of the internet: the Defense Advanced Research Projects Agency (DARPA).”
I regret that I couldn’t get this particular idea into the Action Plan. The Overton Window was not ready. Again, I fully admit I did not warn enough about why I supported these policies, but I did describe policies that were directly relevant to the concern I harbored at the time, and which I now admit I didn’t do enough to raise awareness of.
There are also concerns, which Melancholy Yuga notes, about my posture toward the AI safety community after 2024. I have made many critical comments toward AI safety since I recanted the unfair criticisms I made toward this community early in my writing career. I still hold almost all of the critiques of AI safety I mention at the beginning of this piece, and many others I have articulated since.
When Melancholy Yuga quotes me as saying "No single AI complaint/fear is salient enough to enough people to form a durable political movement," they wonder aloud whether I was being descriptive or normative. The answer is that these words are descriptive. Indeed, I was describing the traps of coalitional politics that can cause an intelligent, high-integrity person to say things they don’t believe, alleging that AI safety advocates could be tempted into making bad arguments about data center water use in the interest of their broader cause. I myself had fallen into this trap earlier in my career. In my effort to point out this trap, I veered into a discursive register that meant to attack a coalition and ended up landing on individuals within that coalition unfairly, and in retrospect I’d phrase them differently without altering the underlying substance of the critique.
Melancholy Yuga also questions whether my description of AI safety issues as “not being salient” was a normative statement or an analytic description. It was a description, made with great frustration, after, by that point, at least 18 earnest months of trying to warn people, in public and in private, about the impending risks of AI.
–
In the first half of 2024, I was largely fair about the existence of AI risks but unfair toward the AI safety community due to my own misapprehension about their alliance with parts of the Biden agenda I viewed (and still view) as pernicious. Since the second half of 2024, I have (a) recanted those unfair criticisms and (b) endeavored to proactively develop and socialize policies that, I hope, meaningfully address AI risk. Along the way I have repeatedly, and with increasing urgency, tried to keep my audience abreast of AI risks in the way I believed they were remotely ready to hear. I did not stay purely within their comfort zone–I routinely pushed beyond it, to the point of losing many friends and making many enemies–yet I tried not to veer so far outside it that my words became ineffective.
I also resisted efforts at undue power centralization, from Biden-era drives by AI safety advocates to ban open-weight models to Trump-era drives to bring AI under the boot of the national-security state.
I regret and apologize for the rhetorical failure I have already addressed in my original essay “On The Loose” and now here. I hope you will see this apology as a product of fastidiousness and high integrity, rather than some admission of a yearslong effort at deception. The opposite is true. I believe I have pulled many people, because of my words spoken and written in public and in private, toward the “AGI pill.”
I don’t believe anyone in AI safety should see me as an enemy. If anything even the furthest flanks of that movement should see the conclusion of “On The Loose” as an olive branch, extended then and still extended now.
Thanks for reading. Talk to you again soon, I hope.