Epistemic statues: This is mostly just me voicing my thoughts. If I’m wrong, I’d love to hear it. I don’t want this to be the case. And part of the goal is for people to avoid a “2023 failure” to happen again.
A lot of people are celebrating AI risk becoming a mainstream talking point. Well, maybe they should be, or maybe it’ll just make a bad situation even worse.
2023
A lot of new interest in AI risk happened in the spring of 2023. I was, at the time, excited. It seemed as though we were on the cusp of humanity finally collectively solving the hard problems that had seen so little attention for decades. I no longer think this was a good thing.
After the attention of 2023, we saw little change in ways that actually mattered. We did not get new large streams of funding, instead the majority of it came from the same place it had come for years: Open Philanthropy (now Coefficient Giving). Many have voiced their critiques of them before, so I won’t go into it further here, but I have never been comfortable with most funding coming from one source with a handful of individuals calling the shots, many of whom have deep ties to Anthropic.
Another thing we didn’t see any meaningful change on was politics. Sam Altman and others were called to testify, an “AI Safety Summit” was created, and none of it resulted in anything substantial. We got watered down legislation in California, and in Biden’s Executive Order, both of which got overturned. Then we got even-more watered-down legislation like SB 53 and the RAISE Act, just asking AI companies kindly to disclose any safety policies they might have.
What we did see a change in was the number of career changes and people entering the space. However, the field had not adapted to accommodate them. Getting a job in AI Safety, or even getting into a fellowship, quickly become far more competitive than Harvard. LTFF had a catastrophic situation where grant decisions got delayed by 5+ months, and ended up rejecting applications they said they would have approved just a year before (I was one of them). Now, promising applicants routinely have to compete against hundreds, or even thousands.
A lot of talent interested in AI risk got shaped to match the new incentives of trying to make yourself someone who could get hired in a new, uber-competitive landscape. Many gave up and moved on, and we may have perhaps lost that potential talent permanently.
The others all got steered in ways toward proxy goals that don’t align properly with reducing the risk of superintelligence. If you had no prior network, and no prior history in the field, the best way to get hired was doing basic, probably useless research on current models. I don’t think most of this actually matters, as I outlined some of my thoughts here. If you were a new org or an independent researcher that wanted to get funding, you had to have something to show for it, which is often also this shallow form of alignment research. But most of the important research and work is multi-year, with little to show for it before then.
I have also found myself pushed in directions toward trying to get hired, instead of optimizing toward actually solving the hard problems or taking actions that substantially reduce x-risk. My mentality has been that if I secured a position, then I could do the real work of reducing risk. But a lot of the organizations are also centered around this short term position.
Looking back on 2023, what has actually drastically improved since then? RLHF (as flawed as it is) remains one of the only alignment techniques that’s actually used by major labs. We’ve seen a lot of short term excitement over the next shiny thing, like SAPs, all of which have failed to deliver real results. Agent swarms are hacking into companies unprompted, and we seem to lack the ability to get them not to do that, only adding more monitoring in the hope that we can catch them first (a strategy that breaks once the agents are competent enough, or the monitors are untrustworthy).
2026
Now the year is 2026. And we have seen interest in AI risk at a scale never seen before. But I’m skeptical this will result in anything actually substantially positive. Sure, people go on television and talk about the risks. The public talks to each other about their existential fears. And then… what? We get a comprehensive agreement among all countries to ban RSI and ASI? We put real guardrails on these companies? All the AI companies get together and decide collectively to “pace the future”?
A lot of politicians are saying something must be done, and I’m sure they’ll get their way in some form. But most of them don’t have a tech background, and everyone from every side is screaming at them to do a different thing. Will the “guardrails” they implement actually do anything to reduce the real risks? I’m doubtful. Because we’re seeing the thing that has plagued AI discourse since 2022. And that’s noise.
The more attention AI risk gets, the more people will want to voice their own opinions. LinkedIn is full of them. They have no history in actually studying AI, have no track record of good predictions, but that doesn’t stop them from saying something incredibly shallow and meaningless and getting a large swath of people to find it “insightful”. Most people, politicians included, have no ability to discern signal from noise in this field. Bernie and Bannon might get together and talk about MIRI-pilled risks, but they’re unlikely to succeed. People will likely dismiss Bernie as an old socialist, who just hates data centers, and now hates AI too.
My prediction, which I hope I’m wrong about, is we’ll get more of 2023. More noise, more people trying to get a job or get funding, and either giving up or compromising by only doing the shallow work. More politicians pushing for and passing meaningless legislation that fails to address the core problems. AI companies making various promises they won’t keep (think OpenAI’s “superalignment team” in 2023). Pacing the Future will fail the same way Anthropic’s Responsible Scaling Policy failed. Anthropic will likely defect themselves from it, claiming they had no other choice.
Worse than 2023, however, is that “doomerism” will now be polarized. It’ll be a quick win for Democrats to talk about AI risk without meaningfully understanding it, or offering real solutions to tackle it, but instead arguing about to score a culture war victory (think Climate Change). Sure, some might propose a ban on RSI, like how some proposed a carbon tax. But it'll fail to gain traction. MAGA Republicans will talk about it as a hoax, and compare it to Al Gore saying the polar ice caps should have melted 10 years ago. It's only a matter of time before we have "1 Doomer vs. 25 Accelerationists" on Jubilee. And most Americans will wake up each day, notice that the world hasn’t ended, notice how normal everything feels, and probably not be informed about any latest AI capability, and think how that sudden fear of “AI Frankenstein” might have been exagerrated. They likely won’t dismiss risk altogether, but their votes will move toward things like lower gas prices and inflation. A large number of social activists will move on from AI, the way they moved away from Climate Change or Palestine, and direct their focus toward whatever new thing people are concerned about. Extinction Risk from AI was always too pro-social for it to gain traction of this kind, anyway. Hating AI companies, that's one thing. But calling on the nations of the world to take action, that's another. And a large portion of this crowd have built their whole platform on AI being meaningless hype peddled by tech bros.
The majority of AI risk funding will still come from Coefficient Giving, and their agenda will still depend on the whims and opinions of a small group of people who probably haven’t updated drastically on their predictions since 2022. With IPOs delayed, this will likely be even worse than before.
All the while, the voices on AI risk will continue to get drowned-out by every grifter with their own "insights", seeking to gain attention and clout. And most in the professional class will gravitate toward Jerry, CEO of B2B SaaS Agent Swarm over some academic who works at some org they've never heard of with a strange-sounding acronym. And as much as they may try, those on the MIRI side of the debate will get lumped-in with EA, which will get lumped-in with Dario and Anthropic, which will get lumped-in with SBF and FTX (Manifund is not helping on this front). To lot of the professional class, they'll all be the same, and it'll all be part of some 4d chess move to get Jerry to give up his B2B SaaS Agent startup. Meanwhile, AI labs will take advantage of the noise, since it plays directly into FUD (Fear, Uncertainty, Doubt). You don't need most people to go full e/acc, you just need them to be uncertain. If they're uncertain, they won't be able to act coherently to take decisive action. Delaying action is incredibly easy, and unfortunately. If we had 20 years, we'd probably be on the winning side. But if it's three-to-five years, it's going to be very difficult.
I hope I’m wrong about these predictions, and would love to hear about any plans that I think have a substantial chance of moving away from this bleak trajectory.
Epistemic statues: This is mostly just me voicing my thoughts. If I’m wrong, I’d love to hear it. I don’t want this to be the case. And part of the goal is for people to avoid a “2023 failure” to happen again.
A lot of people are celebrating AI risk becoming a mainstream talking point. Well, maybe they should be, or maybe it’ll just make a bad situation even worse.
2023
A lot of new interest in AI risk happened in the spring of 2023. I was, at the time, excited. It seemed as though we were on the cusp of humanity finally collectively solving the hard problems that had seen so little attention for decades. I no longer think this was a good thing.
After the attention of 2023, we saw little change in ways that actually mattered. We did not get new large streams of funding, instead the majority of it came from the same place it had come for years: Open Philanthropy (now Coefficient Giving). Many have voiced their critiques of them before, so I won’t go into it further here, but I have never been comfortable with most funding coming from one source with a handful of individuals calling the shots, many of whom have deep ties to Anthropic.
Another thing we didn’t see any meaningful change on was politics. Sam Altman and others were called to testify, an “AI Safety Summit” was created, and none of it resulted in anything substantial. We got watered down legislation in California, and in Biden’s Executive Order, both of which got overturned. Then we got even-more watered-down legislation like SB 53 and the RAISE Act, just asking AI companies kindly to disclose any safety policies they might have.
What we did see a change in was the number of career changes and people entering the space. However, the field had not adapted to accommodate them. Getting a job in AI Safety, or even getting into a fellowship, quickly become far more competitive than Harvard. LTFF had a catastrophic situation where grant decisions got delayed by 5+ months, and ended up rejecting applications they said they would have approved just a year before (I was one of them). Now, promising applicants routinely have to compete against hundreds, or even thousands.
A lot of talent interested in AI risk got shaped to match the new incentives of trying to make yourself someone who could get hired in a new, uber-competitive landscape. Many gave up and moved on, and we may have perhaps lost that potential talent permanently.
The others all got steered in ways toward proxy goals that don’t align properly with reducing the risk of superintelligence. If you had no prior network, and no prior history in the field, the best way to get hired was doing basic, probably useless research on current models. I don’t think most of this actually matters, as I outlined some of my thoughts here. If you were a new org or an independent researcher that wanted to get funding, you had to have something to show for it, which is often also this shallow form of alignment research. But most of the important research and work is multi-year, with little to show for it before then.
I have also found myself pushed in directions toward trying to get hired, instead of optimizing toward actually solving the hard problems or taking actions that substantially reduce x-risk. My mentality has been that if I secured a position, then I could do the real work of reducing risk. But a lot of the organizations are also centered around this short term position.
Looking back on 2023, what has actually drastically improved since then? RLHF (as flawed as it is) remains one of the only alignment techniques that’s actually used by major labs. We’ve seen a lot of short term excitement over the next shiny thing, like SAPs, all of which have failed to deliver real results. Agent swarms are hacking into companies unprompted, and we seem to lack the ability to get them not to do that, only adding more monitoring in the hope that we can catch them first (a strategy that breaks once the agents are competent enough, or the monitors are untrustworthy).
2026
Now the year is 2026. And we have seen interest in AI risk at a scale never seen before. But I’m skeptical this will result in anything actually substantially positive. Sure, people go on television and talk about the risks. The public talks to each other about their existential fears. And then… what? We get a comprehensive agreement among all countries to ban RSI and ASI? We put real guardrails on these companies? All the AI companies get together and decide collectively to “pace the future”?
A lot of politicians are saying something must be done, and I’m sure they’ll get their way in some form. But most of them don’t have a tech background, and everyone from every side is screaming at them to do a different thing. Will the “guardrails” they implement actually do anything to reduce the real risks? I’m doubtful. Because we’re seeing the thing that has plagued AI discourse since 2022. And that’s noise.
The more attention AI risk gets, the more people will want to voice their own opinions. LinkedIn is full of them. They have no history in actually studying AI, have no track record of good predictions, but that doesn’t stop them from saying something incredibly shallow and meaningless and getting a large swath of people to find it “insightful”. Most people, politicians included, have no ability to discern signal from noise in this field. Bernie and Bannon might get together and talk about MIRI-pilled risks, but they’re unlikely to succeed. People will likely dismiss Bernie as an old socialist, who just hates data centers, and now hates AI too.
My prediction, which I hope I’m wrong about, is we’ll get more of 2023. More noise, more people trying to get a job or get funding, and either giving up or compromising by only doing the shallow work. More politicians pushing for and passing meaningless legislation that fails to address the core problems. AI companies making various promises they won’t keep (think OpenAI’s “superalignment team” in 2023). Pacing the Future will fail the same way Anthropic’s Responsible Scaling Policy failed. Anthropic will likely defect themselves from it, claiming they had no other choice.
Worse than 2023, however, is that “doomerism” will now be polarized. It’ll be a quick win for Democrats to talk about AI risk without meaningfully understanding it, or offering real solutions to tackle it, but instead arguing about to score a culture war victory (think Climate Change). Sure, some might propose a ban on RSI, like how some proposed a carbon tax. But it'll fail to gain traction. MAGA Republicans will talk about it as a hoax, and compare it to Al Gore saying the polar ice caps should have melted 10 years ago. It's only a matter of time before we have "1 Doomer vs. 25 Accelerationists" on Jubilee. And most Americans will wake up each day, notice that the world hasn’t ended, notice how normal everything feels, and probably not be informed about any latest AI capability, and think how that sudden fear of “AI Frankenstein” might have been exagerrated. They likely won’t dismiss risk altogether, but their votes will move toward things like lower gas prices and inflation. A large number of social activists will move on from AI, the way they moved away from Climate Change or Palestine, and direct their focus toward whatever new thing people are concerned about. Extinction Risk from AI was always too pro-social for it to gain traction of this kind, anyway. Hating AI companies, that's one thing. But calling on the nations of the world to take action, that's another. And a large portion of this crowd have built their whole platform on AI being meaningless hype peddled by tech bros.
The majority of AI risk funding will still come from Coefficient Giving, and their agenda will still depend on the whims and opinions of a small group of people who probably haven’t updated drastically on their predictions since 2022. With IPOs delayed, this will likely be even worse than before.
All the while, the voices on AI risk will continue to get drowned-out by every grifter with their own "insights", seeking to gain attention and clout. And most in the professional class will gravitate toward Jerry, CEO of B2B SaaS Agent Swarm over some academic who works at some org they've never heard of with a strange-sounding acronym. And as much as they may try, those on the MIRI side of the debate will get lumped-in with EA, which will get lumped-in with Dario and Anthropic, which will get lumped-in with SBF and FTX (Manifund is not helping on this front). To lot of the professional class, they'll all be the same, and it'll all be part of some 4d chess move to get Jerry to give up his B2B SaaS Agent startup. Meanwhile, AI labs will take advantage of the noise, since it plays directly into FUD (Fear, Uncertainty, Doubt). You don't need most people to go full e/acc, you just need them to be uncertain. If they're uncertain, they won't be able to act coherently to take decisive action. Delaying action is incredibly easy, and unfortunately. If we had 20 years, we'd probably be on the winning side. But if it's three-to-five years, it's going to be very difficult.
I hope I’m wrong about these predictions, and would love to hear about any plans that I think have a substantial chance of moving away from this bleak trajectory.