I'm not sure how people in the Dwarkesh basin arrive at their impression of the level of abstraction for the [X]s we need verifiable examples of to train AI to do something, but I wonder what they think were the pre-existing and verifiable examples we used of "AIs that can do the things current AI can do," in order to show the AI that it should be able to do those things. If you respond to that question with, of course it comes from lower level less abstract examples, well, how low can you go? Because at the other end of that line of reasoning lies Deep Thought before its databanks were connected. Seems like we're banking on being and staying in a pretty narrow band.
Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go.
As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped.
If I am quoting directly I use quote marks, otherwise assume paraphrases.
Section titles are from the transcript whenever possible, to aid in navigation.
Introduction
The discussion is interesting throughout, although often frustrating, especially in the (mostly isolated) discussion about ‘aligned to whom?’ As usual, one could expand many responses into full posts, and maybe one should.
This podcast exists in light of recent misalignment and hacking events at OpenAI, Anthropic and UK AISI. You’ll want basic knowledge of that as background.
Ryan and Dwarkesh both have views of the situation different from my own, but are attempting to see where their positions lead, and try to balance educating people who start at zero with having a high level discussion.
Also important background is The Three AI Pills. Dwarkesh cannot be understood here except as someone who is at least somewhat AGI pilled, who realizes that AI is going to be a huge deal and is scary, but that is not ASI pilled. You can also see the reason he rejects that pill, which is I read as, roughly, that he thinks AI learns to do [X] only via examples of [X], and that AI can then only do those [X]s, although combining them and new context might allow modestly new things to happen. He recognizes that already this is kind of a huge deal.
You can see how he came to that over the course of many years and podcasts, if you have been paying attention. There are a lot of influences leading in that direction.
Ryan, who comes from Redwood Research, comes from the ‘models be scheming’ school of misalignment, where when something goes wrong the models become misaligned or start scheming, and the danger is that the models scheme, especially in ways involving the training pipeline. I think this tries to draw a distinction of magisteria that is not there, and overcomplicates and overspecifies, but is not wrong.
I am much closer to Ryan’s position than to Dwarkesh’s, and indeed could be seen as farther past Ryan if we put this on a spectrum.
Having AI do the AI R&D not only means it would go scary fast, it means it would by default focus on what can be measured, and make all the things going horribly wrong go that much more horribly wrong. You end up in a spiral of RLVR for doing RLVR for misaligned models.
Is AI R&D Verifiable Enough To Unlock Recursive Self-Improvement?
Whether or not we will get recursive self-improvement (RSI), and how fast, is the right question. Reading only the title to this section, I want to say ‘wrong sub-question’ but don’t want to jump the gun.
I’m going to group things in ‘logical’ order, not the exact order things were said.
Dwarkesh will be the skeptic. Ryan will make the case for RSI.
Is AI progress bottlenecked by human expert data?
This is what I imagine five more years of AI progress at this pace would enable an AI to be able to do. This is the thing I’m really worried about: ASI that can understand how to do crazy shit in the world, that can do what Kissinger can do, can do what Steve Jobs can do, et cetera, and also his engineers and so on. I’m not sure how you get that without the relevant world data, which is the equivalent of Mythos being really good at coding while not having the coding environments that have improved it relative to GPT-3.”
Flat token prices suggest scaling has been slow
Skills AI can’t train on: does it even need them?
Aligned to whom?
This section is mostly self-contained.
They keep at it, and I want to emphasize I think the position being expressed here – that Anthropic should have Claude cooperate with actively harmful requests – is not merely wrong, it is basically kind of nuts.
I likely need to write an explanatory evergreen post laying out why it is nuts.
Dwarkesh has episode notes and some Twitter posts that go further into his position on reflection, which I’m treating here as beyond scope for now since this is an isolated section. I will be returning to the topic later.
Ryan does offer some other comments on why one might prefer a virtue ethical approach, including that Anthropic might think it is ‘easier to align.’ I agree it is easier in the sense that a deontological approach definitely won’t work, and also can’t be antifragile to all the mistakes you will make, and also does not give you what you want because you cannot specify it, and also pure deontology is not how minds work in practice at least until higher capability levels, and so on. As with many other places, that would be a full long post and I’m going to say that it is beyond scope to go into more detail, other than to say that yes the whole thing is rather overdetermined.
Recent incidents of AIs colluding and deceiving humans
What could possibly go wrong? A concrete scenario
From reward hacking to takeover
Time To Update
Dwarkesh does seem to have several ‘holy ****’ moments throughout the podcast, even if he doesn’t use that language to describe them. That was good to see. In general, this was Dwarkesh being somewhat AGI pilled but not ASI pilled. And then it felt like him trying to advocate for a series of positions and predictions he has built up over the years as a compromise between various different groups and interests, including Leopold, the Tyler Cowen group and those at Mechanize, among others. Except he is realizing that events are not conforming to the plan, and that he must update.
I was happily surprised that I left the podcast entirely sober, in that Dwarkesh did not say the magic words ‘continual learning.’
The central idea behind the Dwarkesh unified perspective, as I understand it, is that AI learns to do [X] if and only if you have a bunch of well-specified verifiable examples of [X] on which you can train, although they can then be combined to produce somewhat new things. Thus, you can’t make that much progress that fast, and the progress will top out, and the impact of it will top out, and so on. And then, if you want to understand impact of AI, you can ask for particular [X]s and add them up.
That is not how I think of even current AIs, let alone future smarter AIs, making ‘which [X]s will we have data for?’ the wrong sub-question.
This then encountered Ryan’s view of all this, which in many ways is a lot closer to my own, but it has ‘scheming’ as kind of a distinct magisteria and active ingredient, as opposed to being a natural thing, and he’s tiptoeing around various aspects to avoid what he fears are landmines.
This comes from his optimism. Ryan’s basic position is that if you had good execution and ‘make no mistakes’ then you get good outcomes. What you have to do is not mess up and do all the work, and make enough incremental safety progress. Not messing up is really hard, we are not doing enough of the work, but you can imagine that changing. I don’t think it is that easy. I think the default is you start out dead, your mistakes make you even more dead but also help you realize how dead you are and thus can help, every time you try to solve your problems in the current ways you push them to be more subtle but also more insidious, and you need to actively find a way out.
Contrary to some whose positions I respect, I do think it is possible for us to find a way out. I just think that is going to be extremely hard to do on the first try, when we cannot even hope to ‘make no mistakes,’ under tremendous pressures.