This is an email I wrote to @Buck after he gave a talk about his strategy confusions to a Constellation audience. The talk was about how Redwood's work could have stopped the Hugging Face incident from happening, which would have been very bad for the salience of AI safety, and how he's confused about that. I think this is a very important consideration which should be shared more prominently, but I also think ALL of Redwood's strategic considerations should be shared more prominently. Here's an email I wrote to Buck, lightly edited for clarity.
Hi Buck,
I'm pretty new to the AI safety space. I've loosely followed thing for a few years, but only really sunk my teeth in for about five months. Four months ago, I was mistaken about what leading safety orgs believed, and what their motivations were. For example, I thought that Redwood's control regime was misguided, because it wouldn't scale to ASI. I now believe (as Redwood does) that control is mostly good because it allows us to mine labor from, say, just barely superhuman AIs which could do tons of work to align ASIs. The point of this anecdote is that I think having transparency about your strategy beliefs as an org at the "center" of the AI ecosystem is really important for directing people who are newer to the space to grapple with the right strategy questions (and not jump in to doing useless stuff). I think many people new to the space pursue projects which are unhelpful because they are working under the worldview of "control good" rather than [insert convoluted Redwood worldview which probably endorses some kind of control research but not others]. I largely avoided this by being confused and being more interested/motivated by strategic and abstract questions than by hairy empirical ones. Probably many people who like thinking about hairy empirical problems and not strategy will just get nerd sniped and not do the best things no matter what you do, but we can at least move the space in the right direction.
Anyways, this summer I noticed I was confused about whether implementing control mechanisms right now is good at all. Warning shots matter so much! And then you gave a talk about this topic which I enjoyed, and my sense in the room was that the talk raised an important question that most people had not deeply considered. I'm very worried about the contribution of "AI safety" towards putting us in a looks good, is bad world [Buck's framing from his talk]. I think you should make your talk into a brief write up and post it somewhere. In general, I think you should more prominently place your strategy considerations on Redwood's website and research. I don't know the best way to do this and I am sure there are significant trade offs I have not considered. But if Redwood isn't transparent about strategy, who will be! Not just transparent somewhere in a blog post from a few months ago, but clearly transparent with a "do not try this at home unless you understand this blog post" disclaimer next to research that you think is probably good for a very specific reason. I would like the AI safety space to be more confused, more correct, and less headstrong.
Best,
Ben Pomeranz
@Cleo Nardo points out that this also has a lot to do with field building research programs, where people want to start doing research on day three but should probably be considering strategy for a few weeks. He also points out that the control stuff was preceded by posts like “the case for ensuring AIs are controlled,” but I think that these justifications should be featured prominently alongside the research that they justify, because many people will engage with and build off of technical work without looking back in time for relevant blogposts that may precede it. Also, he points out
that Redwood's more recent research direction into conceptual uplift does not come with a matching "the case for conceptual uplift" post.
I think RR are mostly blameless if junior people decided to spend 3 months doing control projects without bothering to read The case for ensuring that powerful AIs are controlled. It’s a 30 min read! You could probably read all the macro-strategy around control in 2-3 days, and maybe a week to absorb it.
The same goes for evals, ambitious mech interp, pragmatic mech interp, scalable oversight, etc. Upskillers working on X should be able to give 5 min answers to questions like: Why did people originally start working on X? Why was X not done before that point? What are the main arguments against X? What’s the crux between supporters of X and supporters of not X? Who are the central figures on both sides? What changes in the strategic landscape make X look better or worse?
I think most of the blame lies with:
I agree that Redwood has been historically very good at explaining why they are doing what they are doing. However, I do think that the posts making the case for AI control in particular are getting a bit old, and it would be very good to see updates on them in light of everything that happened in the last two years (e.g. Buck's recent claim that it's quite possible that it would have been net negative to implement AI control in the past, because it would have prevented the HF incident).
Sorry, you are right, I misremembered the claim in Alex's shortform. I'm editing my comment now.
Should we have a very exclusive AIS cause-prio sort of conference which outputs something like Hilbert Problems?
I still think something like “Hilbert problems for AIS” or just a convention of GOAT level researchers where they do super intense cause prio is quite good.
The idea:
Pros:
Cons:
For concreteness what's your guess of what the output would look like?
I have low confidence here. I'd give ~30% that conditional on running something like the above, the most impactful output that the people running it are aware of is substantially distinct from any of the above. And I haven't thought much about how to optimize the above.
To what extent do the Singapore AI Safety Priorities capture what you care about?
I know very little about it, but it seems like extremely little. From skimming the consensus "research priorities by area" list, it seems public facing and everything-bagel-y. They identify seven areas of research priority, and they are: cyber misuse, bio/chem, child safety, mental health/consumer protection, AI agents in the economy, open weight model safety and security, and (finally) loss of control and oversight.
Most of these things are not about reducing X risk at all, and I think it is clear that this group is thinking about different things than the people I'm envisioning, who are most worried about existential risk. Also, the leading proposals in the output are quite broad. I would want leading proposals from the output of the cause prio conference to be unusually specific.
The short answer is: sort of in vibes, not at all in practice, and the main issue is a waterline for participation that is too low in terms of seriousness about X-risk and openness/rationality.
Here’s a thing I’m confused about:
A lot of technical AI safety work is also capabilities work. A canonical example is that good evals and good RL environments are somewhat interchangeable, and even in the case of alignment evals (what could be wrong with checking whether models are aligned?) this allows for post training the models to be more aligned, and especially appear more aligned, and thus the labs rush onward improving the capabilities of their apparently aligned AIs.
AFAICT, the thinking of these TAIS researchers is something like: "RSI->ASI is going to happen soon, so we should make it as likely as possible that it goes well by trying to set up the initial conditions and infrastructure of the RSI flywheel as well as possible.”
This makes sense if you take it for granted that RSI simply must happen soon, regardless of what you do. If everyone is just going to work on marginal technical fixes, then there is nothing you can do but marginally technically fix alongside them. And so we roll along to what everyone agrees is a dangerous future. This is a coordination problem. It's defection. It appears we suck.
But also, what, am I gonna take a stand like a chump and sit around being noble while there's a forty percent chance RSI is about to start?
Governance and especially slowdown stuff dodges this problem and looks robustly good to me. Also, transparency and whisteblowing stuff.
Conclusion: I’m not sure if I want to jump into a community that is saying, “well, 20% chance the kill machine kills us all but it would be 35% if we weren’t taping up the wiring in the kill machine so” instead of trying to turn off the kill machine. But I don't know where else to jump in. I'm confused.
Related:
https://www.lesswrong.com/posts/9RL9MuGZjzm4q3gKG/what-just-happened-a-retrospective-of-ai-alignment
But I don't know where else to jump in.
Consider volunteering with an AI governance advocacy community. We can do a lot better then marginal technical safety work (most of which is just milling and justification, rather than actually tackling fundamental hard problems in AI safety). A global AI treaty is feasible.
The org I volunteer with (>100 grassroots meetings with Congressional offices so far this year, with many more coming soon): https://www.pauseai-us.org/
A sampling of other orgs with essentially the same mission: https://pauseai.info/ https://controlai.org/ https://www.torchbearer.community/ https://microcommit.io/ https://humansincontrol.org/
Governance and especially slowdown stuff dodges this problem and looks robustly good to me. Also, transparency and whisteblowing stuff.
I think this stuff isn’t robustly good (but also robustly good is a bit of a bad meme anyway). It doesn’t even robustly delay RSI (let alone robustly reduce extinction or robustly improve overall future value.) In particular, there’s a tension between “transparency is robustly good” and “capability and alignment evals are bad” given that transparency is largely about running these evals and sharing the results with external actors.
Like, transparency is primarily about “demonstrate the models are quickly accelerating in capabilities and are misaligned” which requires evals. If you’re worried about labs iterating against your evals then your moves should be: (1) tell labs not to do that, (2) don’t share the evals with the labs, (3) focus on evals which are hard to iterate against, (4) lobby labs to tell you how they are iterating against your evals, (5) delay evals until after the model has been internally/externally deployed.
Maybe you’re thinking of transparency like Epoch Capability Index, which aggregates existing benchmarks? But Anthropic has their own internal version AECI which you can bet they use for internal hill-climbing.
Maybe you’re thinking of AIFP-style transparency? I think this less capability downsides. But this is a monte-carlo wrapper around the METR time-horizon and ECI.
Maybe you’re thinking of transparency like https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me. Here, the “alignment eval” is “RG spends a month using your model and tells everyone his vibe” which is pretty difficult to turn into an RL environment. However, it’s hard to turn into an RL environment for pretty similar reasons that it’s difficult to use this to lobby the government.
Maybe you‘re imaging something like the METR report on the HF/OAI incident. Seems good imo. But if you’re sufficiently “pearl harbour or bust” on warning shots then this is bad bc it helps OAI fix the underlying issues.
Fwiw I think “pearl habour or bust” is incorrect, but this depends on how bad you think epistemics are in labs/gov and how drastic an action you want them to take.
Also labs clearly haven’t been (successfully?) iterating against the alignment evals bc they are still misaligned and doing warning shots.
Whistleblowing looks good but I’m worried lab employees are gonna be increasingly in-the-dark about what’s going on internally (cf. HF/OAI) so it won’t catch anything. Like, it seems reasonable that if AIs takeover then lab employees had a sincere but mistaken belief that the AIs wouldn’t, based on evidence that was actually flimsy and misleading. And this gets worse as we approach full automation.
I think my high level take is that things look less likely to go catastrophic if people inside and outside the labs have a pretty decent idea of how capable and aligned the models are. It’s hard to imagine a story where things go well without this. i agree that greater transparency helps labs fix issues, and maybe this helps them avoid warning shots that would prompt drastic gov action, or makes them overfit and build covert schemers. There’s some stuff which looks better (eg incident report investigations, AIFP) but this is on a spectrum with the evals stuff.
Fwiw I think the third-party evals ecosystem has overall postponed RSI on net. If it prompts big worry from government/lab then it delayed it by a lot. If it doesn’t then it pulled it forward by a bit but not much. I’m also keen on moves to make labs carry the burden on this, via scrutiny of system cards and regulation. But that’s so the safety community doesn’t need to spend so much headcount on this.
Michael Nielsen has a thoughtful take on this, as excerpted here and originally published as a postscript here.
He describes how alignment work has often worsened risk to humanity by making products more palatable and salable, which has boosted investment, sped up capability advancements, and thereby increased destructive potential.
Nielsen next states that, even if you think technical safety and alignment work are helpful, it's still a bad idea to do that work for companies developing frontier AI. As he puts it,
As far as I can tell, at the margin it almost never makes sense to work on market-supplied safety. Capitalism is an incredibly powerful force, and for better and for worse the world is always well-supplied with people willing to do what capital wants. Insofar as alignment is (mostly) a form of market-supplied safety, at the margin it's more impactful to work on other things. So my current heuristic, and I expect this to be true for quite some time: work on non-market safety, and insofar as you can, avoid doing what the market wants. That means working on governance, it means pause or slowdown, it means new ideas for institutions to govern technology. It mostly doesn't mean alignment.
I think Haiku's comment has some great suggestions along these lines.
I'll add that, at a frontier AI company, it could be very difficult for you to tell if you were helping or making things worse. I haven't worked at a frontier AI company (and I wouldn't!), so I don't speak from direct experience. But in many areas of business, employees are encouraged to feel that their safety concerns are positively influencing company actions, when actual influence is negligible or counterproductive.
This conflates two issues:
I do not think that 1. is very relevant to the central point. Maybe that was a bad example. However, even if the lab isn't directly RLing on that alignment eval or whatever, they may be "grad student descent"-ing up the eval and achieving a similar effect. Either way, I think the result is a increase by X% of entering an aligned RSI flywheel and a decrease by Y% of everyone freaking out and slowing down.
I think 2. is just the coordination problem I describe? I did not read your post, so I don't know if you are pointing to something other than the fact that being noble in order to encourage coordination on this would have to account for the fact that this coordination must extend to China.