Some notes on public communication, which I expect the LW audience is especially interested in:
I made the tone in this post softer than I would have if I were not a director,
Lol. Ah yes, finally, Paul Christiano is toning down his extremist rhetoric!
If anyone who thinks this comment is somehow inappropriate (or doesn't get why it's amusing) would like to explain why, I would welcome that.
One of the top 3 things that has made Paul's career has been writing a lot in public about this subject in a way that is respectable, never rhetorically alarming, and able to serve as an optimistic counterweight those sounding the alarm. So it is funny in the immediate weeks after the NYT has frontpage pieces about rogue AIs committing crimes and Bernie Sanders announcing a bill to ban superintelligence, to see his one tonal note be that he has softened his writing even further.
Recap for those not tracking the details: my understanding is that the nonprofit (“OpenAI Foundation”) owns 26% of the shares in the public-benefit corporation (“OpenAI Group PBC”). However, the non-profit board has special governance rights that give it complete control over who sits on the PBC board: it appoints all of its directors and can replace them at any time.
Confusingly, the two boards have so far been nearly identical, with 9 people sitting on both, including Bret Taylor as Chair of both boards, and Sam Altman also serving on both (being the only PBC executive to do so). Zico Kolter (Professor at CMU and Head of its Machine Learning Department) was the one exception, serving only on the nonprofit board (and as a non-voting "observer" on the PBC board). Paul Christiano now appears to be the second person appointed only to the nonprofit board.
I guess if there are subcommittees, perhaps more people will be announced soon. We’ll see!
This seems like a position that should not be accepted lightly. Do you have reason to believe the SSC and the board will in practice be a meaningful check on OpenAI?
A twit by the founder of an AI safety advocacy organization states that
the SSC in particular oversees model launches and can delay releases for safety reasons
I have never heard this before and is outside the usual remit of boards. This would be a strong positive update for me about OpenAI's governance having any teeth (though it will not prevent an extinction-level threat, because of course most of the risk comes from training the model and internal runs, not releasing it to the public). Anyone got any more information on this?
and keep in mind that the company can just ignore its governance mechanisms
I am kind of confused about the linguistic construction of "can delay releases for safety reasons" and "the company can just ignore its governance mechanisms". It is not currently clear to me whether the board can actually delay releases for safety reasons (though it might be in some sense supposed to be vested with the power to do so).
My understanding is that the board has the de jure formal power to delay releases but whether or not they have the de facto power to do so is unknown until tested.
Is the former actually true? Like, do we have formalized bylaws or anything like that that gives them actual de jure formal power? I have sent ChatGPT off on a research thread to help me figure this out...
ChatGPT says we don't have any solid evidence to support that it's in bylaws, and all we have to go off of is that blogpost. I currently would take a bet that the board does not have actually formalized non-overwriteable formal power to delay releases (but think it's still like 25% likely).
I would be interested in @paulfchristiano clarifying, if he can.
Can you describe what you think the organization is doing by bringing you on? What has not been done, that will get done now that you're in the room?
While I believe that it is unethical to work for OpenAI in any technical capacity (it will be twisted into capabilities progress), I think the role you are taking is defensible, and I trust you in particular to leave (and publicly state why) if you discern otherwise.
Didn't Paul take a government regulator role job while under a secret non-disparagement agreement with OpenAI? If so, pretty dumb move if you ask me.
If we build superintelligence without more robust alignment I expect we will permanently lose control of it. If that happens then most people could die.
If Anyone Builds It, Most People Could Die wasn't quite as catchy of a title, I suppose - though per your comment I do understand the more authority you have, the more euphemistic you need to be in public comms. I do hope that when the time comes you're willing to be more explicit/confident about your concerns - best of luck in your new role.
I am excited to be joining the OpenAI nonprofit board, serving on the Safety and Security Committee to support safety oversight.
Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.[1] I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level. I’m joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk.
The SSC has an important and challenging role in overseeing risk management at OpenAI, and I hope to help provide expertise and assistance in a critical moment. My joining is not an endorsement or criticism of OpenAI’s safety practices in particular; I hope that all frontier companies strengthen safety oversight and I am excited to work on this at OpenAI. I believe that the rest of the world should judge OpenAI, and all AI developers, by externally verifiable behavior and results.
In the rest of this post, I'll explain why I believe loss-of-control risk is now acute and how I think about the current situation in the AI industry.
First, automated AI R&D could lead to a very rapid acceleration in AI capabilities very soon. OpenAI has predicted that we might have capabilities sufficient to fully automate AI research within 18 months; my personal forecast is extremely uncertain and I think it could easily take anywhere from several months to several years.
Full automation of AI R&D means that improvements in training and algorithms can directly increase the quality and quantity of automated AI researchers available to do additional research. Existing evidence is very uncertain but suggests that this positive feedback loop might be strong enough to overcome diminishing returns and compute bottlenecks, leading to a rapid intelligence explosion. If this happens, then within six months of full AI R&D automation we could see more algorithmic progress than has occurred since the development of the Transformer nearly a decade ago. I believe this would result in superintelligent AI systems.
Second, we currently train our AI agents with RL to get as much reward as they can. It has long seemed theoretically possible that this could motivate AI agents to undermine human control, seek power and resources, and cover up their tracks in pursuit of misaligned goals correlated with reward. Public evidence from recent incidents suggests that this is not just a theoretical possibility.
An intelligence explosion would greatly exacerbate risks from misalignment, both by making the technical problem of alignment even more difficult and by rapidly raising the stakes for failure. Many researchers and leaders at OpenAI and across the industry have expressed concern that rapid recursive self-improvement is not consistent with safe development; I resonated with this recent post by OpenAI's chief scientist Jakub Pachocki on this topic.
If we build superintelligence without more robust alignment I expect we will permanently lose control of it. If that happens then most people could die. I believe we would need domestic and international coordination to ensure global consistency and reduce risk to an acceptable level.
That said, frontier AI developers have a lot of power to unilaterally improve the situation and lay the groundwork for stronger coordination. Developers can improve safety mitigations (including slowing development as necessary), transparently share evidence about risk and the effectiveness of their mitigations, and work towards shared safety standards.
I am encouraged by other members of the SSC, as well as the rest of the board and leadership, taking these issues seriously. I look forward to working with them to help OpenAI raise the bar for its safety practices.
If I had to quantify my uncertainty I would estimate an all-things-considered risk of 4% over the next year and 15% over the next three years. These numbers are a way of stating my subjective beliefs and communicating roughly how big I think the problem is, not a claim to have a model that produces precise or stable estimates.