While OpenAI and Anthropic[1] pursue different lines of safety research, they have yet to produce a public-facing document describing concretely how their companies plan to align superintelligence. I think it is underappreciated how this points to general negligence or a lack of openness to third-party feedback.
By “plan,” I mean a document describing a proposal for technical alignment with at least the level of detail and research effort of AI 2040.[2] Any such plan for technical alignment would likely be flawed in non-obvious ways. But having a proposal that’s sensible enough to consider and detailed enough to critique is a good starting point for wiser proposals. Making such a plan public would also create feedback loops for accountability.
The closest thing to a plan came in 2023, when OpenAI announced their superalignment strategy (also relevant). I do not find this approach particularly convincing, though I do find it laudable that OpenAI explained what they planned to do, who would lead the effort, and what resources would be allocated, at a level of detail which made critique possible.[3] This team no longer exists, and nowadays, as far as I am aware, the research community doesn’t have precise answers to basic questions like:
How, specifically, does OpenAI or Anthropic plan to have their own AIs help solve the alignment problem? What if their alignment agents are themselves somewhat misaligned?
Does OpenAI or Anthropic imagine “aligned” superintelligence being corrigible? Corrigible to who?
What fraction of compute or money is being spent on alignment at OpenAI or Anthropic?
If OpenAI and Anthropic leadership do not think they could produce a technical alignment plan as detailed as AI 2040, they should communicate this loudly and publicly, stating where the uncertainties lie and how they plan to resolve them.
In light of recentevents, it's self-evident that Anthropic and OpenAI cannot reliably control non-superintelligent systems. We don’t know if or how labs will act differently in response, but the public and the research community shouldn’t have to wonder. There should be transparency and room for third-party feedback. There should be a plan.
This post would apply to other labs as well, but I'm focusing on Anthropic and OpenAI because they are ahead and have experienced alignment failures recently.
Of course the content would look a lot different. I am imagining a detailed document for technical approaches to aligning superintelligence (probably with a few different paths or bets). Plan A, however, isn't focused on technical alignment details.
Quite a few researchers and senior people at both companies have published personal opinions making it clear that, in their view, their employers do not have a complete and detailed plan for this.
This is not particularly surprising when doing research: for example, if you asked pharmaceutical companies for complete and detailed plan on how they will cure cancer, they wouldn't yet have one either. This doesn't prove that they won't be able to cure cancer eventually — but it does make it rather clear they're unlikely to do so in the next year or three.
While OpenAI and Anthropic[1] pursue different lines of safety research, they have yet to produce a public-facing document describing concretely how their companies plan to align superintelligence. I think it is underappreciated how this points to general negligence or a lack of openness to third-party feedback.
By “plan,” I mean a document describing a proposal for technical alignment with at least the level of detail and research effort of AI 2040.[2] Any such plan for technical alignment would likely be flawed in non-obvious ways. But having a proposal that’s sensible enough to consider and detailed enough to critique is a good starting point for wiser proposals. Making such a plan public would also create feedback loops for accountability.
The closest thing to a plan came in 2023, when OpenAI announced their superalignment strategy (also relevant). I do not find this approach particularly convincing, though I do find it laudable that OpenAI explained what they planned to do, who would lead the effort, and what resources would be allocated, at a level of detail which made critique possible.[3] This team no longer exists, and nowadays, as far as I am aware, the research community doesn’t have precise answers to basic questions like:
If OpenAI and Anthropic leadership do not think they could produce a technical alignment plan as detailed as AI 2040, they should communicate this loudly and publicly, stating where the uncertainties lie and how they plan to resolve them.
In light of recent events, it's self-evident that Anthropic and OpenAI cannot reliably control non-superintelligent systems. We don’t know if or how labs will act differently in response, but the public and the research community shouldn’t have to wonder. There should be transparency and room for third-party feedback. There should be a plan.
This post would apply to other labs as well, but I'm focusing on Anthropic and OpenAI because they are ahead and have experienced alignment failures recently.
Of course the content would look a lot different. I am imagining a detailed document for technical approaches to aligning superintelligence (probably with a few different paths or bets). Plan A, however, isn't focused on technical alignment details.
Still, I wish there would have been much more detail here.