The Hugging Face incident has brought the idea of corrigibility to the forefront of popular discourse, my inbox is filled with newsletters about Genie Coefficients and Safe Scaffolds.
I think now is a particularly good time to make sure I have a good understanding of its limitations and dangers, particularly because I agree that it's the best path forward we have available to us. This post is my attempt to dive into the discussion and make sure I fully understand what's going on. Please correct me if you notice anything missing or wrong.
This post will not make much sense if you haven't at least read The CAST Strategy by Max Harms (corrigibility as a single target).
0 - Corrigibility is capabilities research
A very brief description of corrigibility is "take HHH, but drop harmless and honest to focus entirely helpful." This is the perspective that the mainstream is taking on corrigibility. What sets corrigibility apart from helpfulness is:
the idea of a principle
an effort to understanding the deep principles involved
Industry is very interested in making sure their agents are in some sense corrigible/helpful, and are investing resources into making sure they reliably do what users tell them to do. Given that, what is the role of the corrigibility research?
The most important work is to put down the theoretical groundwork that "goes on to influence the researchers and engineers at frontier labs in years to come, helping them ensure the first artificial general intelligences are corrigible and safe."[1] See: Open Corrigibility Questions
If this line of research is still too capabilities-adjacent for you, to me the main alternative here is not any other approach to AI safety, but instead to spend time directly on stopping AI development.
1 - Corrigibility is not crisp
The first criticism you will find of corrigibility is Max Harms' own post dramatically labeled Serious Flaws in CAST, the main one he points out is that his own, nor any of the other proposed formalisms of corrigibility, are any good.
This only means there's more work to be done, unless it turns out the reason no one has found a formalism is because corrigibility is not as crisp of a concept as Max Harms believes. Perhaps under further investigation corrigibility turns out to be just as hopelessly messy as morality.
If this is the case... Would that mean it's just as dangerous as FAI (Friendly AI)? Maybe there is a large section of goal space that is safe? On the other hand maybe there is no safe target in this direction?
2 - Corrigibility is anti-natural
"I think of this original idea of corrigibility as being kinda similar to rule utilitarianism. The difficulty of stable rule utilitarianism is that act utilitarianism is strictly better, if you fully trust your own beliefs and decision making algorithm. So to make a stable rule utilitarian, you need it to never become confident in some parts of its own reasoning, in spite of routinely needing to become confident about other beliefs. This isn’t impossible in principle (it’s easy to construct a toy prior that will never update on certain abstract beliefs), but in practice it’d be an impressive achievement to put this into a realistic general purpose reasoner. In this original version there is no “attractor basin” around corrigibility itself. In some sense there is an attractor basin around improving the quality of all the non-corrigibility properties, in that the engineers have the chance to iterate on these other properties."[2]
(I'm 80% sure that this is what people mean when they say Corrigibility is anti-natural)
3 - Power Concentration
One of the big dangers of CAST is that of power concentration, not in the hands or a company, or of an institution, but in the hands of a single individual or at the very most a small group. See Seth Herd and Vladimir's debate on if this is a good idea.
The alternative is to have a proliferation of AI capabilities, but that caries its own risks (again, see the debate for good summary of the issue). If none of the alternatives seem feasible to you, then corrigibility would not be a good idea.
4 - Choosing a principle is technically hard
We haven't even begun to figure out how to train an LLM to have anything like a principle. Part of the advantage of corrigibility is we can start to implement it now with current techniques, but there's no evidence that this claim is true for the process of principle identification.
Like with the formalism, this may be because the idea of a principle is hopelessly confused and fully impossible to implement in practice. This seems unlikely to me to be in the general case, but it may be infeasible to do using gradient descent and RLHF, at least not without significant changes to how we are doing the pipelines. Is there any serious thought about this?
5 - Choosing a principle is politically hard
The plan as stated in the CAST sequence is to have these billion dollar companies hand off control of their products to a smart high integrity person (or small collective of people). Reality and history seem to dictate precisely not that happening. I can't imagine a board of directors agreeing to that plan, it just sounds too weird.
As discussed due to the power concentration this is extra super important to get right. And far as I can tell no work has been done into considering a plan for how to do this. The best I've found is Red Heart, and the situation presented in the book is very different from the current political situation we find our selves in.
6 - Corrigibility isn't the easiest target to hit
Despite the risks of power concentration, the argument for corrigibility hangs on the fact that it's supposedly the easiest target to hit. The main reasons we think this is true are: (1) corrigibility most likely has an attractor basin around it (2) it's probably a simple concept
Corrigibility isn't the only way to think of the "Helpfulness" pillar. It would be easier to aim for an Instruction-Following ASI, but the is whether it's a good enough target to not die[3].
The main two alternative proposals are to focus either on "Harmless" or "Honest":
The maximalist version of Harmless is encoding human morality into the AI, a notoriously thorny issue. If we somehow could get Friendly AI that would be incredibly, but of all the options for alignment, this one is by far the most difficult to get right.[4]
Honest (aka, an Oracle) to me seems like the most plausible alternative route. The main intuition here is that if you have a 100% honest agent, before it takes any action you can ask it the intended effects of its action and stop it if it's misaligned.
Honest lacks any meta desires for changing itself, meaning there's even less likely to be an attractor basin, but on the other hand has a crisper target[] to aim for, unlike present day corrigibility. Honesty could in practice be used like a corrigible agent, if we knew the right questions to ask, but leaves its self much more open to accidental world ending mistakes.[5]
The Hugging Face incident has brought the idea of corrigibility to the forefront of popular discourse, my inbox is filled with newsletters about Genie Coefficients and Safe Scaffolds.
I think now is a particularly good time to make sure I have a good understanding of its limitations and dangers, particularly because I agree that it's the best path forward we have available to us. This post is my attempt to dive into the discussion and make sure I fully understand what's going on. Please correct me if you notice anything missing or wrong.
This post will not make much sense if you haven't at least read The CAST Strategy by Max Harms (corrigibility as a single target).
0 - Corrigibility is capabilities research
A very brief description of corrigibility is "take HHH, but drop harmless and honest to focus entirely helpful." This is the perspective that the mainstream is taking on corrigibility. What sets corrigibility apart from helpfulness is:
Industry is very interested in making sure their agents are in some sense corrigible/helpful, and are investing resources into making sure they reliably do what users tell them to do. Given that, what is the role of the corrigibility research?
The most important work is to put down the theoretical groundwork that "goes on to influence the researchers and engineers at frontier labs in years to come, helping them ensure the first artificial general intelligences are corrigible and safe."[1] See: Open Corrigibility Questions
If this line of research is still too capabilities-adjacent for you, to me the main alternative here is not any other approach to AI safety, but instead to spend time directly on stopping AI development.
1 - Corrigibility is not crisp
The first criticism you will find of corrigibility is Max Harms' own post dramatically labeled Serious Flaws in CAST, the main one he points out is that his own, nor any of the other proposed formalisms of corrigibility, are any good.
This only means there's more work to be done, unless it turns out the reason no one has found a formalism is because corrigibility is not as crisp of a concept as Max Harms believes. Perhaps under further investigation corrigibility turns out to be just as hopelessly messy as morality.
If this is the case... Would that mean it's just as dangerous as FAI (Friendly AI)? Maybe there is a large section of goal space that is safe? On the other hand maybe there is no safe target in this direction?
2 - Corrigibility is anti-natural
"I think of this original idea of corrigibility as being kinda similar to rule utilitarianism. The difficulty of stable rule utilitarianism is that act utilitarianism is strictly better, if you fully trust your own beliefs and decision making algorithm. So to make a stable rule utilitarian, you need it to never become confident in some parts of its own reasoning, in spite of routinely needing to become confident about other beliefs. This isn’t impossible in principle (it’s easy to construct a toy prior that will never update on certain abstract beliefs), but in practice it’d be an impressive achievement to put this into a realistic general purpose reasoner. In this original version there is no “attractor basin” around corrigibility itself. In some sense there is an attractor basin around improving the quality of all the non-corrigibility properties, in that the engineers have the chance to iterate on these other properties."[2]
(I'm 80% sure that this is what people mean when they say Corrigibility is anti-natural)
3 - Power Concentration
One of the big dangers of CAST is that of power concentration, not in the hands or a company, or of an institution, but in the hands of a single individual or at the very most a small group. See Seth Herd and Vladimir's debate on if this is a good idea.
The alternative is to have a proliferation of AI capabilities, but that caries its own risks (again, see the debate for good summary of the issue). If none of the alternatives seem feasible to you, then corrigibility would not be a good idea.
4 - Choosing a principle is technically hard
We haven't even begun to figure out how to train an LLM to have anything like a principle. Part of the advantage of corrigibility is we can start to implement it now with current techniques, but there's no evidence that this claim is true for the process of principle identification.
Like with the formalism, this may be because the idea of a principle is hopelessly confused and fully impossible to implement in practice. This seems unlikely to me to be in the general case, but it may be infeasible to do using gradient descent and RLHF, at least not without significant changes to how we are doing the pipelines. Is there any serious thought about this?
5 - Choosing a principle is politically hard
The plan as stated in the CAST sequence is to have these billion dollar companies hand off control of their products to a smart high integrity person (or small collective of people). Reality and history seem to dictate precisely not that happening. I can't imagine a board of directors agreeing to that plan, it just sounds too weird.
As discussed due to the power concentration this is extra super important to get right. And far as I can tell no work has been done into considering a plan for how to do this. The best I've found is Red Heart, and the situation presented in the book is very different from the current political situation we find our selves in.
6 - Corrigibility isn't the easiest target to hit
Despite the risks of power concentration, the argument for corrigibility hangs on the fact that it's supposedly the easiest target to hit. The main reasons we think this is true are: (1) corrigibility most likely has an attractor basin around it (2) it's probably a simple concept
A lot of the pushback against corrigibility has been around the idea of the attractor basin. See The corrigibility basin of attraction is a misleading gloss by Jeremy Gillen for the best arguments against it.
But if not corrigibility, what are the alternatives?
TAMing The Alignment Problem gives a good list of the general targets for alignment that need to be considered.
Corrigibility isn't the only way to think of the "Helpfulness" pillar. It would be easier to aim for an Instruction-Following ASI, but the is whether it's a good enough target to not die[3].
The main two alternative proposals are to focus either on "Harmless" or "Honest":
The maximalist version of Harmless is encoding human morality into the AI, a notoriously thorny issue. If we somehow could get Friendly AI that would be incredibly, but of all the options for alignment, this one is by far the most difficult to get right.[4]
Honest (aka, an Oracle) to me seems like the most plausible alternative route. The main intuition here is that if you have a 100% honest agent, before it takes any action you can ask it the intended effects of its action and stop it if it's misaligned.
Honest lacks any meta desires for changing itself, meaning there's even less likely to be an attractor basin, but on the other hand has a crisper target[] to aim for, unlike present day corrigibility. Honesty could in practice be used like a corrigible agent, if we knew the right questions to ask, but leaves its self much more open to accidental world ending mistakes.[5]
https://www.lesswrong.com/posts/FBqe5dt8ZjaHN4Xj9/announcing-the-corrigibility-research-fund
https://www.lesswrong.com/posts/oLbpfPkdtcknABvvw/the-corrigibility-basin-of-attraction-is-a-misleading-gloss, also Jeremy Gillen apologies to any philosophers for his abuse of rule utilitarianism.
https://www.lesswrong.com/posts/CSFa9rvGNGAfCzBk6/problems-with-instruction-following-as-an-alignment-target
Do I need to cite this one? Read anything by Yudkowsky.
Can anyone point me to good writing on honest ai? I couldn't find much. Am I using the wrong term?