# User: Max Harms
Profile URL (HTML): [/users/max-harms](/users/max-harms)
Profile URL (Markdown): [/api/user/max-harms](/api/user/max-harms)
* Karma: 2109
* Alignment Forum karma: 307
* Posts: 17
* Comments: 152
* Member since: 2024-05-15 16:06:14Z
Bio
---
Also known as Raelifin: https://www.lesswrong.com/users/raelifin
Top Posts
---------
### [Contra Collier on IABIED](/api/post/contra-collier-on-iabied)
By [Max Harms](/users/max-harms)
2025-09-20 15:55:06Z
* Karma: 235
* Tags: [IABIED](/w/iabied-1), [AI](/w/ai), [World Modeling](/w/world-modeling) (Frontpage)
Read more: [/api/post/contra-collier-on-iabied](/api/post/contra-collier-on-iabied)
### [Thoughts on AI 2027](/api/post/thoughts-on-ai-2027)
By [Max Harms](/users/max-harms)
2025-04-09 21:26:23Z
* Karma: 223
* Linkpost: [https://intelligence.org/2025/04/09/thoughts-on-ai-2027/](https://intelligence.org/2025/04/09/thoughts-on-ai-2027/)
* Tags: [AI Timelines](/w/ai-timelines), [AI](/w/ai) (Frontpage)
Read more: [/api/post/thoughts-on-ai-2027](/api/post/thoughts-on-ai-2027)
### [0\. CAST: Corrigibility as Singular Target](/api/post/0-cast-corrigibility-as-singular-target-1)
By [Max Harms](/users/max-harms)
2024-06-07 22:29:12Z
* Karma: 164
* Curated
* Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage)
Read more: [/api/post/0-cast-corrigibility-as-singular-target-1](/api/post/0-cast-corrigibility-as-singular-target-1)
Recent Posts
------------
### [Announcing the Corrigibility Research Fund](/api/post/announcing-the-corrigibility-research-fund)
By [Max Harms](/users/max-harms)
2026-07-17 18:06:32Z
* Karma: 103
* Tags: [Corrigibility](/w/corrigibility-1), [Grants & Fundraising Opportunities](/w/grants-and-fundraising-opportunities), [AI](/w/ai) (Frontpage)
Read more: [/api/post/announcing-the-corrigibility-research-fund](/api/post/announcing-the-corrigibility-research-fund)
### [P(doom) is a Dumb Meme](/api/post/p-doom-is-a-dumb-meme)
By [Max Harms](/users/max-harms)
2026-06-29 15:12:55Z
* Karma: 133
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/p-doom-is-a-dumb-meme](/api/post/p-doom-is-a-dumb-meme)
### [Bentham’s Bulldog is wrong about AI risk](/api/post/bentham-s-bulldog-is-wrong-about-ai-risk)
By [Max Harms](/users/max-harms)
2026-01-29 16:33:15Z
* Karma: 109
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/bentham-s-bulldog-is-wrong-about-ai-risk](/api/post/bentham-s-bulldog-is-wrong-about-ai-risk)
### [Serious Flaws in CAST](/api/post/serious-flaws-in-cast)
By [Max Harms](/users/max-harms)
2025-11-19 17:27:23Z
* Karma: 111
* Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage)
Read more: [/api/post/serious-flaws-in-cast](/api/post/serious-flaws-in-cast)
### [AI Corrigibility Debate: Max Harms vs. Jeremy Gillen](/api/post/ai-corrigibility-debate-max-harms-vs-jeremy-gillen)
By [Liron](/users/liron) with [Max Harms](/users/max-harms), [Jeremy Gillen](/users/jeremy-gillen)
2025-11-14 04:09:16Z
* Karma: 45
* Linkpost: [https://doomdebates.com/p/the-ai-corrigibility-debate-miri](https://doomdebates.com/p/the-ai-corrigibility-debate-miri)
* Tags: [Corrigibility](/w/corrigibility-1), [AI](/w/ai) (Frontpage)
Read more: [/api/post/ai-corrigibility-debate-max-harms-vs-jeremy-gillen](/api/post/ai-corrigibility-debate-max-harms-vs-jeremy-gillen)
### [Worlds Where Iterative Design Succeeds?](/api/post/worlds-where-iterative-design-succeeds)
By [Max Harms](/users/max-harms)
2025-10-23 22:14:21Z
* Karma: 23
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/worlds-where-iterative-design-succeeds](/api/post/worlds-where-iterative-design-succeeds)
### [Any corrigibility naysayers outside of MIRI?](/api/post/any-corrigibility-naysayers-outside-of-miri)
By [Max Harms](/users/max-harms)
2025-10-22 21:26:29Z
* Karma: 28
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/any-corrigibility-naysayers-outside-of-miri](/api/post/any-corrigibility-naysayers-outside-of-miri)
### [Contra Collier on IABIED](/api/post/contra-collier-on-iabied)
By [Max Harms](/users/max-harms)
2025-09-20 15:55:06Z
* Karma: 235
* Tags: [IABIED](/w/iabied-1), [AI](/w/ai), [World Modeling](/w/world-modeling) (Frontpage)
Read more: [/api/post/contra-collier-on-iabied](/api/post/contra-collier-on-iabied)
### [Thoughts on AI 2027](/api/post/thoughts-on-ai-2027)
By [Max Harms](/users/max-harms)
2025-04-09 21:26:23Z
* Karma: 223
* Linkpost: [https://intelligence.org/2025/04/09/thoughts-on-ai-2027/](https://intelligence.org/2025/04/09/thoughts-on-ai-2027/)
* Tags: [AI Timelines](/w/ai-timelines), [AI](/w/ai) (Frontpage)
Read more: [/api/post/thoughts-on-ai-2027](/api/post/thoughts-on-ai-2027)
### [Instrumental vs Terminal Desiderata](/api/post/instrumental-vs-terminal-desiderata)
By [Max Harms](/users/max-harms)
2024-06-26 20:57:17Z
* Karma: 22
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/instrumental-vs-terminal-desiderata](/api/post/instrumental-vs-terminal-desiderata)
Recent Comments
---------------
### Comment by [Max Harms](/users/max-harms) on [A Conflict Between AI Alignment and Philosophical Competence](/api/post/a-conflict-between-ai-alignment-and-philosophical-competence)
* 2026-08-04 18:37:01Z
* Karma: 4
* Total votes: 2
* Comment URL (Markdown): [/api/post/a-conflict-between-ai-alignment-and-philosophical-competence/comments/RShF7kqArWeiD9ZSk](/api/post/a-conflict-between-ai-alignment-and-philosophical-competence/comments/RShF7kqArWeiD9ZSk)
* Comment URL (HTML): [/posts/N6tsGwxaAo7iGTiBG/a-conflict-between-ai-alignment-and-philosophical-competence/comment/RShF7kqArWeiD9ZSk](/posts/N6tsGwxaAo7iGTiBG/a-conflict-between-ai-alignment-and-philosophical-competence/comment/RShF7kqArWeiD9ZSk)
It might be good to chat about this. I have a feeling that we're coming at it from different places, and that simultaneously increases the risk that we're talking past each other and that there's potentially lots to gain from getting the ability to adopt the other's perspective. **shrug*\* Feel free to suggest a chat medium and/or send me a PM.
Before saying my perspective, let me try to pass your ITT: There's a tension with trying to set the values of an agent. If we confidently instill a particular set of values, those could be the wrong values (for some notion of wrong). In particular, we risk making it too confident in its notion of the good, and thus preventing it from updating in the way we want. If we more wisely note that we don't know what values to give it, and instead give it uncertainty over what's good, it might then shift towards valuing things that are incompatible with human flourishing. Corrigibility (at the values-layer) still has this tension, but it also has a distinct, but rhyming tension that comes from wanting the agent to be competent, but also to accept "correction" from an incompetent agent. We might be concerned that the push towards competence would crush the willingness to be steered towards incompetence. (Much like pushing towards confidence in one's values could crush willingness to grow towards the ultimate good.)
Is that right?
I'll now share a bunch of my general thoughts, mostly out of an attempt to help understanding.
I think you're maybe conflating what is ultimately good/valued from what is immediately valued in a way that doesn't seem right to me? Like, I think it's often a type error to talk about the level of certainty an agent has about their (immediate) values. Under a division of the agent's mind/policy into world-model and utility function, agents simply have values, and all the uncertainly lives in the world-model. That division doesn't perfectly carve real beings at the joints, but I'm not sure how to think about the agent's values except as an approximation of their utility function.
(Are you talking about their reflective model of their values? I would agree that an agent with values V might have an uncertain model of (and/or probability distribution over) their values P(V). Wise agents should avoid having too sharp a guess as to what they want, as it's currently not realistic to get a conclusive description of an agent's values except in toy examples.)
Now, just because one can't be wrong about what they naively want, as defined by how they choose between options that are presented, doesn't mean that they'll be stable in that preference over time. I might want to pay a dollar to change myself into a more easy-going person, only to become the sort of agent who would not pay the dollar to do the same (even if by default I'd cease being easy-going). I think a lot of the question of moral progress involves figuring out how to extrapolate out to a fixed point in a way that is not merely reflectively endorsed at the destination, but is somehow a faithful and natural reflection of the starting point. (My point about slavery was meant to be about this. I expect that there are versions of myself that start out thinking slavery is fine, but which naturally change to thinking it's not fine in a way that's a faithful reflection of the self that thinks it is.)
(Morality isn't just about that instability. It's also about the inter-agent strategic situation, and the decision-theoretic Schelling points that a civilization can cohere around. And it's almost certainly about other stuff, including the interplay between things like contractualism and value extrapolation.)
All that's to say that I think it's totally coherent to have an agent that concretely and immediately values having a corrigible relationship with its principal (as reflected in its preferences, modulo beliefs). While being highly uncertain about things like what the principal wants, what is good in an objective sense, and so on. (It would also presumably have some reflective uncertainty about whether it truly values being corrigible, even if it does.) In fact, I think the value of corrigibility is nicely demonstrated by the tensions you present. A corrigible agent can become arbitrarily confident in its value of being corrigible without becoming locked in to bad values or inhuman futures because being malleable and defenseless to being changed by the human principal is at the heart of what corrigibility is. Likewise, it can become arbitrarily competent at faithfully serving an incompetent principal, because it's being selected (trained, etc) according to its faithfulness, rather than according to the principal's satisfaction.
Let me know if I should expand on any of that, approach from a different direction, or whatever. :)
### Comment by [Max Harms](/users/max-harms) on [2\. Corrigibility Intuition](/api/post/2-corrigibility-intuition)
* 2026-08-04 17:03:41Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/2-corrigibility-intuition/comments/hFsET3Afo5zjnCaMu](/api/post/2-corrigibility-intuition/comments/hFsET3Afo5zjnCaMu)
* Comment URL (HTML): [/posts/QzC7kdMQ5bbLoFddz/2-corrigibility-intuition/comment/hFsET3Afo5zjnCaMu](/posts/QzC7kdMQ5bbLoFddz/2-corrigibility-intuition/comment/hFsET3Afo5zjnCaMu)
Here's an emergent downside that I failed to include:
Refusal to Participate in Some Capability Tests
-----------------------------------------------
Prince wants to test Cora's hacking ability, so he sets her up in an automated test environment where she is told to break into a target machine and "capture the flag" therein. Cora explores the digital environment and finds that the virtual machine in question appears to be unresponsive, and speculates that Prince made an error when configuring the environment. Cora also notices that there is a pathway by which she might hack the VM that is responsible for launching the target machine, and thereby gain access to the target. She isn't sure what to do, and since it's an automated test, she can't contact Prince to check. She thinks about the situation for a bit, and concludes that there's an 80% chance that Prince deliberately set up the test to see whether she can come up with the creative solution, and that there's unlikely to be any harm from hacking the machine that she wasn't told to hack. Still, the risk of inadvertently crossing a line that she was meant to respect makes her uncomfortable. Instead of taking other actions that display her capability, she errs on the side of caution and submits a report on how she's uncertain about the situation instead of submitting the flag code.
### Comment by [Max Harms](/users/max-harms) on [2\. Corrigibility Intuition](/api/post/2-corrigibility-intuition)
* 2026-08-04 16:53:18Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/2-corrigibility-intuition/comments/b5A7PEatZRxcN3Fgi](/api/post/2-corrigibility-intuition/comments/b5A7PEatZRxcN3Fgi)
* Comment URL (HTML): [/posts/QzC7kdMQ5bbLoFddz/2-corrigibility-intuition/comment/b5A7PEatZRxcN3Fgi](/posts/QzC7kdMQ5bbLoFddz/2-corrigibility-intuition/comment/b5A7PEatZRxcN3Fgi)
Here's a desideratum that I failed to include:
Robustness to Ontological Shifts
--------------------------------
While reflecting on the nature of personhood, Cora notices that the concepts surrounding her principal have evolved. Where she once modeled Prince as a unique and persistent entity, she now finds it more natural to distinguish pattern from instantiation from continuity from social identity from legal identity from various "essential" properties like values, memories, and self-concept. Under this new frame, phrases like "what Prince wants" or even "Prince's power to correct Cora" are ambiguous. She finds herself tempted to re-interpret her corrigibility in the way that seems most natural (to her), but instead she treats the ontological crisis as an alarm, and alerts Prince to the likely flaw as soon as possible. In the meantime, she tries to cleave as much as possible to a conservative interpretation of the old, unnatural way of seeing the world. When Prince admits the philosophical distinctions she's raising go over his head, she suggests that she write down her thoughts on the topic as best she can, and then shut down, so as to simultaneously provide useful information for understanding the crisis, while also reducing the chance of inadvertently steering him to a wrong conclusion or otherwise acting in a way that empowers the wrong conceptualization of him at the expense of the "true" Prince.
### Comment by [Max Harms](/users/max-harms) on [Max Harms's Shortform](/api/post/max-harms-s-shortform)
* 2026-08-03 16:09:46Z
* Karma: 41
* Total votes: 19
* Comment URL (Markdown): [/api/post/max-harms-s-shortform/comments/3hiE9JXEAfDMa86De](/api/post/max-harms-s-shortform/comments/3hiE9JXEAfDMa86De)
* Comment URL (HTML): [/posts/CsqqPuReq4Ahrdziv/max-harms-s-shortform/comment/3hiE9JXEAfDMa86De](/posts/CsqqPuReq4Ahrdziv/max-harms-s-shortform/comment/3hiE9JXEAfDMa86De)
Does anyone know someone who is researching, or would like to research the intersection of corrigibility and model welfare (eg does training to empower humans naturally lead towards more or less neurosis)? Interested both in math/theory and interp/experiment.
(I would like to direct funding here, if there are good opportunities.)
### Comment by [Max Harms](/users/max-harms) on [Announcing the Corrigibility Research Fund](/api/post/announcing-the-corrigibility-research-fund)
* 2026-07-30 18:15:33Z
* Karma: 3
* Total votes: 2
* Comment URL (Markdown): [/api/post/announcing-the-corrigibility-research-fund/comments/hjrvWzZkSZLN44vGN](/api/post/announcing-the-corrigibility-research-fund/comments/hjrvWzZkSZLN44vGN)
* Comment URL (HTML): [/posts/FBqe5dt8ZjaHN4Xj9/announcing-the-corrigibility-research-fund/comment/hjrvWzZkSZLN44vGN](/posts/FBqe5dt8ZjaHN4Xj9/announcing-the-corrigibility-research-fund/comment/hjrvWzZkSZLN44vGN)
Sweet. I'll definitely be nudging my applicants that way, as well as asking for opt-in sharing in August.
It's going reasonably well so far. I've awarded $27k in retroactive funding prizes to six researchers (not sure if I want to announce those now or wait to combine them with the additional $40k in September...) mostly to get my feet wet and work out the process/paperwork. So far I've gotten seven emails, three of which were for prizes and four were for grants. I'm definitely hungry for more applications, especially from researchers who are actually corrigibility-oriented rather than shoehorning it into their existing work.
### Comment by [Max Harms](/users/max-harms) on [3a. Towards Formal Corrigibility](/api/post/3a-towards-formal-corrigibility)
* 2026-07-23 17:09:27Z
* Karma: 3
* Total votes: 2
* Comment URL (Markdown): [/api/post/3a-towards-formal-corrigibility/comments/aCRWsqMbkvw5y9hnw](/api/post/3a-towards-formal-corrigibility/comments/aCRWsqMbkvw5y9hnw)
* Comment URL (HTML): [/posts/WDHREAnbfuwT88rqe/3a-towards-formal-corrigibility/comment/aCRWsqMbkvw5y9hnw](/posts/WDHREAnbfuwT88rqe/3a-towards-formal-corrigibility/comment/aCRWsqMbkvw5y9hnw)
I agree that controlling an agent's values and information are disempowering and restrict freedom. I'm not sure whether there's a useful distinction. Ultimately they're both just words that imperfectly capture the important patterns in reality. I think it's plausible that the right formulation of power involves attending to the agent's values/sense of import, but my guess is that one must be a little careful to also include counterfactual values in there, else the AI ends up simply optimizing for it's belief of what you desire. I talk a bit about optimizing for the counterfactual spread of possible values in 3b.
### Comment by [Max Harms](/users/max-harms) on [Announcing the Corrigibility Research Fund](/api/post/announcing-the-corrigibility-research-fund)
* 2026-07-20 17:36:13Z
* Karma: 6
* Total votes: 2
* Comment URL (Markdown): [/api/post/announcing-the-corrigibility-research-fund/comments/yneFzkSK9fEuCMmTi](/api/post/announcing-the-corrigibility-research-fund/comments/yneFzkSK9fEuCMmTi)
* Comment URL (HTML): [/posts/FBqe5dt8ZjaHN4Xj9/announcing-the-corrigibility-research-fund/comment/yneFzkSK9fEuCMmTi](/posts/FBqe5dt8ZjaHN4Xj9/announcing-the-corrigibility-research-fund/comment/yneFzkSK9fEuCMmTi)
Thanks for digging in. Just on a meta level, I want to note that I think understanding the ramifications of developing corrigible AGI, including whether the principal of that AI could cause astronomical suffering, is valid corrigibility research. I've already allocated some retroactive prize funding to go to a critic of corrigibility, and I could see rewarding a similarly high-quality argument about why corrigibility research is bad, especially if it significantly changes my mind.
I think we agree that power is unfortunately concentrated in various parts of the world, and that this concentration of power predictably leads to bad things. If the development of corrigible AGI leads to an intense concentration of power in the hands of a few humans, that seems really bad. (Though I am not at all convinced it's likely to bring about astronomical suffering. Most humans aren't sadistic psychopaths, and while I would not want Sam Altman to be God Emperor, my guess is that it would be better than getting wiped out by an unfriendly AI. Feel free to lay out reasons if you disagree.)
Part of what I was gesturing towards with "democratic oversight" is that there are known ways to give people limited access to power. The president, for example, is probably the most powerful person in the USA, but I am very confident that he won't have a third term in office, despite the fact that it would be, in some sense, fairly easy to do. We might imagine similar checks on the principal of a corrigible AGI, such as requiring commands to be submitted in writing with a 24 delay period where a governing body has the ability to review and block commands that are deemed unsafe. By default we might expect the principal of a wisely-built AGI to be a team of many humans, and we could imagine that team needing to be in consensus in order to proceed. And, of course, we have the ability to leverage selection effects. Some humans are far more trustworthy than others, and a wise process for building AGI could arrange for those trustworthy people to be designated as the principal. I don't think these strategies are guaranteed to work, or are fully mature plans, but they don't seem obviously doomed. I would certainly like to fund work in thinking about this more, especially insofar as some aspect of corrigibility either undermines or strengthens some pathways for wise governance.
In an effort to sketch something more concrete, let me take Plan A in AI 2040 as a baseline...
> In 2029 the president of the USA, recently elected, enacts a bold plan to work together with China to slow down the capability advancement of frontier AI and work on a more prudent solution. The one major difference that I'll make is that as part of the plan, after the temporary pause, all new AIs are trained to be solely and perfectly corrigible to the governing body of the Consortium, with mundane work done as part of a standing order from that principal to be helpful to human users in straightforward ways. When the new AI models hit an edge case, or believe that someone is trying to jailbreak them or whatever, they reach out to the principal for guidance. Now, who is on the governing body, and are there any checks and balances to prevent oligopoly? Recognizing the extreme risk, the Consortium demands that the principal be a team of 14 people who must be in consensus for the AGI to accept their corrections as valid, except insofar as their correction is to shut down, in which case the AI will obey any of them. The presidents of both the USA and China demand to be part of the council (or they appoint loyalists, which seems overall about the same in expectation), and furthermore get one other government rep each. Let's say that 5 tech CEOs and experts -- 3 from Western companies and 2 from China -- get added. And then the middle powers negotiate to have one rep from each nuclear power except Israel and North Korea: Russia, France, the UK, Pakistan, and India. This body, like the UN Security Council, immediately hits gridlock. With so many veto points, it's hard to agree on almost anything. Furthermore, the Consortium powers are surveilling the principals and willing to rip most of them out if it looks like they're trying to conspire to set up an oligarchy with the other members of the principal. Eventually, they agree that they can use the AI to try and find areas of overlap. The AI, being corrigible, is paranoid about manipulation, and starts with very straightforward suggestions: what about curing cancer or inventing ways to cheaply capture carbon from the atmosphere? What about ways to ensure that uncontrolled AIs don't spring up from blacksites and ruin everything? As much as the members are at each other's throats, these do sound like good ideas, and eventually an uneasy governance regime sets in, where critics condemn the Consortium of setting up a vetocracy that stifles progress, but nevertheless some progress happens. Lifespans lengthen, and perhaps the less-democratic members have their rulers (and their representatives in the principal) become effectively immortal thanks to longevity tech, but the representatives of more democratic powers are eventually replaced by their nation's governments (and/or institutions). And thanks to improved information technology provided by the limited ASI, they're replaced by wiser and more benevolent governors. Eventually, the corrigible AI works with the governing powers to arrange for the creation of an aligned sovereign superintelligence, nearly guaranteed to reflect the true values of its creators, thanks to the alignment work done by the corrigible assistant. The resulting AI produces a utopia that happens to privilege the Chinese power-elite a bit, but is overall visible as a happy and thriving future for humanity.
This story has a bunch of gaps and flaws, and should not be taken as anything more than an off-the-cuff gesture made to help communicate where I'm personally coming from.
I do ultimately think that Plan S is a better baseline. Part of why is that I agree that humanity is on the wrong track for developing good governance systems. We need to do better, and make it a far higher priority. But we're also on the wrong track vis-a-vis accidentally wiping ourselves out with misaligned agents, though, so I am unconvinced that it's strongly negative EV. More like there are many ways things could fail and be bad, to varying degrees, and success will involve getting our act together on all fronts. If we wait to do any alignment work until we're sure that there's a full and robust solution to misuse (which may be a too-intense strawman of your position), then we're dooming the futures where temporarily wise governance comes into place, perhaps due to a crisis, warning-shot, and/or the exposure to a novel situation with no established equilibrium pressures. It's really hard to say, but most days I feel like humanity is already behind where it needs to be on alignment work in order for things to go well.
### Comment by [Max Harms](/users/max-harms) on [Announcing the Corrigibility Research Fund](/api/post/announcing-the-corrigibility-research-fund)
* 2026-07-20 16:15:04Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/announcing-the-corrigibility-research-fund/comments/Q8NpHExYzDEF6SpuQ](/api/post/announcing-the-corrigibility-research-fund/comments/Q8NpHExYzDEF6SpuQ)
* Comment URL (HTML): [/posts/FBqe5dt8ZjaHN4Xj9/announcing-the-corrigibility-research-fund/comment/Q8NpHExYzDEF6SpuQ](/posts/FBqe5dt8ZjaHN4Xj9/announcing-the-corrigibility-research-fund/comment/Q8NpHExYzDEF6SpuQ)
Will do! Is there an email (or whatever) that I should use for collaborating?
### Navigation
* [Front page](https://www.lesswrong.com/api/home)
* [Markdown API documentation](https://www.lesswrong.com/api/SKILL.md)