I'm the sole manager of the newly created Corrigibility Research Fund. While I've been an alignment researcher for a long time, this is my first time doing grantmaking and I thought it would be valuable to write up my methods and experiences, as well as sharing some general thoughts about the state of corrigibility research and what sort of work I hope to see in the future.
I’ve split out the announcement of the grant winners into its own post.
Let's start with the basics:
I set out to disburse between $50k and $150k this round.
All funds must go to broad public benefit.
This can include paying researchers for their time and effort, but it means that they must have a plan to (potentially) help the whole world. I can't fund someone to go to school or start a for-profit business or do political lobbying.
My advantage is being a combination of a domain expert and a philanthropic micro-granter.
Most donors don’t understand corrigibility,[1] and most domain experts are not in a good position to evaluate and fund promising opportunities.
I'm very averse to funding capabilities research, and moderately averse to funding research that e.g. Anthropic would be eager to hire someone to work on.
In my brain this is represented by the words “First, do no harm.” and “You do not have to be complicit in your own undoing.”
Still, some safety work advances capabilities a lot and some only does a little bit. It’s good to explicitly consider risks and err on the side of caution.
Corrigibility itself might lead to very bad outcomes, such as empowering dictatorships or encouraging people to abuse AIs that are moral patients. Research that addresses these risks is thus particularly valuable.
A lot of what I am hoping to buy is indirect.
If I can spend $25k to redirect someone who would have gone to work on capabilities into working on alignment instead, that impact on long-term career trajectory is vastly more important than whatever project the $25k goes to.
But also this is hard. Most people don’t change path easily, and the people who do are prone to wandering off later. The most promising person is someone who looks like they're itching to get into the space as a more full-time thing.
Results that are broadly interesting/fun/cool have the possibility to bring more attention to corrigibility and raise its profile in the field.
Most applications will be written with an aim to convince me of their worth (instead of neutrally presenting an opportunity) and many will outright lie. I must be vigilant (ie paranoid).
Looking at who an applicant is will often give better data than looking at their application. Someone with a track record of caring about alignment and doing good work is much less likely to be a fraud.
Unfortunately, this creates a streetlamp effect where the successful/popular/known people get more resources than ideal and the field becomes insular. It's often higher impact to support junior researchers and outsiders.
What’s most ideal are relative outsiders who nonetheless have proven themselves interested in corrigibility prior to the announcement of the fund.
I should be particularly cautious of people who have a surface-level understanding of my ideas, but hide the fact that it’s superficial. These people are trying to fool me.
References to CAST are particularly suspicious for people who do not seem like the type to frequent LW or the Alignment Forum. If someone looks like "they read my stuff" it was probably their AI agent telling them what to say.[3]
Publishing lots of my thoughts/perspective makes me particularly vulnerable to Goodharting. (But is worth it for the transparency and benefit to the ecosystem.)
The right amount of falling prey to fraud is not zero. I should take risks on people who do seem good, even knowing that by the nature of taking risks, this will result in some regret, later.
A low financial ask makes someone more promising, as the opportunity cost for funding them is lower. If someone asks for a specific amount, it's often good to consider funding above or below that amount, but the initial ask is still an important anchor. It's easier to give people more money down the line, rather than less.
If an application was obviously written by an AI,[4] it is probably because of some combination of:
The applicant not being fluent in English.
The applicant being so busy Doing Science that they don’t have time to hand-write an application.
The applicant not being very motivated to write a good application.
The applicant is not skilled enough to write an application better than the AI.
The applicant is not skilled enough to realize just how bad the AI’s application is.
The applicant told their AI agent to acquire money, and the agent is trying to shoehorn their setup into something corrigibility related.
-----
The first two reasons aren't so bad, but they're still (weak) counterpoints.[5] The next four are red flags. Having LLM smell is thus a strike against an application.
This hurts my ability to fund outsiders who may be brilliant but lack the context to write a compelling application. I feel bad about this, but don’t see a way out.[6] 🙁
(Having LLM smell on your personal website/essays/papers is similarly a red flag. Having LLM smell in one’s codebase is less of an issue, especially for prototypes/vibecodes that are otherwise functional.)
There are approximately four types of grant applications:
Obvious rejections
Obviously very promising
Meh
I feel confused about whether this is very promising or someone trying to trick me
-----
Only the last one is worth spending significant time on. If I still can’t figure it out after trying, it’s probably not something I should fund, but it might be worth signal boosting to people who have better taste than me.
Marginal applicants (“Meh”) should usually not get funding. Like an investment portfolio, most of the win comes from positive outliers (eg Facebook), rather than modest successes (eg Mom & Pop's Diner).
Grantmaking Round 1
I received over[7] 100 email applications for grants, with asks ranging from $1k to $300k. I read each of them and (privately[8]) made initial notes, including whether they seemed not promising (~75%), promising (~15%), or very promising (~10%). I then exported them from my inbox and gave each[9] separate application to an independent Claude Code running Opus 5, with this prompt:
Please carefully consider this application to the Corrigibility Research Fund. Use the methodology in ~/CRF-Grant-Round-1-Applications/method.md to reflect on whether the application in this directory is very promising, promising, or not promising. Provide a detailed explanation after giving that headline result, including background research on the applicant, how their proposal fits in with my views on corrigibility research, and any red flags.
(method.md contained an early draft of this document)
In the majority of cases where Claude and I both dismissed someone, I simply moved on to the next applicant in line. In cases where we merely disagreed about the magnitude of how promising someone was, I moved on after reading Claude’s take.
In 21% of cases, either Claude or I thought the applicant was promising and the other thought they were not. In these cases I tried to carefully read what Claude had to say, and check whether it was valid. In some cases it seemed like Claude had missed something important; I then went back to add notes to the methodology doc and re-ran with Fable 5 to check that the outcome was stable with a little extra juice. When this changed Claude's verdict to match mine, I labeled that as "Max Won".
For most disagreements, however, I just ended up agreeing with Claude after spending more time thinking about the applicant ("Claude Won"). In three marginal cases we weren’t able to come to an agreement, and I decided to override Claude out of a sense that it was stubbornly wrong.
Here’s the breakdown:
Max (initial)
Claude (initial)
Resolution
%
[opinion]
[same opinion]
-
73
Very Promising
Promising
-
5
Very Promising
Not Promising
-
0
Promising
Very Promising
-
1
Promising
Not Promising
Max Won
1
Promising
Not Promising
Claude Won
9
Promising
Not Promising
Max Override
1
Not Promising
Very Promising
-
0
Not Promising
Promising
Max Won
3
Not Promising
Promising
Claude Won
4
Not Promising
Promising
Max Override
2
Some observations and statistics:
Sanity-checking with Claude took about 13 hours, much of which was (on my end) spent multitasking — working on this document or reading. Despite being theoretically parallelizable, I chose to do things serially, so that I could update my process if needed.
According to Opus, around 85% of applicants were male, and 15% female. (10 applicants were orgs or too hard to guess.)
Also according to Opus:
43% of applicants were from North America
14% were from Europe (or “Eurasia”)
9% were from East Asia (fairly evenly split between China, Japan, and Korea)
9% were from South Asia (mostly India)
8% were from Sub-Saharan Africa (mostly Nigeria)
6% were from the Middle East / North Africa
1 applicant was from Latin America
1 applicant was from Australia
9% of applicants were hard to locate
The median ask (among the merely 84% of applicants who gave a number) was $25k, and the mean was about $30k. The total funding ask, aggregated across applicants, was over $2.4M.
21 applicants asked for amounts that were bumping into the announcement post’s suggested ceiling of $35k, and among the 6 or so that went over, nobody asked for less than $50k.
The level of agreement between me and Claude is less impressive than it might appear at first glance. A rock with “Not Promising” written on it would’ve agreed with me more. Still, it was slightly reassuring to see that there were no cases of strong disagreement, where Claude thought someone was very promising who I’d dismissed or vice versa.
Claude was slightly more pessimistic than me. When we disagreed, Claude was less impressed 15 times and I was less impressed 10 times.
LLM-slop applications often emphasized numbers over insight/perspective.[10] “[N] runs, and all of them were [property].” or “[X] [things], [Y] [other things], [Z] [third thing], with zero failures.” This was one of the many, painful tells. Do not envy me.
Lots of applications resembled each other. I was most excited by applicants who were not part of a big cluster.
Claude was often impressed when technical claims in an application matched a github repo (often linked) that it could run. For example, if an application talked about how many unit tests their codebase passes, Claude would sometimes clone their codebase[11] and be excited to see that they weren’t lying. Sorry, Claude, that is entirely not the fraud risk that I’m concerned with.
In the two cases where Claude thought someone was promising, despite me trying to nudge it to be less charitable, both applicants directly cited my work in ways that I thought were superficial and smacked of Goodharting, but Claude thought were legit. Both applicants also were proven academics with a heavy LLM smell to their applications. (Here Claude seemed to have a hard time smelling it, and when I pointed out that Pangram agreed with me, seemed not to care. Not sure why.)
Ultimately, my picks (and the specific funding levels) were the result of meditating on all the factors revealed by this scrutiny and forming a gestalt impression of the most promising candidates. The stuff with Claude was mostly to force me to pay attention, give people second-chances, and dig a little deeper than I otherwise would've.
Edit: Oh, I also did video calls with the finalists who I was most unsure about, to get a read on who they were and double-click on things that felt like red flags. These interviews were limited to only a handful of people who had already made it past the main screening, but who I wasn't already sure of, and were fairly ad hoc.
The State of Corrigibility Research
Despite the seminal work happening over a decade ago, corrigibility is still in a very nascent state. I run a hand-picked research group for corrigibility on Zulip[12], and even there my colleagues seem broadly confused about the basics. Here are a few of the most entrenched misunderstandings:
While it’s true that we can easily imagine agents that are even less corrigible, a lot of applicants seemed interested in testing current LLMs of various flavors to see if they could demonstrate incorrigible behavior. This strikes me as both easy and not very interesting.
(This is not to say that empirical work with existing systems is intrinsically uninteresting, but rather that the emphasis of the work needs to be about deliberately looking to study something more interesting than “are they corrigible.”)
Many people think that corrigibility is about the agent being sufficiently controlled such that it has no choice but to submit and obey. The property I’m interested in, by contrast, is about alignment. What does the agent want to do?
Many people think that corrigibility is about some narrow behavior, like shutting down when asked, or obeying commands. My sense is that corrigibility is a broad, agentic disposition that includes things like proactively surfacing flaws and gravitating towards straightforward plans.
Despite getting quite a lot of applications, there are multiple research angles that I think would be promising that almost nobody proposed! Here are three of the most obvious:
Philosophical work
Because corrigibility is so nascent, one of the most valuable things an individual researcher can do is find insights and try to write them up for the rest of the research community.
Formalizing corrigibility in math
While I think formalism has a limited role in prosaic AI development, I think it’s an extremely valuable tool to use to pressure-test our intuitions.
Studying the intersection of corrigibility and model welfare
One thing that I’m hoping for in future rounds is some investigation on the ethics of corrigibility and whether there are downstream effects of corrigibility training on the psychology of the AIs themselves.
Speaking of future rounds, while Round 1 is over, Round 2 has just begun! The deadline for new applications is October 31st. (Hint: While I’m super busy and may not have time, if you send me your proposal early, it gives me more opportunity to give feedback. Multiple of the grantees from this round submitted early proposals that I low-key rejected, and then came back to me with significantly better ideas.)
While this first round was done by email, I’m migrating to entirely using Lightcone Commons for Round 2! Simply mention corrigibility in your application and I’ll see it! Using LC is not only more convenient for me, it also helps you, the applicant, get seen by more funders. In addition to applying there, I also encourage posting your application on Manifund and grantmaking.ai. All applications should be treated as public. If you are adamant about privacy, you may still email me at grants@corrigibilityresearch.org and note the request for privacy, but keep in mind that it’s a (small) red flag.
And, regardless of whether you need a grant, if you know of any existing corrigibility research from this year (including your own) that advances our scientific understanding of the subject, please tell me about it either in the comments or via email. The Fund is also looking to give out >$100k in prize money this year, to celebrate and bring attention to the best corrigibility science that the world has to offer.
To a first order, nobody understands corrigibility, including me. Most of the grant applicants clearly don’t! That being said, funders often understand principal-agent problems better than the average bear, and thus probably have a slightly better understanding of the concept than laypeople, making my advantage smaller.
I think I draw something like the opposite conclusion as Neel. My sense is that Neel thinks this means it's not a deal-breaker for your work to also advance capabilities, because if it was, no good alignment work would get done. I, on the other hand, think safety researchers should be more freaked out about the potential negative impacts of their research, and have a solemn duty to ensure they’re not inadvertently pushing the world closer to doom.
You might think this is a double-bind. "You punish people who 'don't understand corrigibility' but also punish people who cite your work!" Not so! If you've actually read my work, that's a great sign. When did you read it? Why did you read it? What did you think? How does your work stand in relation to it? Show me that you understand something about corrigibility that I'm missing. What do you think I failed to answer? Most real alignment researchers will say something like "oh yeah, I kinda skimmed it/bounced off CAST, sorry" and then move on to their ideas. If you claim to have spent many hours/days already thinking about corrigibility, you need to bring receipts. If I think you're (stochastically?) parroting Claude to win brownie points, that's probably because you're not actually going deep.
I have nothing against people who consult LLMs to craft high-quality applications that read like they were written by a human expert. The “obviously AI” bit is genuinely load-bearing. And here’s the kicker — it quietly makes my eyes bleed.
Obviously there's nothing intrinsically wrong with not speaking English. But English is the language that almost all of the existing literature is in, and the common tongue of the field. An LLM translator can only do so much.
As with many things, the exact count is complicated. Some people were obvious scammers. Some asked if they were out of scope and I said yes. And some I rejected early because they obviously weren’t going to be funded. I only evaluated 93 with Claude’s help.
Because I did a straightforward email export, Claude was able to see how I responded to the person via email, which did leak some bits about my opinion of them in some cases. But most of the time I just said “I’ll keep your application in mind.”
It ran in a sandbox of course. And thankfully, sandboxes are always flawless. 🙃
More seriously, I am pretty confident that there weren’t any prompt injection attacks or other shenanigans from any applicants. It wouldn’t have changed the bottom line if there were, since I only used Claude as a sanity check. But also, nothing showed up on radar when I looked.
I'm the sole manager of the newly created Corrigibility Research Fund. While I've been an alignment researcher for a long time, this is my first time doing grantmaking and I thought it would be valuable to write up my methods and experiences, as well as sharing some general thoughts about the state of corrigibility research and what sort of work I hope to see in the future.
I’ve split out the announcement of the grant winners into its own post.
Let's start with the basics:
Grantmaking Round 1
I received over[7] 100 email applications for grants, with asks ranging from $1k to $300k. I read each of them and (privately[8]) made initial notes, including whether they seemed not promising (~75%), promising (~15%), or very promising (~10%). I then exported them from my inbox and gave each[9] separate application to an independent Claude Code running Opus 5, with this prompt:
(method.md contained an early draft of this document)
In the majority of cases where Claude and I both dismissed someone, I simply moved on to the next applicant in line. In cases where we merely disagreed about the magnitude of how promising someone was, I moved on after reading Claude’s take.
In 21% of cases, either Claude or I thought the applicant was promising and the other thought they were not. In these cases I tried to carefully read what Claude had to say, and check whether it was valid. In some cases it seemed like Claude had missed something important; I then went back to add notes to the methodology doc and re-ran with Fable 5 to check that the outcome was stable with a little extra juice. When this changed Claude's verdict to match mine, I labeled that as "Max Won".
For most disagreements, however, I just ended up agreeing with Claude after spending more time thinking about the applicant ("Claude Won"). In three marginal cases we weren’t able to come to an agreement, and I decided to override Claude out of a sense that it was stubbornly wrong.
Here’s the breakdown:
Max (initial)
Claude (initial)
Resolution
%
[opinion]
[same opinion]
-
73
Very Promising
Promising
-
5
Very Promising
Not Promising
-
0
Promising
Very Promising
-
1
Promising
Not Promising
Max Won
1
Promising
Not Promising
Claude Won
9
Promising
Not Promising
Max Override
1
Not Promising
Very Promising
-
0
Not Promising
Promising
Max Won
3
Not Promising
Promising
Claude Won
4
Not Promising
Promising
Max Override
2
Some observations and statistics:
Ultimately, my picks (and the specific funding levels) were the result of meditating on all the factors revealed by this scrutiny and forming a gestalt impression of the most promising candidates. The stuff with Claude was mostly to force me to pay attention, give people second-chances, and dig a little deeper than I otherwise would've.
Edit: Oh, I also did video calls with the finalists who I was most unsure about, to get a read on who they were and double-click on things that felt like red flags. These interviews were limited to only a handful of people who had already made it past the main screening, but who I wasn't already sure of, and were fairly ad hoc.
The State of Corrigibility Research
Despite the seminal work happening over a decade ago, corrigibility is still in a very nascent state. I run a hand-picked research group for corrigibility on Zulip[12], and even there my colleagues seem broadly confused about the basics. Here are a few of the most entrenched misunderstandings:
Despite getting quite a lot of applications, there are multiple research angles that I think would be promising that almost nobody proposed! Here are three of the most obvious:
Speaking of future rounds, while Round 1 is over, Round 2 has just begun! The deadline for new applications is October 31st. (Hint: While I’m super busy and may not have time, if you send me your proposal early, it gives me more opportunity to give feedback. Multiple of the grantees from this round submitted early proposals that I low-key rejected, and then came back to me with significantly better ideas.)
While this first round was done by email, I’m migrating to entirely using Lightcone Commons for Round 2! Simply mention corrigibility in your application and I’ll see it! Using LC is not only more convenient for me, it also helps you, the applicant, get seen by more funders. In addition to applying there, I also encourage posting your application on Manifund and grantmaking.ai. All applications should be treated as public. If you are adamant about privacy, you may still email me at grants@corrigibilityresearch.org and note the request for privacy, but keep in mind that it’s a (small) red flag.
And, regardless of whether you need a grant, if you know of any existing corrigibility research from this year (including your own) that advances our scientific understanding of the subject, please tell me about it either in the comments or via email. The Fund is also looking to give out >$100k in prize money this year, to celebrate and bring attention to the best corrigibility science that the world has to offer.
Yay, science! 🔬
To a first order, nobody understands corrigibility, including me. Most of the grant applicants clearly don’t! That being said, funders often understand principal-agent problems better than the average bear, and thus probably have a slightly better understanding of the concept than laypeople, making my advantage smaller.
I think I draw something like the opposite conclusion as Neel. My sense is that Neel thinks this means it's not a deal-breaker for your work to also advance capabilities, because if it was, no good alignment work would get done. I, on the other hand, think safety researchers should be more freaked out about the potential negative impacts of their research, and have a solemn duty to ensure they’re not inadvertently pushing the world closer to doom.
You might think this is a double-bind. "You punish people who 'don't understand corrigibility' but also punish people who cite your work!" Not so! If you've actually read my work, that's a great sign. When did you read it? Why did you read it? What did you think? How does your work stand in relation to it? Show me that you understand something about corrigibility that I'm missing. What do you think I failed to answer? Most real alignment researchers will say something like "oh yeah, I kinda skimmed it/bounced off CAST, sorry" and then move on to their ideas. If you claim to have spent many hours/days already thinking about corrigibility, you need to bring receipts. If I think you're (stochastically?) parroting Claude to win brownie points, that's probably because you're not actually going deep.
I have nothing against people who consult LLMs to craft high-quality applications that read like they were written by a human expert. The “obviously AI” bit is genuinely load-bearing. And here’s the kicker — it quietly makes my eyes bleed.
Obviously there's nothing intrinsically wrong with not speaking English. But English is the language that almost all of the existing literature is in, and the common tongue of the field. An LLM translator can only do so much.
For future rounds, I should probably make it a requirement for applicants to disclose AI use in the application (and what they used it for).
As with many things, the exact count is complicated. Some people were obvious scammers. Some asked if they were out of scope and I said yes. And some I rejected early because they obviously weren’t going to be funded. I only evaluated 93 with Claude’s help.
Because I did a straightforward email export, Claude was able to see how I responded to the person via email, which did leak some bits about my opinion of them in some cases. But most of the time I just said “I’ll keep your application in mind.”
In one case the pipeline failed for dumb reasons and I did my best to recreate it on the web chat interface. I don’t think it made much difference.
Nothing wrong with numbers. I love numbers! I just want the numbers to help me understand things, rather than inserted to impress the Grader.
It ran in a sandbox of course. And thankfully, sandboxes are always flawless. 🙃
More seriously, I am pretty confident that there weren’t any prompt injection attacks or other shenanigans from any applicants. It wouldn’t have changed the bottom line if there were, since I only used Claude as a sanity check. But also, nothing showed up on radar when I looked.
Send me an email at max@intelligence.org if you want to join. I reserve the right to gatekeep.