August 23rd was the deadline for the first round of grant applications for the Corrigibility Research Fund. I received, read, and analyzed over 100 applications, asking for around 2.5 million dollars in total. Given that the fund was only able to allocate between $50k and $150k, I was unfortunately only able to award grants[1] to eight applicants, totaling $132k. I wrote a separate post about my experience and process.
There are many more applicants that I thought were worthy of funding, and I've advised many of them to re-submit applications to Lightcone Commons, which is also the platform that the Corrigibility Research Fund will be using, moving forward. If you're a funder, I suggest browsing the applications there. (I also recommend Manifund and grantmaking.ai, both for funders, and for those seeking funding.) If you're particularly interested in funding corrigibility and alignment research, feel free to reach out to me about regranting through the fund at grants@corrigibilityresearch.org.
In an earlier draft of this post I had a dedicated section for each grantee where I considered drawbacks and risks, but I found that I was mostly repeating myself. The risks I see fall into general patterns, that aren't very specific to one applicant or another:
Many of these grantees are not established in the AI safety landscape, and there's a good chance that they won't deliver much of import.
But also, helping outsiders get into the field and focus on corrigibility is one of the primary things that the fund is trying to do, so I think this risk is a good sign.
Being not-very-established, there's a risk that the most promising people will go on to work at Anthropic, or otherwise contribute to AI capabilities.
While I tried to steer away from proposals that seemed like they were significantly about capabilities, I was less concerned with funding people who have some track record of working on doing cool things with LLMs. In part this is because it's a good indicator of interest and skill, but also because I think it's good for alignment work to siphon off the most promising AI researchers, assuming that they would want to work on alignment, all else being equal.
Those who are more established have various disagreements with me about the nature of corrigibility. Thus, there's some chance that by funding them I'm giving their ideas legitimacy and that'll end up setting the field back.
This is also the right sort of problem to have; groupthink is bad in a multitude of ways. I trust in science and reason to show each of us the ways in which we're mistaken, if we live long enough to be so lucky.
Work that grantees do to build benchmarks or otherwise advance our understanding of corrigibility might get adopted by frontier labs in a way that either makes their AIs seem corrigible without being corrigible, or they succeed and then use their corrigibile AIs to do bad things.
This is just the general risk of corrigibility research. I continue to think it's important to learn about corrigibility, for the sake of having a long-term plan, but I'm sensitive to the risks here, and would love it if there were better guarantees about how research gets used.
Okay! With all that out of the way, let's meet this round's grantees! (In random order.) With each, I'll be presenting a summary of their application, along with a commentary about why I chose them and what I'm excited about.
Ian is building the first comprehensive behavioral benchmark of corrigibility and will attempt to fine-tune open-weight models with corrigibility as a singular target (CAST). The benchmark combines an expert-validated multiple-choice test, model self-reports, and, as its core component, agentic scenarios in realistic simulated environments scored by an LLM judge against hand-authored rubrics. The fine-tuning work uses a synthetic data pipeline and a purpose-built CAST constitution to instill corrigibility as a stable identity and test whether theoretically predicted properties such as obedience, transparency, and resistance to instrumentally convergent drives actually emerge.
Ian first reached out to me in March, after I appeared on the 80k hours podcast. He was pursuing corrigibility as part of his graduate thesis, and soon became a primary colleague in discussing CAST. While he is new to the field and has no substantial public work on the topic, what he's shared with me directly has impressed me, and I'm excited to support him working on corrigibility full time.
Jared Glover (MIT PhD '14) founded CapSen Robotics in 2014 and is transitioning his career into AI alignment. He has self-funded two research alignment projects this year: one on augmenting LLMs with emotional memory, and another on LLM honesty detection and steering. In this new project, he will use both goal-oriented and open-ended scenarios to map the activation-level geometry of corrigibility and related abstractions in LLMs, along with the impact of corrigibility and corrigibility-adjacent steering on the geometry of LLM values and goals.
Jared reached out to me with an initial proposal, building off his earlier work, to test how the use of corrigibility steering vectors would change the behavior of models, and what impact that might have on things like honesty. I was initially a bit skeptical; much of what he was describing, from my perspective, looked more like generic interp and control work — the sort of stuff Anthropic has been doing. Still, I invited him to my research group and to chat with me more, and I was pleasantly surprised by his thoughtful, independent perspective and impressed by his technical understanding. And after talking about how to refine his proposal to hit corrigibility more directly and scientifically, I'm very excited to see what he finds, both in the near future and in years to come.
Alfaxad is an AI engineer based in Tokyo, Japan. He is exploring the use of AI-native games as a new medium for AI safety research. He will develop Corrigibility Games, a set of AI-native games designed as environments for studying the principal-agent correction relationship.
Alfaxad has a very strong hacker/builder track record, given his age. (I don't know how old he is, but c'mon. Nobody my age is this cool.) He impressed me with his quick uptake on the subtleties of corrigibility when I provided feedback on his initial proposal, and his (amended) proposal to work on games that demonstrate aspects of corrigibility (as distinct from nearby concepts like pure obedience) impressed me with its novelty. Just as I wrote Red Heart to try to provide an accessible introduction to corrigibility and its nuances, I am excited by the prospect of one or more games that can do the same, in a fast and engaging way. With luck, Alfaxad's work will draw attention and build understanding in other young hackers who can help make our AI safer.
Florian, the author of Split Personality Training, is looking to study novel training architectures where models have a dedicated communication channel for flagging when they see their behavior as misaligned or otherwise problematic. The presence of this channel may allow models to more effectively surface their flaws, allow for improved training pipelines, and potentially resist learning misaligned behaviors even when presented with poisoned or problematic off-policy training data.
I enjoyed reading the Split Personality Training paper when it first came out, and am generally a fan of Florian, as an alignment researcher. His approaches have a good mixture of being concrete/prosaic/empirical and clever/interesting/theoretical, and he clearly has a lot of ideas. While I am primarily interested in the possibility of training a dedicated communication channel as a way to facilitate corrigibility on the architectural level, I'm also excited to simply have him thinking about the topic. As an independent researcher, he has both a good track record and a dependency on grants like this one. I hope that he gets a more substantial level of support from a big funder, so that he can contribute even more insight.
Ben, a math/philosophy undergrad at the University of Michigan, is (in addition to his conceptual work on corrigibility) looking to train an LLM to be corrigible as its singular target. This effort will likely involve developing a dataset that emphasizes corrigibility over nearby concepts and fine-tuning an open-source model on that data. The goal of the project is to collect evidence about whether current LLMs can be trained to behave significantly more corrigibly, to see how far the learned concept generalizes, and to identify the nearby concepts that distract from corrigibility.
Ben first reached out to me, like Ian Kahn, after my appearance on the 80k podcast, and was interested in studying and writing about corrigibility (from a purely theoretical direction) over the summer as a way to dip his toes into the field. His thoughts and essays over the last few months have improved my understanding of the subject, and I'm very excited to have him continue on that path. His application to train an LLM came as a little bit of a shock, and I'm hoping that it doesn't distract him from his more theoretical pursuits. But, as school starts up again, I selfishly want him to spend lots of time continuing to think about corrigibility, and I expect that funding him to play around with fine-tuning is a way to help make that happen. And who knows, maybe he'll beat Ian to the finish line on 1/12th the budget![2]
Xuanyi Wang and Shuo Li Liu will spend eight weeks attempting to formalize corrigibility. Their approach will be to ground corrigibility in observable quantities that can be empirically measured, using artifacts like system prompts, correction transcripts, constitutions, and rubrics as operational proxies for what the principal wants. They will also attempt to prove that corrigibility yields downstream properties such as shutdownability and safe instruction-following, calibrate the boundary between corrigibility and nearby alignment concepts, and verify their proofs in Lean. The goal is to turn corrigibility into a well-defined, falsifiable object that experts can debate and that can eventually be measured against real pipeline data.
Xuanyi is a recent philosophy graduate in Beijing with industry experience working on LLMs, and a particular interest in AI governance. Shuo is an economist at Princeton specializing in decision theory. While I'm unsure that they can make headway on a formalism, I was impressed by their enthusiasm and ambition. (Special thanks to Xuanyi for staying up past 2:00am to jump on a call with me with almost no advance notice.) Their first submission was pretty marginal, but after I nudged them to engage more with corrigibility as a concept, Shuo in particular seems to have gotten nerd-sniped by the topic and their revised proposal is very promising, to my eyes. I'm also hopeful that this will be the start of a pivot for one or the both of them towards working in AI alignment full-time.
Nick R.
Benevolent Override Pilot Study — $1.5k
Nick, a physician with clinical experience in palliative care, will run a small pilot study of (simulated) cases where an AI overrides a person's decision, or an authorized constraint on its behavior, because it predicts doing so will substantially improve human welfare. This study will likely involve constructing and testing various scenarios for frontier LLMs, as well as engaging with existing literature. The pilot's goal is to determine whether the concepts involved are coherent and empirically measurable, and to share any resulting insights with the broader research community, likely in the form of one or more essays, papers, or memos.
Nick has been following AI safety from afar since 2023, but is a relative novice in terms of doing work in the field. When he reached out to me in August, I was impressed by his careful and transparent communication, such as the way he flagged possible reasons not to fund him, marked the parts of his application that were the product of conversing with LLMs, and sought my feedback. He also asked for a remarkably small grant. This pilot project is, from my perspective, mostly[4] a speculative bet on him upskilling and pivoting his career towards additional safety work.
Rubi Hudson, author of Corrigibility Transformation: Constructing Goals That Accept Updates, recently launched Principled Agents, a research group focusing on AI alignment from a largely corrigibility-centric angle. Their current research centers on eliciting predictions that cover only outcomes relevant to a given decision, without extraneous details, helping identify the considerations that drive an agent’s actions and then misalignments between agents. The next project they have planned is investigating incentives for taking actions that are reversible in the relevant ways, in the vein of impact regularization. Success would lead to agents attempting to preserve their own corrigibility and that of subagents, as well as preventing other undesirable actions based on mistaken goals (e.g. launching nuclear weapons).
I'm a bit embarrassed to note that I missed Rubi's corrigibility paper when it came out last year. It's very impressive in many respects, and I want a hundred more like it! (I also have several quibbles with Rubi's take on corrigibility, and am currently debating it with him on LessWrong.) This $20k is something of a vote-of-confidence to give them a little more runway and to nudge other, bigger funders to send more substantial support to their newborn research org. They're doing actual alignment research! What a concept! Send funding!
Honorable Mentions
Other applicants who I decided not to fund this round, but who were promising enough for me to strongly consider, and who I hope other funders consider[5] backing:
Hou-Hsien Hsu — Effort Asymmetry in Multi-turn LLM Interaction
Danielle Franklin — Survey of Lay Intuitions About Corrigibility
Maryam Ilegbodu — Measuring Model Correction-Responsiveness Across Capability
Saman Seshadri — Does deep corrigibility training make models harder to correct?
Patrick Keogh — Sycophancy as an Oversight-Channel Failure
Sankalp Gilda — Corrigibility failures are rare events, and we average them away.
Zhen Wang — Corrigibility Evals for Long-Horizon Research Agents
Darius Kianersi — Do AI delegates stay corrigible to their principals?
Ekaterina Kalugina — Impossible Task Eval for Corrigibility
I'm sure many other qualified people applied. My sincerest apologies to anyone who ought to have been funded, but whom I overlooked. Again, I encourage everyone who is interested in being funded for Round 2 (deadline: October 31st), or for AI alignment funding more generally, to apply on Lightcone Commons.
Technically I only advise Lightcone Infrastructure on how to disburse the funds, and they have the final call. But in practice I'm the fund's sole manager.
More realistically, I hope they, and the other researchers looking to experiment with training for corrigibility, share what they have, work together when possible, and help each other advance the broader field.
I think they mean something more like a formal model of how a corrigible agent makes decisions, rather than the kind of philosophical decision theory usually focused on in LW circles. On a call with Shuo, I got the impression that he's been familiar with FDT/LDT (as well as CDT & EDT) for many years, but usually thinks about the subject from more of an economics perspective, which (combined with linguistic barriers) might explain the use of "decision theory" as a central term in a project that I don't really think of as being centrally about decision theory.
"Benevolent override" scenarios don't strike me as the most interesting topic, but I still think the object-level is interesting enough to warrant a small pilot project. Even just getting more thoughts from junior researchers seriously engaging with corrigibility would be valuable for identifying primary confusions/obstacles.
August 23rd was the deadline for the first round of grant applications for the Corrigibility Research Fund. I received, read, and analyzed over 100 applications, asking for around 2.5 million dollars in total. Given that the fund was only able to allocate between $50k and $150k, I was unfortunately only able to award grants[1] to eight applicants, totaling $132k. I wrote a separate post about my experience and process.
There are many more applicants that I thought were worthy of funding, and I've advised many of them to re-submit applications to Lightcone Commons, which is also the platform that the Corrigibility Research Fund will be using, moving forward. If you're a funder, I suggest browsing the applications there. (I also recommend Manifund and grantmaking.ai, both for funders, and for those seeking funding.) If you're particularly interested in funding corrigibility and alignment research, feel free to reach out to me about regranting through the fund at grants@corrigibilityresearch.org.
In an earlier draft of this post I had a dedicated section for each grantee where I considered drawbacks and risks, but I found that I was mostly repeating myself. The risks I see fall into general patterns, that aren't very specific to one applicant or another:
Okay! With all that out of the way, let's meet this round's grantees! (In random order.) With each, I'll be presenting a summary of their application, along with a commentary about why I chose them and what I'm excited about.
Ian Kahn
Testing Corrigibility as a Singular Target — $35k
Ian first reached out to me in March, after I appeared on the 80k hours podcast. He was pursuing corrigibility as part of his graduate thesis, and soon became a primary colleague in discussing CAST. While he is new to the field and has no substantial public work on the topic, what he's shared with me directly has impressed me, and I'm excited to support him working on corrigibility full time.
Jared Glover
The Geometry of Corrigibility — $20k
Jared reached out to me with an initial proposal, building off his earlier work, to test how the use of corrigibility steering vectors would change the behavior of models, and what impact that might have on things like honesty. I was initially a bit skeptical; much of what he was describing, from my perspective, looked more like generic interp and control work — the sort of stuff Anthropic has been doing. Still, I invited him to my research group and to chat with me more, and I was pleasantly surprised by his thoughtful, independent perspective and impressed by his technical understanding. And after talking about how to refine his proposal to hit corrigibility more directly and scientifically, I'm very excited to see what he finds, both in the near future and in years to come.
Alfaxad Eyembe
Corrigibility Games — $12.5k
Alfaxad has a very strong hacker/builder track record, given his age. (I don't know how old he is, but c'mon. Nobody my age is this cool.) He impressed me with his quick uptake on the subtleties of corrigibility when I provided feedback on his initial proposal, and his (amended) proposal to work on games that demonstrate aspects of corrigibility (as distinct from nearby concepts like pure obedience) impressed me with its novelty. Just as I wrote Red Heart to try to provide an accessible introduction to corrigibility and its nuances, I am excited by the prospect of one or more games that can do the same, in a fast and engaging way. With luck, Alfaxad's work will draw attention and build understanding in other young hackers who can help make our AI safer.
Florian Dietz
Dedicated Misalignment Flag — $20k
I enjoyed reading the Split Personality Training paper when it first came out, and am generally a fan of Florian, as an alignment researcher. His approaches have a good mixture of being concrete/prosaic/empirical and clever/interesting/theoretical, and he clearly has a lot of ideas. While I am primarily interested in the possibility of training a dedicated communication channel as a way to facilitate corrigibility on the architectural level, I'm also excited to simply have him thinking about the topic. As an independent researcher, he has both a good track record and a dependency on grants like this one. I hope that he gets a more substantial level of support from a big funder, so that he can contribute even more insight.
Ben Saudek
Training a Corrigible Model — $3k
Ben first reached out to me, like Ian Kahn, after my appearance on the 80k podcast, and was interested in studying and writing about corrigibility (from a purely theoretical direction) over the summer as a way to dip his toes into the field. His thoughts and essays over the last few months have improved my understanding of the subject, and I'm very excited to have him continue on that path. His application to train an LLM came as a little bit of a shock, and I'm hoping that it doesn't distract him from his more theoretical pursuits. But, as school starts up again, I selfishly want him to spend lots of time continuing to think about corrigibility, and I expect that funding him to play around with fine-tuning is a way to help make that happen. And who knows, maybe he'll beat Ian to the finish line on 1/12th the budget![2]
Xuanyi Wang and Shuo Li Liu
CAST Decision Theory[3] — $20k
Xuanyi is a recent philosophy graduate in Beijing with industry experience working on LLMs, and a particular interest in AI governance. Shuo is an economist at Princeton specializing in decision theory. While I'm unsure that they can make headway on a formalism, I was impressed by their enthusiasm and ambition. (Special thanks to Xuanyi for staying up past 2:00am to jump on a call with me with almost no advance notice.) Their first submission was pretty marginal, but after I nudged them to engage more with corrigibility as a concept, Shuo in particular seems to have gotten nerd-sniped by the topic and their revised proposal is very promising, to my eyes. I'm also hopeful that this will be the start of a pivot for one or the both of them towards working in AI alignment full-time.
Nick R.
Benevolent Override Pilot Study — $1.5k
Nick has been following AI safety from afar since 2023, but is a relative novice in terms of doing work in the field. When he reached out to me in August, I was impressed by his careful and transparent communication, such as the way he flagged possible reasons not to fund him, marked the parts of his application that were the product of conversing with LLMs, and sought my feedback. He also asked for a remarkably small grant. This pilot project is, from my perspective, mostly[4] a speculative bet on him upskilling and pivoting his career towards additional safety work.
Rubi Hudson
Principled Agents — $20k
I'm a bit embarrassed to note that I missed Rubi's corrigibility paper when it came out last year. It's very impressive in many respects, and I want a hundred more like it! (I also have several quibbles with Rubi's take on corrigibility, and am currently debating it with him on LessWrong.) This $20k is something of a vote-of-confidence to give them a little more runway and to nudge other, bigger funders to send more substantial support to their newborn research org. They're doing actual alignment research! What a concept! Send funding!
Honorable Mentions
Other applicants who I decided not to fund this round, but who were promising enough for me to strongly consider, and who I hope other funders consider[5] backing:
I'm sure many other qualified people applied. My sincerest apologies to anyone who ought to have been funded, but whom I overlooked. Again, I encourage everyone who is interested in being funded for Round 2 (deadline: October 31st), or for AI alignment funding more generally, to apply on Lightcone Commons.
Technically I only advise Lightcone Infrastructure on how to disburse the funds, and they have the final call. But in practice I'm the fund's sole manager.
More realistically, I hope they, and the other researchers looking to experiment with training for corrigibility, share what they have, work together when possible, and help each other advance the broader field.
I think they mean something more like a formal model of how a corrigible agent makes decisions, rather than the kind of philosophical decision theory usually focused on in LW circles. On a call with Shuo, I got the impression that he's been familiar with FDT/LDT (as well as CDT & EDT) for many years, but usually thinks about the subject from more of an economics perspective, which (combined with linguistic barriers) might explain the use of "decision theory" as a central term in a project that I don't really think of as being centrally about decision theory.
"Benevolent override" scenarios don't strike me as the most interesting topic, but I still think the object-level is interesting enough to warrant a small pilot project. Even just getting more thoughts from junior researchers seriously engaging with corrigibility would be valuable for identifying primary confusions/obstacles.
Funders, please do your own due diligence. This is not a blanket endorsement.