Linkpost for my Substack piece, adapted a reasonable amount for EA specifically.
Almost all EA projects would benefit from better monitoring, but AI governance most of all, in my experience.
Most donors in EA never find out whether their grants worked. They model the impact before the money goes out, but don’t check if the models were accurate. Evaluating impact is hard, but monitoring grant progress is simple: agree indicators in writing before you send the money, ask the grantee to put a probability on each, and score them once a year.
Charities usually send grant reports anyway. Unless you ask for something specific, they rarely contain the most useful information.
A recent report to a client touted the project’s success: it had made 20 policy recommendations. When pressed, the grantee responded that only 20% had been implemented, even partially.
A 20% implementation rate may or may not be a good outcome. Either way, it wasn’t the outcome they reported, or what they were asked to report by the donor.
Ironically, this is pervasive in ‘evidence-based giving’, and particularly in AI governance work.
How much do donors know about what their grants achieved?
When it comes to individual donors, often nothing.
‘Evidence-based giving’ typically applies before the grant. We are very good at modelling what a grant might achieve. We’re surprisingly bad at checking what it did.
This varies by sector. Frontline global health interventions routinely monitor outputs - clinic visits, vaccines distributed - and typically push further to monitor outcomes and impact, because the whole sector expects it.
In harder-to-measure areas, with no strong standard for self-monitoring, it’s the Wild West. I see policy, advocacy and governance interventions struggle the most.
Berlin-based think tank Future Matters recently found that 13 out of 22 major earning-to-give donors couldn’t tell whether AI governance work achieves anything, including four out of six who work at frontier AI labs.
If even the people closest to the technology aren’t sure what their grants are achieving, what hope do the rest of us have?
Future Matters blames legibility. I’d go further: illegibility is a choice donors make, by not asking the right questions at the right time.
The effective giving space currently inverts best practice: monitoring is often most rigorous where uncertainty is lowest (bednets), and often absent where it’s highest (policy and advocacy).
Why does monitoring and evaluation of donations matter?
Monitoring, evaluation and learning (MEL) is how we understand what is working, to improve both projects and funders.
A useful MEL plan starts with its purpose, and what questions it seeks to answer.
As a project, you might want to know whether your services are effective and reach the right people; how to improve them; or what you can use to fundraise.
As a funder, you have a stake in whether your individual grantees are succeeding. However, there are many other benefits to you: improving your funding decisions; managing risk across your portfolio; and making a meaningful difference to the problem (even more than a single project, as you hold a variety of bets). You might also want to share lessons with the wider ecosystem or even build the field.
Defining which of these questions matters focuses our activities and makes our MEL plans useful.
MEL is the only way to answer these questions systematically. Without it, we risk wasting funding and time. In the absence of evidence, projects can go astray and donors can withhold grants. Consequently, we willstruggle to solve the life-or-death problems we are working on.
What about evaluation specifically?
Monitoring is the ongoing tracking of execution at the programme level. Evaluation is the wider assessment of what was achieved.
Although this post focuses on monitoring, because evaluation is much more resource-intensive and difficult, effective giving is also surprisingly weak at evaluating its grants. Even GiveWell, the field-leader in terms of rigour and transparency, has conducted very few retrospective evaluations, only starting this process in 2025.
Interestingly, when they did, they found that some grants had performed much better than modelled and some had performed worse.
I think this is notable in a field that traditionally looks for RCTs from its grantees. We very rarely apply the same rigourous measurement to our own grantmaking.
If we do more post hoc evaluation in future, it is likely that we can check how well calibrated our models were and improve them over time. I won't dwell on this in this post but would like to see more of this work, especially outside global health, which tends to be stronger on all aspects of MEL than other cause areas.
Why don’t donors ask for more reporting from their grantees?
One motivation is noble: not overburdening the project. This has been a positive shift away from intensive, performative reporting that almost nobody read anyway.
This change has been driven by two very different camps: cost-effective giving at one end and trust-based philanthropy at the other.
The problem is that each risks having the right diagnosis and the wrong cure. The fact that traditional reporting istoo burdensome and not useful doesn’t mean that you shouldn’t ask anything at all.
Any strong organisation should actively want to know whether its work is effective. If it is happy to run purely on anecdotes and intuition, that’s a red flag.
Is policy work too hard to measure?
A second reason to forgo monitoring and evaluation is the belief that some interventions are too hard to measure.
This is half right. Evaluating the impact of policy, advocacy and governance work is extremely difficult. You can’t usually run RCTs, so counterfactuals are tricky. Results can take years to arrive. Several groups will usually work on the same issue, without a clean way to assign the credit. As Future Matters points out, publicly claiming attribution can burn your most important relationships.
This is particularly acute in AI safety and governance, where it isn’t even obvious what makes a good regulation. If reasonable people disagree about whether slowing down AI development is net positive, no monitoring and evaluation plan will solve that. It’s just the risk you take by funding this work.
There are nonetheless proxy indicators of effectiveness - certainly enough to distinguish organisations from one another in the short-term. For example, bills sometimes use language drawn directly from a project’s publications; lawmakers sometimes acknowledge conversations with particular groups; and it can be surprisingly effective to ring congressional staff and ask how influential an advocacy project was.
In addition, monitoring policy work is really not that complex, even when impact measurement is hard. Any MEL professional would find it routine: ask an organisation what it plans to do and how likely it is to succeed, and then track whether it did.
The best projects I have assessed could point me clearly to language in a bill that passed into law, and show how it was derived or quoted from their work, which is a pretty serious achievement. These organisations should be rewarded with more fundingand support.
How can I monitor my policy and advocacy donations?
A strong monitoring plan includes a handful of indicators, with the project assigning each a probability of success. These can be scored once a year, with any failures left in. It should take the project around half a day to complete and track things it wants to know anyway.
My advisory, Ultra Philanthropy, has an in-house MEL specialist with 12 years’ experience advising development programmes. We have worked successfully with some admirably responsive, transparent organisations to produce these sorts of plans, even in thorny areas of advocacy.
For example, we worked with a climate advocacy group in 2025 to define the following:
Defend climate provisions in reconciliation and agency reorganisation
Protect Loan Programs Office credit subsidy at Department of Energy and defend other DOE funding - 25% [chance of success]
Protect advanced geothermal and nuclear access to tax credit through at least 2029 with viable Foreign Entities of Concern restrictions - 60%
Achieve trade legislative objectives
PROVE IT Act on emissions intensity to pass the House (20%) and Senate (30%) in 119th Congress
Pass reworked carbon border adjustment mechanism in the House (20%) and Senate (25%) during 120th Congress
Success in Energy Acts
Pass an Energy Act with meaningful regulatory reform in 2027 (House, 15%; Senate, 30%) or 2029 (House, 20%; Senate, 30%)
Success in permitting reform
Pass bipartisan bill in 119th Congress (65% of slim or full package passing)
If 4a fails, groundwork laid to pass slim or full package in 120th Congress (75%, provided Democrats control the House)
It would have been very easy for the project to focus on controllable activities (‘we’ll have 50 meetings with lawmakers’), or to claim that their work was powerful magic that defied tracking (a depressingly common claim, which I have universally found to be false).
Instead, its plan gave our client (and the organisation itself) a clear, measurable monitoring framework. Even though the impact of these legislative achievements remains uncertain (laws still have to be enforced, might have unintended consequences, take a long time to have an effect, etc.), it’s easy to see whether or not they moved the ball down the field.
Unfortunately, this clarity is extremely rare in the advocacy and policy work that I encounter, especially in AI safety and governance.
Why ask grantees to put a probability on each target?
An aspect of our monitoring that goes beyond the norm is assigning probabilities to each output.
One benefit is that this exposes overconfident fundraisers. Often, fundraising teams send donors several impressive milestones that a grant could enable. Ask their programme team to put a success percentage against each, though, and it’s clear how likely you are to see them.
It also prevents sandbagging. If a project suggests indicators that it has a 95% chance of achieving, we look more closely at its ambition and risk-tolerance, and check whether those outcomes are actually meaningful.
Sometimes success is both highly likely and meaningful - ‘we’re 85% confident of reaching 10,000 people with oral rehydration solution this year’. But this can also be playing it safe, or not focusing on the most important changes.
Wielded together, these opposing concerns create something exciting: a portfolio that blends solid gains with stretching, ambitious goals, which tell the donor and the project how close they are to getting the hard thing done.
Donors then need to focus on calibration, not merely success rate. An ambitious miss is usually more useful than a 100% hit rate. It’s essential if we are going to embrace more risk in our giving.
What should I agree with a grantee before I send the money?
For most large donors (>$1m/year), it makes sense to engage a MEL consultant or an advisory with in-house expertise to do this work, since it usually involves negotiating with grantees, and a reasonable amount of data collection and analysis across a portfolio.
For donors wanting to implement this themselves, however, there are some things to watch out for:
Start by defining your purpose and related questions. If you know what you want to learn and what you want to do with the information, you increase your chances of asking the right questions, focusing on what is useful over what you might find interesting.
Make sure to agree targets in writing before sending the money. It is much harder (and sometimes unfair) to negotiate the goals of a grant after sending the funds.
Don’t accept activities and outputs dressed as outcomes. Activities are things a project does (write policy briefs, convene conferences). Outputs are things produced by the activities (number of policy briefs, number of conferences). Neither is meaningful unless it achieves an outcome, i.e. a positive change (such a bill being passed into law). It is extremely common for projects to substitute activities and outputs for outcomes - don’t let them.
Agree a suitable range of targets. One way to look great is to list many possible outcomes and then report only the one or two that landed. A written framework makes projects focus on things they can actually affect, and acknowledge the things they tried that didn’t work out. Hits-based giving is fine - you might only need one success in eight - as long as both sides knew that going in.
Include confidence levels. This prevents overclaiming and sandbagging, and lets you gauge how ambitious the project is, and how much risk you have across your whole portfolio.
Take refusal as a signal. If the project won’t agree to ambitious and clearly defined indicators (or, even worse, tries to explain why this doesn’t apply to its work because it’s unique), that’s a huge red flag.
Make decisions based on the results. Look at the monitoring data and update your beliefs when making future decisions. It’s perfectly fine to continue funding difficult work, even if it missed every objective - presumably most successful vaccines came after tens or even hundreds of failed candidates. What’s harmful is to burden projects with reporting and then not to use it to improve your grantmaking.
I’m expecting an AI windfall. What should I do?
Establishing a consistent process to monitor your grants and learn from them is essential if you are about to become a significant philanthropist.
The Funding Anthropalypse will create many large donors who have little or no experience in grantmaking or MEL. This makes feedback loops a high priority, so that you can improve your grants, risk-tolerance and general understanding over time.
You may also give a significant amount to AI safety and governance work, which is much more uncertain and harder to measure than many cause areas, with greater risk of accidental harm. This makes effective monitoring even more important.
Use the pointers above to shape your monitoring strategy, and talk to a professional if you can. You can also defer to evaluators, advisories or funds you trust to do this work for you.
Illegibility is achoice. Make a different one.
Jack Lewars is the founder of Ultra Philanthropy, an independent advisory that helps major donors give for maximum impact, and is the fund manager of its Mid-Stage Global Health Fund. He advises donors giving up to nine figures a year, and is Chair of Trustees at High Impact Athletes. Talk to him about your giving.
Thanks to Olivia Kaye and Samuel Verbi for feedback on the draft. I used Claude to help structure my thoughts and to suggest improvements and flag gaps, as well as for proofreading; all views, primary drafting and final edits are mine.
Linkpost for my Substack piece, adapted a reasonable amount for EA specifically.
Almost all EA projects would benefit from better monitoring, but AI governance most of all, in my experience.
Most donors in EA never find out whether their grants worked. They model the impact before the money goes out, but don’t check if the models were accurate. Evaluating impact is hard, but monitoring grant progress is simple: agree indicators in writing before you send the money, ask the grantee to put a probability on each, and score them once a year.
Charities usually send grant reports anyway. Unless you ask for something specific, they rarely contain the most useful information.
A recent report to a client touted the project’s success: it had made 20 policy recommendations. When pressed, the grantee responded that only 20% had been implemented, even partially.
A 20% implementation rate may or may not be a good outcome. Either way, it wasn’t the outcome they reported, or what they were asked to report by the donor.
Ironically, this is pervasive in ‘evidence-based giving’, and particularly in AI governance work.
How much do donors know about what their grants achieved?
When it comes to individual donors, often nothing.
‘Evidence-based giving’ typically applies before the grant. We are very good at modelling what a grant might achieve. We’re surprisingly bad at checking what it did.
This varies by sector. Frontline global health interventions routinely monitor outputs - clinic visits, vaccines distributed - and typically push further to monitor outcomes and impact, because the whole sector expects it.
In harder-to-measure areas, with no strong standard for self-monitoring, it’s the Wild West. I see policy, advocacy and governance interventions struggle the most.
Berlin-based think tank Future Matters recently found that 13 out of 22 major earning-to-give donors couldn’t tell whether AI governance work achieves anything, including four out of six who work at frontier AI labs.
If even the people closest to the technology aren’t sure what their grants are achieving, what hope do the rest of us have?
Future Matters blames legibility. I’d go further: illegibility is a choice donors make, by not asking the right questions at the right time.
The effective giving space currently inverts best practice: monitoring is often most rigorous where uncertainty is lowest (bednets), and often absent where it’s highest (policy and advocacy).
Why does monitoring and evaluation of donations matter?
Monitoring, evaluation and learning (MEL) is how we understand what is working, to improve both projects and funders.
A useful MEL plan starts with its purpose, and what questions it seeks to answer.
As a project, you might want to know whether your services are effective and reach the right people; how to improve them; or what you can use to fundraise.
As a funder, you have a stake in whether your individual grantees are succeeding. However, there are many other benefits to you: improving your funding decisions; managing risk across your portfolio; and making a meaningful difference to the problem (even more than a single project, as you hold a variety of bets). You might also want to share lessons with the wider ecosystem or even build the field.
Defining which of these questions matters focuses our activities and makes our MEL plans useful.
MEL is the only way to answer these questions systematically. Without it, we risk wasting funding and time. In the absence of evidence, projects can go astray and donors can withhold grants. Consequently, we will struggle to solve the life-or-death problems we are working on.
What about evaluation specifically?
Monitoring is the ongoing tracking of execution at the programme level. Evaluation is the wider assessment of what was achieved.
Although this post focuses on monitoring, because evaluation is much more resource-intensive and difficult, effective giving is also surprisingly weak at evaluating its grants. Even GiveWell, the field-leader in terms of rigour and transparency, has conducted very few retrospective evaluations, only starting this process in 2025.
Interestingly, when they did, they found that some grants had performed much better than modelled and some had performed worse.
I think this is notable in a field that traditionally looks for RCTs from its grantees. We very rarely apply the same rigourous measurement to our own grantmaking.
If we do more post hoc evaluation in future, it is likely that we can check how well calibrated our models were and improve them over time. I won't dwell on this in this post but would like to see more of this work, especially outside global health, which tends to be stronger on all aspects of MEL than other cause areas.
Why don’t donors ask for more reporting from their grantees?
One motivation is noble: not overburdening the project. This has been a positive shift away from intensive, performative reporting that almost nobody read anyway.
This change has been driven by two very different camps: cost-effective giving at one end and trust-based philanthropy at the other.
The problem is that each risks having the right diagnosis and the wrong cure. The fact that traditional reporting is too burdensome and not useful doesn’t mean that you shouldn’t ask anything at all.
Any strong organisation should actively want to know whether its work is effective. If it is happy to run purely on anecdotes and intuition, that’s a red flag.
Is policy work too hard to measure?
A second reason to forgo monitoring and evaluation is the belief that some interventions are too hard to measure.
This is half right. Evaluating the impact of policy, advocacy and governance work is extremely difficult. You can’t usually run RCTs, so counterfactuals are tricky. Results can take years to arrive. Several groups will usually work on the same issue, without a clean way to assign the credit. As Future Matters points out, publicly claiming attribution can burn your most important relationships.
This is particularly acute in AI safety and governance, where it isn’t even obvious what makes a good regulation. If reasonable people disagree about whether slowing down AI development is net positive, no monitoring and evaluation plan will solve that. It’s just the risk you take by funding this work.
There are nonetheless proxy indicators of effectiveness - certainly enough to distinguish organisations from one another in the short-term. For example, bills sometimes use language drawn directly from a project’s publications; lawmakers sometimes acknowledge conversations with particular groups; and it can be surprisingly effective to ring congressional staff and ask how influential an advocacy project was.
In addition, monitoring policy work is really not that complex, even when impact measurement is hard. Any MEL professional would find it routine: ask an organisation what it plans to do and how likely it is to succeed, and then track whether it did.
The best projects I have assessed could point me clearly to language in a bill that passed into law, and show how it was derived or quoted from their work, which is a pretty serious achievement. These organisations should be rewarded with more funding and support.
How can I monitor my policy and advocacy donations?
A strong monitoring plan includes a handful of indicators, with the project assigning each a probability of success. These can be scored once a year, with any failures left in. It should take the project around half a day to complete and track things it wants to know anyway.
My advisory, Ultra Philanthropy, has an in-house MEL specialist with 12 years’ experience advising development programmes. We have worked successfully with some admirably responsive, transparent organisations to produce these sorts of plans, even in thorny areas of advocacy.
For example, we worked with a climate advocacy group in 2025 to define the following:
It would have been very easy for the project to focus on controllable activities (‘we’ll have 50 meetings with lawmakers’), or to claim that their work was powerful magic that defied tracking (a depressingly common claim, which I have universally found to be false).
Instead, its plan gave our client (and the organisation itself) a clear, measurable monitoring framework. Even though the impact of these legislative achievements remains uncertain (laws still have to be enforced, might have unintended consequences, take a long time to have an effect, etc.), it’s easy to see whether or not they moved the ball down the field.
Unfortunately, this clarity is extremely rare in the advocacy and policy work that I encounter, especially in AI safety and governance.
Why ask grantees to put a probability on each target?
An aspect of our monitoring that goes beyond the norm is assigning probabilities to each output.
One benefit is that this exposes overconfident fundraisers. Often, fundraising teams send donors several impressive milestones that a grant could enable. Ask their programme team to put a success percentage against each, though, and it’s clear how likely you are to see them.
It also prevents sandbagging. If a project suggests indicators that it has a 95% chance of achieving, we look more closely at its ambition and risk-tolerance, and check whether those outcomes are actually meaningful.
Sometimes success is both highly likely and meaningful - ‘we’re 85% confident of reaching 10,000 people with oral rehydration solution this year’. But this can also be playing it safe, or not focusing on the most important changes.
Wielded together, these opposing concerns create something exciting: a portfolio that blends solid gains with stretching, ambitious goals, which tell the donor and the project how close they are to getting the hard thing done.
Donors then need to focus on calibration, not merely success rate. An ambitious miss is usually more useful than a 100% hit rate. It’s essential if we are going to embrace more risk in our giving.
What should I agree with a grantee before I send the money?
For most large donors (>$1m/year), it makes sense to engage a MEL consultant or an advisory with in-house expertise to do this work, since it usually involves negotiating with grantees, and a reasonable amount of data collection and analysis across a portfolio.
For donors wanting to implement this themselves, however, there are some things to watch out for:
I’m expecting an AI windfall. What should I do?
Establishing a consistent process to monitor your grants and learn from them is essential if you are about to become a significant philanthropist.
The Funding Anthropalypse will create many large donors who have little or no experience in grantmaking or MEL. This makes feedback loops a high priority, so that you can improve your grants, risk-tolerance and general understanding over time.
You may also give a significant amount to AI safety and governance work, which is much more uncertain and harder to measure than many cause areas, with greater risk of accidental harm. This makes effective monitoring even more important.
Use the pointers above to shape your monitoring strategy, and talk to a professional if you can. You can also defer to evaluators, advisories or funds you trust to do this work for you.
Illegibility is a choice. Make a different one.
Jack Lewars is the founder of Ultra Philanthropy, an independent advisory that helps major donors give for maximum impact, and is the fund manager of its Mid-Stage Global Health Fund. He advises donors giving up to nine figures a year, and is Chair of Trustees at High Impact Athletes. Talk to him about your giving.
Thanks to Olivia Kaye and Samuel Verbi for feedback on the draft. I used Claude to help structure my thoughts and to suggest improvements and flag gaps, as well as for proofreading; all views, primary drafting and final edits are mine.