In Q2 2026, we ran the first MATS x Coefficient Giving (CG) Pitch Week: a four-week program designed to take fellowship researchers from an early idea to a funding decision. The program was a joint effort between MATS, CG, Constellation, and Catalyze Impact.
38 expressions of interest; all were invited to a two-week Exploration Sprint. 17 teams applied and 15 teams pitched.
$4.3M was awarded across six grants (47% of teams that pitched) with five org-founding grants and two independent research grants.
Funding outcomes were not strongly correlated with the founder characteristics we could observe. What distinguished successful pitches was the maturity and evaluability of the idea: a named first user, a concrete first deliverable, a bounded scope, evidence of technical feasibility, and a legible reason the work required a new organization.
Why we ran this, and why we're writing it up
We believe that supporting founders could be unusually valuable. AI safety orgs founded by MATS alumni, such as Apollo and Timaeus, likely account for an outsized share of MATS’s impact. New organizations can also act as capacity multipliers, especially now that funding is outstripping the supply of organizations ready to absorb it, making founder support one of our highest-leverage activities.
At the same time, AI safety needs more people who can work independently, set strategy, manage teams, and build organizations. It remains an open question whether the field is short of such people or short of the infrastructure to help technically strong researchers make the transition. Talent programs such as MATS primarily select and train for research ability, not for founder-specific skills and experience. (Note: the new MATS Founding & Field-Building track, the GovAI Entrepreneur-in-Residence, and the Generator Residency are exceptions.)
Pitch Week was our first attempt at scaffolding part of that transition. We did not expect to produce experienced leaders in a few weeks, just to give technically strong people time and structure to develop an organization idea and consider co-founders, organization structure, theory of impact, and fundraising strategies.
We're writing it up for two reasons. First, Pitch Week may become a recurring MATS offering and we want an account of the first run to be compared against future versions. Second, other AI safety research fellowship programs might see similar interest in founder support from their fellows, but have little published process or outcome data. We hope this report helps other AI safety programs support founders; please reach out if you are interested!
What we ran
The target group for the program comprised first-time technical founders coming out of fellowships: current MATS fellows and alumni, plus current Astra and Anthropic Fellows Program (AFP) fellows. There were four stages:
Expression of Interest (released 30 Apr, deadline 10 May). Thirty-eight fellows signed up, and all of them were admitted to the Exploration Sprint. We did not filter at this stage, as the marginal cost of including each participant was low, and we used the Pitch Week application as the main check point. We thought that keeping the sprint open also gave participants more opportunities to find potential co-founders.
Exploration Sprint (12–25 May). We opened with a kickoff session, then ran recurring office hours alongside coaching calls from Catalyze Impact and light-touch support from two MATS Research Managers. There were also opportunities for conversations with prospective users and advisors, and introductions to external collaborators.
Pitch Week application (deadline 25 May). Twenty-eight of the 38 sprint participants submitted memo-style applications across 17 teams; the other 10 did not apply. We used the application as a checkpoint rather than a competitive gate. Fifteen teams went on to participate in Pitch Week, either at the MATS London office or remotely. One team dropped out, while another was sufficiently advanced to move into a seed-stage conversation.
Between the sprint and Pitch Week, we also held two founder Q&As (27 and 28 May): Ronak Mehta, Co-Founder of Coordinal Research, gave a postmortem of his org, and Rajashree Agrawal of Theorem shared her co-founding story. These sessions covered something we had otherwise underweighted: what running an organization involves, as distinct from doing research. We also encouraged participants to attend EA Global London, which offered additional opportunities to meet potential advisors and collaborators.
Pitch Week (1–5 Jun). From Monday through Wednesday, teams prepared and stress-tested their plans through a theory-of-change workshop, red-teaming with invited AI safety researchers, and a round ofIntelligence Rising. Pitches took place on Thursday and Friday: each team had a separate 25-minute 1:1 with CG grantmakers Jake Mendel (technical AI safety) and Catherine Brewer (governance and cybersecurity). These were run as research conversations rather than presentations. BlueDot Impact was running its own Incubation Week in the MATS London office at the same time, and the two programs closed with a joint AI Safety Founders Dinner organized by BlueDot.
What kinds of orgs were pitched
The 15 teams covered a broad range of problem areas:
Evaluations and auditing (5 teams). This is the largest cluster, spanning third-party auditing and assurance, evaluation infrastructure for agents, and domain-specific evaluation proposals.
Research infrastructure (4 teams). This includes benchmarks, datasets, environments, and tooling intended to make safety research faster or more trustworthy.
Security (2 teams). This includes defending the training pipeline, and verifying properties of training runs.
Policy and field-building (2 teams). Helping the public and policy makers understand emerging AI risks.
Biosecurity (1 team). Safeguards against hazardous capabilities in frontier biological models.
The problem area did not predict the outcome. Evaluations, infrastructure, and public-facing proposals each appeared among both funded and non-funded teams. What differentiated funded and non-funded projects seems to be the proposal shape, as we discuss below.
What the data shows
1. The founder characteristics we observed did not meaningfully distinguish outcomes. Reference strength and responses to our application question regarding leadership experience were similar across funded and non-funded teams. The same was true for prior founding experience and demonstrated ability to ship products/research: several non-funded teams had strong technical credentials and had previously deployed products. This is not to say founder quality does not matter. Rather, within this selected group, the characteristics we measured seemed to carry little information about who would be funded.
2. Funded org proposals were specific about what they would build first, and for whom. The four funded org proposals shared several features. Each proposal identified a specific first user or audience rather than "the field of AI safety, broadly". Each proposal had a first deliverable that could plausibly exist by the end of the grant period, and that deliverable was an artifact someone could use rather than merely "more research". Finally, each proposal made a clear case for why this should become an organization, rather than one person's research agenda, with a plan that detailed ongoing delivery, adoption, and building a team.
3. Non-funded proposals tended to lack specificity about what they would build first, for whom, and how it would lead to impact, rather than founder quality. In many cases, the target user, initial deliverable, or near-term execution plan was not well defined, or the proposal’s key assumptions had not been sufficiently validated. The pitch notes returned to three questions: Who exactly is the user, and why would they adopt the work? Will the central technical assumption hold? Why does the work require a new organization rather than becoming a project at an existing organization?
4. More advisory conversations did not predict funding. We recommended that fellows take further advisory calls based on the assumption that more calls increased the odds of finding the one insight that mattered. Participants followed that advice: non-funded proposals reported a median of 30 advisory conversations compared with 12 among funded ones. This could mean that, generally, participants who received critical feedback from advisory calls were unable to improve their pitches with further calls.
Much of the feedback from these calls simply confirmed that advisors considered the cause area important. Because this kind of vague positive signal is common across many cause areas, it carries little useful information for grantmakers to distinguish promising pitches. Rather than generic endorsements of the cause area, positive endorsements of the specific founders seemed more important to getting funded. Founders were more likely to receive funding when an advisor thought that a particular team was well-suited to the work, committed to advising the team over time, or positively updated their view of the proposal over the conversation.
We might have overoptimized for quantity of advisory calls. Separately, feedback from several participants suggests that the startup-oriented interview framing we supplied (the "Mom Test") might not translate well to research organizations.
Operational lessons
Lock and communicate the design before applications open. Compressed timelines and late communication were the dominant complaints among participants. They did not learn the program structure until the sprint began, and this made their planning around other commitments very difficult. When the opportunity to collaborate with CG arose, we moved quickly to make Pitch Week happen. Given those constraints, we were still encouraged by what the program was able to accomplish, but with more lead time we would finalize and communicate the design before opening applications.
Redesign the guidance for advisory conversation. Rather than emphasizing a high volume of startup-style customer interviews, encourage participants to build ongoing advisor relationships that can provide more specific insight into their fit for the proposed work.
Several parts of the Pitch week worked well. The deadline and prospect of a funding decision within two weeks gave participants momentum to make progress, and several participants reported that the sprint changed how they thought about their work, including some who came in skeptical of founding. Office hours appeared to be quite helpful and received the most consistent praise from participants. Participants also valued the kickoff and the memo-style application as a device for committing ideas to paper.
Open questions for future iterations
How long the sprint should be. Two weeks felt too short. This was the most consistent piece of feedback we received, even from participants who joined with well-developed ideas. A longer sprint or tighter weekly milestones might give more time and structure. We also found that screening participants between the Exploration Sprint and Pitch Week created considerable uncertainty. For future iterations, we might instead screen participants based on idea readiness before exploration, with those selected participating in both Exploration Sprint and Pitch Week.
How much structure to impose. Participants had different preferences. Some wanted more programming; others found the same programming a distraction from the conversations that mattered. A more flexible program design may be the best way to accommodate these different preferences.
Whether the conversion rate is good. We surfaced five fundable org-founding teams and two fundable research agendas from 38 interested researchers in one fellowship community in one quarter. Whether that reads as encouraging depends on assumptions about how many new AI safety organizations the field can absorb and the funding pipeline can support. We would be interested to see how these numbers compare with those of similar programs. As a rough benchmark based on statistics one year ago: 10% of MATS alumni who completed the program before 2025 had co-founded an AI safety organization or team by Aug 2025.
Whether the funded organizations prove to be the right calls. A funding decision is only an early signal. We do not yet know whether the teams will stay together, whether their original proposals will survive contact with the work, or whether they will secure follow-up funding. Tracking these outcomes will be necessary to evaluate the program properly.
If you run something adjacent (e.g., an incubator, a fellowship, or a grantmaking program) we'd like to compare notes, particularly on the parts that didn't work. And if you think we're misreading our own data, we'd like to hear that too.
Limitations
The cohort was small (28 fellows in 15/17 teams).
Funding decisions may reflect considerations that are not fully captured in our program data or retrospective notes.
Most measures were imperfect proxies. For example, the number of advisory conversations says little about their quality.
Our account is shaped by the information available to the program team and by our interest in improving future versions.
Accordingly, we treat these results as directional. The observations are useful for program design, but they should not be generalized to founder selection or AI safety incubation without further evidence.
Acknowledgments
This report was produced by Machine Alignment, Transparency, and Security (MATS) Research. Jinghua Ou and Patryk Wielopolski were the primary authors of this report with the assistance of Claude 5; Ryan Kidd and John Teichman helped with editing. Patryk Wielopolski and Jinghua Ou ran the Pitch Week with support from other MATS team members.
Thanks to Jake Mendel and Catherine Brewer at Coefficient Giving for the grantmaking and committing time to meet with each team, evaluate each proposal in detail and provide valuable feedback to them; to Clarissa Lam at Constellation, who ran participant logistics, and to Olivia Benoit, whose in-person day contributed a non-technical org-building perspective we were missing, and to both for bringing Astra and AFP fellows into the program; to Hugo Walrand at Catalyze Impact for providing 1-1 coaching for participants; to BlueDot Impact for the founders dinner; and to the participants who spent four weeks seriously considering a change of direction, including those we could not fund.
TL;DR
In Q2 2026, we ran the first MATS x Coefficient Giving (CG) Pitch Week: a four-week program designed to take fellowship researchers from an early idea to a funding decision. The program was a joint effort between MATS, CG, Constellation, and Catalyze Impact.
Why we ran this, and why we're writing it up
We believe that supporting founders could be unusually valuable. AI safety orgs founded by MATS alumni, such as Apollo and Timaeus, likely account for an outsized share of MATS’s impact. New organizations can also act as capacity multipliers, especially now that funding is outstripping the supply of organizations ready to absorb it, making founder support one of our highest-leverage activities.
At the same time, AI safety needs more people who can work independently, set strategy, manage teams, and build organizations. It remains an open question whether the field is short of such people or short of the infrastructure to help technically strong researchers make the transition. Talent programs such as MATS primarily select and train for research ability, not for founder-specific skills and experience. (Note: the new MATS Founding & Field-Building track, the GovAI Entrepreneur-in-Residence, and the Generator Residency are exceptions.)
Pitch Week was our first attempt at scaffolding part of that transition. We did not expect to produce experienced leaders in a few weeks, just to give technically strong people time and structure to develop an organization idea and consider co-founders, organization structure, theory of impact, and fundraising strategies.
We're writing it up for two reasons. First, Pitch Week may become a recurring MATS offering and we want an account of the first run to be compared against future versions. Second, other AI safety research fellowship programs might see similar interest in founder support from their fellows, but have little published process or outcome data. We hope this report helps other AI safety programs support founders; please reach out if you are interested!
What we ran
The target group for the program comprised first-time technical founders coming out of fellowships: current MATS fellows and alumni, plus current Astra and Anthropic Fellows Program (AFP) fellows. There were four stages:
Expression of Interest (released 30 Apr, deadline 10 May). Thirty-eight fellows signed up, and all of them were admitted to the Exploration Sprint. We did not filter at this stage, as the marginal cost of including each participant was low, and we used the Pitch Week application as the main check point. We thought that keeping the sprint open also gave participants more opportunities to find potential co-founders.
Exploration Sprint (12–25 May). We opened with a kickoff session, then ran recurring office hours alongside coaching calls from Catalyze Impact and light-touch support from two MATS Research Managers. There were also opportunities for conversations with prospective users and advisors, and introductions to external collaborators.
Pitch Week application (deadline 25 May). Twenty-eight of the 38 sprint participants submitted memo-style applications across 17 teams; the other 10 did not apply. We used the application as a checkpoint rather than a competitive gate. Fifteen teams went on to participate in Pitch Week, either at the MATS London office or remotely. One team dropped out, while another was sufficiently advanced to move into a seed-stage conversation.
Between the sprint and Pitch Week, we also held two founder Q&As (27 and 28 May): Ronak Mehta, Co-Founder of Coordinal Research, gave a postmortem of his org, and Rajashree Agrawal of Theorem shared her co-founding story. These sessions covered something we had otherwise underweighted: what running an organization involves, as distinct from doing research. We also encouraged participants to attend EA Global London, which offered additional opportunities to meet potential advisors and collaborators.
Pitch Week (1–5 Jun). From Monday through Wednesday, teams prepared and stress-tested their plans through a theory-of-change workshop, red-teaming with invited AI safety researchers, and a round of Intelligence Rising. Pitches took place on Thursday and Friday: each team had a separate 25-minute 1:1 with CG grantmakers Jake Mendel (technical AI safety) and Catherine Brewer (governance and cybersecurity). These were run as research conversations rather than presentations. BlueDot Impact was running its own Incubation Week in the MATS London office at the same time, and the two programs closed with a joint AI Safety Founders Dinner organized by BlueDot.
What kinds of orgs were pitched
The 15 teams covered a broad range of problem areas:
The problem area did not predict the outcome. Evaluations, infrastructure, and public-facing proposals each appeared among both funded and non-funded teams. What differentiated funded and non-funded projects seems to be the proposal shape, as we discuss below.
What the data shows
1. The founder characteristics we observed did not meaningfully distinguish outcomes. Reference strength and responses to our application question regarding leadership experience were similar across funded and non-funded teams. The same was true for prior founding experience and demonstrated ability to ship products/research: several non-funded teams had strong technical credentials and had previously deployed products. This is not to say founder quality does not matter. Rather, within this selected group, the characteristics we measured seemed to carry little information about who would be funded.
2. Funded org proposals were specific about what they would build first, and for whom. The four funded org proposals shared several features. Each proposal identified a specific first user or audience rather than "the field of AI safety, broadly". Each proposal had a first deliverable that could plausibly exist by the end of the grant period, and that deliverable was an artifact someone could use rather than merely "more research". Finally, each proposal made a clear case for why this should become an organization, rather than one person's research agenda, with a plan that detailed ongoing delivery, adoption, and building a team.
3. Non-funded proposals tended to lack specificity about what they would build first, for whom, and how it would lead to impact, rather than founder quality. In many cases, the target user, initial deliverable, or near-term execution plan was not well defined, or the proposal’s key assumptions had not been sufficiently validated. The pitch notes returned to three questions: Who exactly is the user, and why would they adopt the work? Will the central technical assumption hold? Why does the work require a new organization rather than becoming a project at an existing organization?
4. More advisory conversations did not predict funding. We recommended that fellows take further advisory calls based on the assumption that more calls increased the odds of finding the one insight that mattered. Participants followed that advice: non-funded proposals reported a median of 30 advisory conversations compared with 12 among funded ones. This could mean that, generally, participants who received critical feedback from advisory calls were unable to improve their pitches with further calls.
Much of the feedback from these calls simply confirmed that advisors considered the cause area important. Because this kind of vague positive signal is common across many cause areas, it carries little useful information for grantmakers to distinguish promising pitches. Rather than generic endorsements of the cause area, positive endorsements of the specific founders seemed more important to getting funded. Founders were more likely to receive funding when an advisor thought that a particular team was well-suited to the work, committed to advising the team over time, or positively updated their view of the proposal over the conversation.
We might have overoptimized for quantity of advisory calls. Separately, feedback from several participants suggests that the startup-oriented interview framing we supplied (the "Mom Test") might not translate well to research organizations.
Operational lessons
Several parts of the Pitch week worked well. The deadline and prospect of a funding decision within two weeks gave participants momentum to make progress, and several participants reported that the sprint changed how they thought about their work, including some who came in skeptical of founding. Office hours appeared to be quite helpful and received the most consistent praise from participants. Participants also valued the kickoff and the memo-style application as a device for committing ideas to paper.
Open questions for future iterations
How long the sprint should be. Two weeks felt too short. This was the most consistent piece of feedback we received, even from participants who joined with well-developed ideas. A longer sprint or tighter weekly milestones might give more time and structure. We also found that screening participants between the Exploration Sprint and Pitch Week created considerable uncertainty. For future iterations, we might instead screen participants based on idea readiness before exploration, with those selected participating in both Exploration Sprint and Pitch Week.
How much structure to impose. Participants had different preferences. Some wanted more programming; others found the same programming a distraction from the conversations that mattered. A more flexible program design may be the best way to accommodate these different preferences.
Whether the conversion rate is good. We surfaced five fundable org-founding teams and two fundable research agendas from 38 interested researchers in one fellowship community in one quarter. Whether that reads as encouraging depends on assumptions about how many new AI safety organizations the field can absorb and the funding pipeline can support. We would be interested to see how these numbers compare with those of similar programs. As a rough benchmark based on statistics one year ago: 10% of MATS alumni who completed the program before 2025 had co-founded an AI safety organization or team by Aug 2025.
Whether the funded organizations prove to be the right calls. A funding decision is only an early signal. We do not yet know whether the teams will stay together, whether their original proposals will survive contact with the work, or whether they will secure follow-up funding. Tracking these outcomes will be necessary to evaluate the program properly.
If you run something adjacent (e.g., an incubator, a fellowship, or a grantmaking program) we'd like to compare notes, particularly on the parts that didn't work. And if you think we're misreading our own data, we'd like to hear that too.
Limitations
Accordingly, we treat these results as directional. The observations are useful for program design, but they should not be generalized to founder selection or AI safety incubation without further evidence.
Acknowledgments
This report was produced by Machine Alignment, Transparency, and Security (MATS) Research. Jinghua Ou and Patryk Wielopolski were the primary authors of this report with the assistance of Claude 5; Ryan Kidd and John Teichman helped with editing. Patryk Wielopolski and Jinghua Ou ran the Pitch Week with support from other MATS team members.
Thanks to Jake Mendel and Catherine Brewer at Coefficient Giving for the grantmaking and committing time to meet with each team, evaluate each proposal in detail and provide valuable feedback to them; to Clarissa Lam at Constellation, who ran participant logistics, and to Olivia Benoit, whose in-person day contributed a non-technical org-building perspective we were missing, and to both for bringing Astra and AFP fellows into the program; to Hugo Walrand at Catalyze Impact for providing 1-1 coaching for participants; to BlueDot Impact for the founders dinner; and to the participants who spent four weeks seriously considering a change of direction, including those we could not fund.