In May 2025 I met Yo Shavit,
who was working on national security policy at OpenAI and was thinking
about how to prepare for a future in which models could seriously assist
attackers in creating pandemics. We had a call, and when I shared notes
with my team their main response was: "maybe start with not making
models that can do that?"
Which is, in many ways, fair: by continuing to push the frontier in
biological capabilities, OpenAI's actions were making things worse on
many of the problems SecureBio is trying to solve. But OpenAI stopping
wouldn't have resolved the problem: other firms were pushing quickly
too, and the economic incentives strongly favored rapid capability
advancement. Making the world more resilient to pandemics needed to be
a high priority regardless, especially in light of models' increasing ability to help
people with biology.
When I thought about what our initial conversations might turn into,
however, my primary concerns were whether that might (a) compromise
SecureBio's ability to independently assess and criticize OpenAI's work
or (b) make the world less safe via reducing model developers'
motivation to improve safeguards. I do think there's something to both
of these, which I get into more below, but it seemed well worth it to
talk with Yo about how we could work together.
I explained how we were building an early warning system to
flag outbreaks, especially engineered
ones that could otherwise spread
widely before detection. I described how metagenomic sequencing
lets you see what nucleic acid sequences (DNA and RNA) are present in
a sample, without choosing in advance which sequences to look for, and
how we were piloting this on wastewater and nasal swabs. Over the next
few months we discussed opportunities to accelerate our work, he
introduced us to OpenAI cofounder Wojciech Zaremba, and both Yo and
Wojciech left OpenAI PBC (the for-profit) for the OpenAI Foundation (OAIF). We
continued talking to them in these new roles, and these discussions,
plus a lot of due diligence, led to the $17.2M
OAIF grant which we announced today. I'm incredibly excited about
this grant, which will allow us to expand our monitoring system,
reduce our turnaround time substantially, [1] and generally reduce the
risk that something could spread widely before detection.
Which brings me back to the two concerns I mentioned above. On (a),
independence, SecureBio has two divisions, Detection and AI. This grant
funds Detection, while the AI side of SecureBio evaluates models from
many firms, including OpenAI. The grant does not give OpenAI or OAIF
any formal control over what anyone at SecureBio can say publicly about
any models. That was important to us, but it was also key to OAIF: Yo
was very clear that they did not want to influence SecureBio AI's work,
and wanted to review our policies to make sure they were sufficiently
robust.
That said, we should pay attention to incentives. While it looks to
me like OAIF is making its own decisions, the two organizations are
closely linked in a way that goes beyond the shared "OpenAI" name: the
OpenAI Foundation's endowment is a ~1/4 stake in OpenAI PBC, and they
have almost identical boards. Might the AI side of SecureBio pull
punches in criticizing the PBC to increase the chances that OAIF gives
Detection more money in the future?
SecureBio has systemic controls to mitigate conflicts of interests,
including separate leadership, budgets, and deliverables, and the AI
team has written up public docs on their principles and conflict of
interest policy. But I think the strongest evidence here is from
April, when the AI team was looking at GPT 5.5. The grant was at what I
would consider its most sensitive stage: we had been working on it for
months with very positive signals, but we still didn't have an answer.
This was public internally, and AI leadership was looped in, but there
was never a question of this affecting what the AI team published, and
in their assessment
they documented a range of concerns. The biggest was that across
several benchmarks the model would appear to refuse a dangerous question
by deflecting, but actually it would go on to provide the requested
information by giving a highly-transferable related solution. On these
benchmarks the safeguards were illusory, and the AI team released their
evaluation while the grant was pending.
This concern with incentives, however, is not new with this grant.
When the AI side evaluates a company's models, that company typically
pays for the evaluation. For example, OpenAI PBC covered SecureBio's
costs for the GPT 5.5 evaluation above. This is common with AI evals,
and it's a tighter connection than this grant because there's no
AI-Detection division insulating evaluators from financial
incentives.
Still, this is a place where it takes continued effort to uphold
standards, and if there's any indication that funding for Detection is
being used as a lever to pressure our AI team on evals, I'll say so
publicly and use whatever leverage I have to stop it (up to and
including resignation). But I'm not expecting this, and I think the
real worry is a subtle drift towards being more generous without anyone
explicitly asking for anything. This is also a concern with funding from
Anthropic employees, since SecureBio also evaluates Anthropic's work
(ex).
So if you see SecureBio put out anything unfairly positive towards
OpenAI or Anthropic, or unfairly negative in overcorrecting for these
incentives, please say so.
On (b), the question is whether this will let model developers take
more risk. If we help them sleep better at night, knowing defenses are
stronger, will they just push ahead faster during the day? I want us to
be a complement to the frontier firms' internal safeguards, but what if
we become a substitute?
A world sufficiently robust against catastrophic biorisk, where it
doesn't matter what models are willing to explain because real-world
protections are a full substitute for model safeguards, would be a
fantastic place to be. We and manyotherprojects are working towards that
world! But there's a ton of work to get there. The worry is that AI
firms perceive risk as lower and ease up on their safeguards when the
risk is still unacceptably high.
I think this is directionally real, but as a direct substitute it's
small compared to the reputational, legal (liability + risk of
directives), and moral forces pushing firms to invest in safeguards.
This grant will allow us to flag attacks earlier, but doesn't come close
to mitigating the full impact of an attack and only covers one of
several paths to large-scale biological harm. The first-order positive
effects of making the world more robust to catastrophic biorisk are
really very likely to outweigh the second-order effects of reducing
safeguards; if I thought the other way around it would suggest that I
should instead do or fund work (virus hunting?) that visibly increases
risk in order to motivate others to step up, which seems like a terrible
idea.
Separately but relatedly, there's also a "political cover" angle,
where the PBC might point at this grant (even though it was made by
OAIF) to say that they're doing something, taking off some external
pressure. To the extent that this reduction in pressure lets them avoid
costly actions that would do more to reduce risk, this is a loss, and
it's my largest concern with this grant. Philanthropic funding should
not be a license to act recklessly, and if they offer it as an excuse we
shouldn't accept it. Please keep the pressure on OpenAI, and all the
other frontier firms, to slow down and prioritize reducing the risk that
their work leads to catastrophe. Even more important than pressuring
firms, however, is advocacy for thoughtful AI regulation: this is a
coordination problem where it's not in any individual firm's interest to
slow down even though it is in humanity's interest collectively.
On balance I think this grant is strongly positive, and I'm much more
worried that we won't do the best possible job pushing this work forward
than that our efforts will let OpenAI and others ease up on their own
work.
In May 2025 I met Yo Shavit, who was working on national security policy at OpenAI and was thinking about how to prepare for a future in which models could seriously assist attackers in creating pandemics. We had a call, and when I shared notes with my team their main response was: "maybe start with not making models that can do that?"
Which is, in many ways, fair: by continuing to push the frontier in biological capabilities, OpenAI's actions were making things worse on many of the problems SecureBio is trying to solve. But OpenAI stopping wouldn't have resolved the problem: other firms were pushing quickly too, and the economic incentives strongly favored rapid capability advancement. Making the world more resilient to pandemics needed to be a high priority regardless, especially in light of models' increasing ability to help people with biology.
When I thought about what our initial conversations might turn into, however, my primary concerns were whether that might (a) compromise SecureBio's ability to independently assess and criticize OpenAI's work or (b) make the world less safe via reducing model developers' motivation to improve safeguards. I do think there's something to both of these, which I get into more below, but it seemed well worth it to talk with Yo about how we could work together.
I explained how we were building an early warning system to flag outbreaks, especially engineered ones that could otherwise spread widely before detection. I described how metagenomic sequencing lets you see what nucleic acid sequences (DNA and RNA) are present in a sample, without choosing in advance which sequences to look for, and how we were piloting this on wastewater and nasal swabs. Over the next few months we discussed opportunities to accelerate our work, he introduced us to OpenAI cofounder Wojciech Zaremba, and both Yo and Wojciech left OpenAI PBC (the for-profit) for the OpenAI Foundation (OAIF). We continued talking to them in these new roles, and these discussions, plus a lot of due diligence, led to the $17.2M OAIF grant which we announced today. I'm incredibly excited about this grant, which will allow us to expand our monitoring system, reduce our turnaround time substantially, [1] and generally reduce the risk that something could spread widely before detection.
Which brings me back to the two concerns I mentioned above. On (a), independence, SecureBio has two divisions, Detection and AI. This grant funds Detection, while the AI side of SecureBio evaluates models from many firms, including OpenAI. The grant does not give OpenAI or OAIF any formal control over what anyone at SecureBio can say publicly about any models. That was important to us, but it was also key to OAIF: Yo was very clear that they did not want to influence SecureBio AI's work, and wanted to review our policies to make sure they were sufficiently robust.
That said, we should pay attention to incentives. While it looks to me like OAIF is making its own decisions, the two organizations are closely linked in a way that goes beyond the shared "OpenAI" name: the OpenAI Foundation's endowment is a ~1/4 stake in OpenAI PBC, and they have almost identical boards. Might the AI side of SecureBio pull punches in criticizing the PBC to increase the chances that OAIF gives Detection more money in the future?
SecureBio has systemic controls to mitigate conflicts of interests, including separate leadership, budgets, and deliverables, and the AI team has written up public docs on their principles and conflict of interest policy. But I think the strongest evidence here is from April, when the AI team was looking at GPT 5.5. The grant was at what I would consider its most sensitive stage: we had been working on it for months with very positive signals, but we still didn't have an answer. This was public internally, and AI leadership was looped in, but there was never a question of this affecting what the AI team published, and in their assessment they documented a range of concerns. The biggest was that across several benchmarks the model would appear to refuse a dangerous question by deflecting, but actually it would go on to provide the requested information by giving a highly-transferable related solution. On these benchmarks the safeguards were illusory, and the AI team released their evaluation while the grant was pending.
This concern with incentives, however, is not new with this grant. When the AI side evaluates a company's models, that company typically pays for the evaluation. For example, OpenAI PBC covered SecureBio's costs for the GPT 5.5 evaluation above. This is common with AI evals, and it's a tighter connection than this grant because there's no AI-Detection division insulating evaluators from financial incentives.
Still, this is a place where it takes continued effort to uphold standards, and if there's any indication that funding for Detection is being used as a lever to pressure our AI team on evals, I'll say so publicly and use whatever leverage I have to stop it (up to and including resignation). But I'm not expecting this, and I think the real worry is a subtle drift towards being more generous without anyone explicitly asking for anything. This is also a concern with funding from Anthropic employees, since SecureBio also evaluates Anthropic's work (ex). So if you see SecureBio put out anything unfairly positive towards OpenAI or Anthropic, or unfairly negative in overcorrecting for these incentives, please say so.
On (b), the question is whether this will let model developers take more risk. If we help them sleep better at night, knowing defenses are stronger, will they just push ahead faster during the day? I want us to be a complement to the frontier firms' internal safeguards, but what if we become a substitute?
A world sufficiently robust against catastrophic biorisk, where it doesn't matter what models are willing to explain because real-world protections are a full substitute for model safeguards, would be a fantastic place to be. We and many other projects are working towards that world! But there's a ton of work to get there. The worry is that AI firms perceive risk as lower and ease up on their safeguards when the risk is still unacceptably high.
I think this is directionally real, but as a direct substitute it's small compared to the reputational, legal (liability + risk of directives), and moral forces pushing firms to invest in safeguards. This grant will allow us to flag attacks earlier, but doesn't come close to mitigating the full impact of an attack and only covers one of several paths to large-scale biological harm. The first-order positive effects of making the world more robust to catastrophic biorisk are really very likely to outweigh the second-order effects of reducing safeguards; if I thought the other way around it would suggest that I should instead do or fund work (virus hunting?) that visibly increases risk in order to motivate others to step up, which seems like a terrible idea.
Separately but relatedly, there's also a "political cover" angle, where the PBC might point at this grant (even though it was made by OAIF) to say that they're doing something, taking off some external pressure. To the extent that this reduction in pressure lets them avoid costly actions that would do more to reduce risk, this is a loss, and it's my largest concern with this grant. Philanthropic funding should not be a license to act recklessly, and if they offer it as an excuse we shouldn't accept it. Please keep the pressure on OpenAI, and all the other frontier firms, to slow down and prioritize reducing the risk that their work leads to catastrophe. Even more important than pressuring firms, however, is advocacy for thoughtful AI regulation: this is a coordination problem where it's not in any individual firm's interest to slow down even though it is in humanity's interest collectively.
On balance I think this grant is strongly positive, and I'm much more worried that we won't do the best possible job pushing this work forward than that our efforts will let OpenAI and others ease up on their own work.
[1] Allowing me to finally answer a question that I asked four years ago, a few months before I quit Google to join the project.