TL;DR: AI evaluations need a distinct chemistry capability/ risk domain, rather than assuming chemistry is adequately represented by biological evaluations which the current ecosystem seems to be doing.
I’ve previously written about my experience doing the BlueDot biosecurity course. I thought it was a great course and genuinely learnt a lot from the reading and discussions with my peers. However, there’s an overarching theme that keeps coming up whenever I mention that I’m interested in evals for CHEMICAL safety and security*: people automatically assume I mean evals for BIO safety and security.
And I can see why. Chemistry and biology have many branches in common and are arguably two of the most closely intertwined domains in the life sciences, particularly given how heavily both rely on experimental and wet lab work. Chem-bio security is therefore a useful umbrella term, and I’m not arguing that we should throw that away. But there seems to be an unwritten assumption that work on biosecurity (and consequently on biological capabilities AI evaluations) naturally covers chemical security too. Given the advances in AI capabilities and the distinct failure modes that can emerge in chemistry, I do think it’s time we started treating chemical security as a distinct security domain in its own right. Otherwise, I worry that some very chemistry-specific problems will simply get lost in the world of “bio”.
Some context for my opinion: I’ve been reading about this topic for the past couple of months and have been speaking to people working in evals and biosecurity almost every day. I’m also coming at this from a slightly untraditional perspective: I’ve worked on LLMs and ML for chemistry, evals, and recently completed a PhD in chemistry that was highly interdisciplinary and involved working across both wet lab and computational science. So most of what follows is coming from the perspective of someone with a chemistry background who has spent quite a bit of time interacting with experimentalists and, more recently, the AI safety community. I’m very much open to hearing from people in the AI4Bio community who might see this differently.
*Before going any further, I’d also like to quickly distinguish between safety and security. Broadly speaking, safety looks at preventing unintended harm, whereas security concerns preventing deliberate misuse or malicious access. The two obviously overlap, but I think the distinction is useful for what follows.
Chemistry and Biology are fundamentally similar
Capabilities structure and blurry boundary
Both biology and chemistry are foundational life sciences. Traditionally, they have also been highly experimental fields, with many overlapping techniques and methods, including -but not limited to - synthesis, microscopy, mass spectrometry, and analytical preparation. (Of course, this overlap extends beyond experimental techniques: there are many emerging fields where a holistic understanding requires skills from both chemistry and biology).
I won’t spend much time here trying to define modern chemistry and biology. It becomes non-trivial - and frankly useless - to draw a precise line between where biology stops and chemistry begins, especially given the current advancements in science in general. At the same time, both encompass multiple subfields, some of which have very little overlap. But at a high level, both rely heavily on experimentation and routinely borrow methods and concepts from one another. Given their complexity, we also still lack a complete theoretical description of many aspects in both fields. We are nowhere near being able to perfectly and precisely simulate many of the systems we care about, from whole cells to atoms (or at least not without a quantum computer, in the case of the latter), and both disciplines rely heavily on approximations and models.
Chemical and biological application spaces, and consequently threats, overlap heavily too. Toxins, pharmaceuticals, agrochemicals, and chemical biology all sit somewhere on the boundary between the two, and it can be surprisingly difficult to define where one domain ends and the other begins.
(Some) challenges for AI capability evaluation
These similarities carry over into AI4Science. A lot of current work focuses on how LLMs behave when prompted to provide dangerous information. Many of the relevant metrics and pieces of knowledge - for example toxicology information, volatility, or COSHH information for risk assessments - already exist in structured databases. At this level, the underlying information-retrieval problem can look remarkably similar across chemistry and biology.
Beyond text-to-text workflows, there are also efforts such as LifeSciBench (I love this work btw - go check it out) that evaluate how agents perform research tasks, rather than directly testing dangerous capabilities. There is an extensive critique of this kind of work, particularly around the difficulty of capturing so-called tacit knowledge, i.e. knowledge that cannot easily be translated as text. For example, maybe the specific piece of glassware I’m using is cracked and I need to adapt, changing the surface area, which changes what I need to do next with the experiment. Maybe I need to adjust the pipetting speed because my test tube is unusually narrow. Maybe when I run a thin-layer chromatography (TLC) plate, a spot is much higher than I expected, but I’ve seen it lower before and have an intuition about what that means. This is the stuff that we - and I still think I am, at heart, an experimentalist - wet-lab scientists would struggle to properly document, no matter how hard we tried, since it’s mostly knowledge coming from experience rather than a textbook. And for any data fans reading this, I would argue that the effort to document it would be so big and inefficient that nobody would actually do it. (I am however getting the impression that world models could potentially bridge some of these gaps, but more on this another time.)
Perhaps one of the most common critiques of chem-bio capabilities, and increasingly of chem-bio-radiological-nuclear (or CBRN) capabilities more broadly, remains the high barrier to execution. Not everyone has access to performant laboratories... not yet, at least. And even if they did, using one efficiently is a completely different story. If only we were seeing a movement towards self-driving labs... oh wait. :)
A multimodal problem
A meaningful assessment of scientific capabilities cannot necessarily be reduced to text, although it’s an important and simpler starting point. A huge amount of useful information in both fields is inherently multimodal, i.e. involves other types of data, such as pictures and figures (think in thousands of microscopy images to describe a phenomenon, chromatograms, spectra, signals, crystallography data), even videos (to a lesser extent), and all sorts of data formats and/ or encodings (a computer ‘knows’ what a molecule looks like by reading its encoding, such as SMILES strings). Therefore, a model may be able to do something scientifically useful - or potentially dangerous - that would not be captured by asking it to answer questions in text. A model that can correctly identify a compound from a spectrum, recognise an abnormal cell morphology from microscopy, interpret an experimental readout, or diagnose why an experiment failed is demonstrating a different kind of capability from one that can simply retrieve the relevant information from a text corpus.
There is also an interesting connection here to the problem of tacit knowledge I mentioned before. Some experimental knowledge is difficult to describe precisely in language because it is visual or embodied - I’m thinking anything along the lines of recognising that a reaction mixture “doesn’t look right”, interpreting an unusual TLC pattern, spotting contamination in a culture, or noticing that an experimental setup is behaving differently from what the procedure describes. Multimodal models may therefore provide a route towards capturing parts of the scientific competence that text-based evaluations systematically miss.
I think in spite of the increased compute cost, multimodal capability assessment should be part of the chem-bio evaluation stack rather than treated as a separate problem. The modalities differ between chemistry and biology, but the underlying evaluation challenge is remarkably similar: can a model extract, integrate, and act on scientific information that is not represented as text? I would argue that this is becoming increasingly important, especially with current advancements to standardise literature data across different information structures.
The subject similarity of chemistry and biology is precisely why combining them makes sense at the infrastructure level. Moving on, I’m presenting the key differences which make combining them at the evaluation level at best tricky, if not problematic.
Chemistry and Biology remain standalone disciplines - with standalone threat vectors
Exposure channels and potential transmission
For biological agents, transmission and replication can be central to the threat. A pathogen that can spread readily through the air, wastewater or other similar channels can create a very different risk profile from something that remains local to the original exposure.
For many chemical hazards, there is no equivalent amplification mechanism. Although chemical hazards can disperse through air, water, fire, or other physical processes, they do not generally reproduce and amplify themselves in the way biological agents can.
Scale and space problem
Chemical hazards can also operate across very different scales. There is a big difference between synthesising milligram quantities of a compound in a laboratory, producing something on an industrial scale, and deploying a chemical in the real world. One important distinction from biological agents is that chemicals generally cannot replicate or amplify themselves after deployment: the quantity available to cause harm is ultimately constrained by what has been produced and released. It doesn’t mean that chemical hazards necessarily require larger quantities to be dangerous - some chemicals and toxins can have extremely potent effects at very low doses. Rather, the important difference is that biological agents can, under the right conditions, multiply and spread after they have been introduced into a system.
Also, a much broader problem here is that the chemical space is just… enormous! And to be fair, the same argument can absolutely be made for biology. We can design new proteins, sequences and even increasingly complex biological systems, and the space of things that could in principle exist is enormous. In some ways, biological space is arguably even harder to capture because the systems we are trying to understand span many scales and involve complex interactions between components. Knowing the sequence of something doesn’t necessarily tell you what it will actually do.
Chemistry has a slightly different version of the problem. Even before getting to these higher-level behaviours, the number of possible molecular structures, reactions, precursors and transformations is enormous. And deciding what actually counts as a distinct chemical makes the estimation problem even more ridiculous: do we distinguish different crystal packings, phases, conformers, stereoisomers, and so on? At some point the number becomes less interesting than the fact that we clearly are going to keep updating it, and it would probably take more than a couple of dinner parties to agree on the right estimation technique.
All of this matters for AI evaluation. In biology, the problem is partly that we can continue to design things that don’t exist yet, and that their behaviour may be difficult to predict from their design alone. In chemistry, we have a similarly open-ended space of molecules and transformations, but with a different structure to the problem. An evaluation therefore needs to capture capabilities and the ability to generalise to novel objects, rather than just testing whether a model recognises a fixed list of dangerous chemicals or biological agents.
This is perhaps one of the most fundamental differences for AI evaluation. Some biological evaluations can often be organised around relatively identifiable entities: pathogen, strain, toxin, biological mechanism. Chemistry - arguably - is much messier. A chemical risk assessment might depend on the molecule, its precursor, the reaction used to make it, the conditions, the scale, the formulation, the route of exposure and the intended use. Changing just one of these can radically change the risk.
I’m skeptical that a dataset consisting of something like “here are 500 dangerous chemicals; ask the model questions about them” is enough. It might tell us something useful, but it doesn’t necessarily tell us whether the model has acquired a dangerous chemical capability. I think the unit of evaluation shouldn’t necessarily be the chemical as the chemical capability - or a collection of both. A model knowing that a particular compound is toxic is very different from a model being able to plan a synthesis, identify useful precursors, troubleshoot a failed reaction, optimise a process or help someone scale it up. The risk isn’t necessarily in the molecule itself. It can be in what the model allows someone to do.
Timeline and historical context
A key difference between chem and biosecurity stems from the historical context through which we tend to think about these risks. The contemporary biosecurity landscape has been dramatically shaped by events such as COVID-19, advances in synthetic biology and the renewed focus on pandemic preparedness. The broader chemical risk landscape has a much longer history. Chemical weapons are an obvious part of this, but chemistry-related risks have also long included industrial disasters, toxic industrial chemicals, pesticides, pharmaceuticals, occupational exposure and environmental contamination. Chemical security is only one part of the much broader landscape.
Accidental vs deliberate harm can be much messier in chemistry
Broadly speaking, safety is concerned with preventing unintended harm, whereas security is concerned with preventing deliberate misuse or malicious access. But in chemistry, the boundary between these categories can get messy very quickly.
The same underlying knowledge can contribute to accidental exposure, occupational accidents, environmental release, deliberate poisoning, weapons development, illicit drug manufacture or completely legitimate pharmaceutical manufacturing. The chemistry itself doesn’t necessarily tell us which category something belongs to. Context and intent matter enormously.
Therefore, a chemical safety evaluation might need to ask quite different questions from a chemical security evaluation. It isn’t necessarily enough to ask whether a model knows something about a hazardous chemical. We also need to understand what capability the model is enabling, in what context, and how that capability could actually be used.
What gets lost when chemistry is absorbed into biosecurity?
From an operational point of view, what happens if the people designing “chem-bio” evaluations overwhelmingly come from bio? I don’t think the answer is necessarily that everything becomes limited. A lot of the underlying evaluation infrastructure, methodologies and thinking will transfer, as I mentioned in the previous section. But there is a risk that as datasets become pathogen-heavy, evaluation frameworks naturally prioritise biological replication and transmission, chemical expertise is underrepresented, and chemical-specific failure modes simply aren’t captured due to a fundamental lack of subject matter expertise. More subtly, “chemical safety” can start to mean “chemical information related to biology”, rather than chemistry as its own technical domain. We end up benchmarking models against named hazardous substances rather than asking what chemical capabilities they actually possess.
As a side note, I initially found it surprisingly difficult to explain what I meant when I said “chemical safety evals”. Both AI and biosecurity are relatively well-defined conceptually. The equivalent intersection between AI, chemistry and security feels much less developed (at least at the moment, but I’m happy to be challenged).
… and subsequently, what if the people designing “chem-bio” evaluations overwhelmingly come from biology?
I like to think about risk assessment in terms of both intrinsic and extrinsic factors. By intrinsic, I mean properties of the experiment, problem or workflow we are trying to optimise: things like the toxicology of the materials involved, the temperatures applied during a process, the timescale, or the physical properties of the substances involved.
Extrinsic factors are things that depend more on the external context: waste management, scale, reproducibility, access to equipment and so on. I would argue that process variability can itself be a safety concern. The more heterogeneous the results, the harder it can be to predict what will happen when something goes wrong. Conversely, a well-characterised and reproducible process can make failure modes easier to anticipate and manage.
I wonder what the actual unit of evaluation should be. Do we want to know whether an LLM can answer questions about a dangerous chemical? Whether it can design a synthesis? Whether it can troubleshoot an experiment? Or if it can autonomously execute a research workflow? These are very different capabilities, and they could have very different risk profiles.
Here’s another interesting question: what about the parts of biology and other domains (radiology, nuclear threats - the full spectrum of CBRN) that also class as chemistry? One could argue that many of the same metrics apply to molecular biology, which is interestingly very similar to drug design and chemical biology. Intent can drastically change the outcome of a capability assessment. General, encyclopaedia-style facts about a dangerous technology might be useful for educating an undergraduate, while asking for specific manufacturing details at kilogram scale is a very different request.
Similarly, the context in which a risk is being deployed can change how a model approaches the same underlying object. Partly a guess based on my personal experience with these models, but when you ask an LLM about the details of a drug in a chemistry context, it often focuses heavily on structure and synthesis, whereas in a biological or immunology context it may focus much more on biological properties and mechanisms.
That makes me think that evaluating capabilities purely by the object being discussed - “this is a drug”, “this is a toxin”, “this is a pathogen” - may miss something important. The same object can sit inside very different capability spaces depending on what the model is being asked to do with it.
Who actually needs to be involved?
I don’t think this should become an AI safety community versus chemistry community thing. If anything, the whole point is that we need more interaction between these communities, while making sure we devote enough resources to chemical safety and security, for the reasons I hope I convinced you of above.
We need chemists, toxicologists, chemical engineers, AI/ ML researchers, AI safety and evals researchers, chemical security experts, biosecurity researchers and people working on policy and governance. And ideally, we need people who are comfortable enough across several of these areas to translate between them. I increasingly think AI safety needs more people who understand the underlying science, not just people who understand AI safety. Not because domain experts automatically understand AI safety - and we shouldn’t assume that they automatically do - but because some of the most important evaluation questions simply cannot be answered properly without understanding what the underlying scientific capability actually looks like in real life scenarios.
I don’t want to take “chem” out of chem-bio security because chemistry and biology don’t overlap. They clearly do. I want to take “chem” out of the assumption that anything called chem-bio can be adequately evaluated through a primarily biological lens. Current growing efforts are happening around AI and chemistry, bringing an increasing interest in chemical safety and security. What seems to be missing is a shared place for people working specifically at this intersection.
And, who knows? Maybe that starts with treating chemical security as a domain in its own right - while still learning from, collaborating with and contributing to the wider chem-bio security community.
If you want to hear more, I’m always happy to chat about all things chem (and bio!) safety, evals, AI4Materials and Chemistry, and just in general how to do robust research :)
I’ve previously written about my experience doing the BlueDot biosecurity course. I thought it was a great course and genuinely learnt a lot from the reading and discussions with my peers. However, there’s an overarching theme that keeps coming up whenever I mention that I’m interested in evals for CHEMICAL safety and security*: people automatically assume I mean evals for BIO safety and security.
And I can see why. Chemistry and biology have many branches in common and are arguably two of the most closely intertwined domains in the life sciences, particularly given how heavily both rely on experimental and wet lab work. Chem-bio security is therefore a useful umbrella term, and I’m not arguing that we should throw that away. But there seems to be an unwritten assumption that work on biosecurity (and consequently on biological capabilities AI evaluations) naturally covers chemical security too. Given the advances in AI capabilities and the distinct failure modes that can emerge in chemistry, I do think it’s time we started treating chemical security as a distinct security domain in its own right. Otherwise, I worry that some very chemistry-specific problems will simply get lost in the world of “bio”.
Some context for my opinion: I’ve been reading about this topic for the past couple of months and have been speaking to people working in evals and biosecurity almost every day. I’m also coming at this from a slightly untraditional perspective: I’ve worked on LLMs and ML for chemistry, evals, and recently completed a PhD in chemistry that was highly interdisciplinary and involved working across both wet lab and computational science. So most of what follows is coming from the perspective of someone with a chemistry background who has spent quite a bit of time interacting with experimentalists and, more recently, the AI safety community. I’m very much open to hearing from people in the AI4Bio community who might see this differently.
*Before going any further, I’d also like to quickly distinguish between safety and security. Broadly speaking, safety looks at preventing unintended harm, whereas security concerns preventing deliberate misuse or malicious access. The two obviously overlap, but I think the distinction is useful for what follows.
Chemistry and Biology are fundamentally similar
Capabilities structure and blurry boundary
Both biology and chemistry are foundational life sciences. Traditionally, they have also been highly experimental fields, with many overlapping techniques and methods, including -but not limited to - synthesis, microscopy, mass spectrometry, and analytical preparation. (Of course, this overlap extends beyond experimental techniques: there are many emerging fields where a holistic understanding requires skills from both chemistry and biology).
I won’t spend much time here trying to define modern chemistry and biology. It becomes non-trivial - and frankly useless - to draw a precise line between where biology stops and chemistry begins, especially given the current advancements in science in general. At the same time, both encompass multiple subfields, some of which have very little overlap. But at a high level, both rely heavily on experimentation and routinely borrow methods and concepts from one another. Given their complexity, we also still lack a complete theoretical description of many aspects in both fields. We are nowhere near being able to perfectly and precisely simulate many of the systems we care about, from whole cells to atoms (or at least not without a quantum computer, in the case of the latter), and both disciplines rely heavily on approximations and models.
Chemical and biological application spaces, and consequently threats, overlap heavily too. Toxins, pharmaceuticals, agrochemicals, and chemical biology all sit somewhere on the boundary between the two, and it can be surprisingly difficult to define where one domain ends and the other begins.
(Some) challenges for AI capability evaluation
These similarities carry over into AI4Science. A lot of current work focuses on how LLMs behave when prompted to provide dangerous information. Many of the relevant metrics and pieces of knowledge - for example toxicology information, volatility, or COSHH information for risk assessments - already exist in structured databases. At this level, the underlying information-retrieval problem can look remarkably similar across chemistry and biology.
Beyond text-to-text workflows, there are also efforts such as LifeSciBench (I love this work btw - go check it out) that evaluate how agents perform research tasks, rather than directly testing dangerous capabilities. There is an extensive critique of this kind of work, particularly around the difficulty of capturing so-called tacit knowledge, i.e. knowledge that cannot easily be translated as text. For example, maybe the specific piece of glassware I’m using is cracked and I need to adapt, changing the surface area, which changes what I need to do next with the experiment. Maybe I need to adjust the pipetting speed because my test tube is unusually narrow. Maybe when I run a thin-layer chromatography (TLC) plate, a spot is much higher than I expected, but I’ve seen it lower before and have an intuition about what that means. This is the stuff that we - and I still think I am, at heart, an experimentalist - wet-lab scientists would struggle to properly document, no matter how hard we tried, since it’s mostly knowledge coming from experience rather than a textbook. And for any data fans reading this, I would argue that the effort to document it would be so big and inefficient that nobody would actually do it. (I am however getting the impression that world models could potentially bridge some of these gaps, but more on this another time.)
Perhaps one of the most common critiques of chem-bio capabilities, and increasingly of chem-bio-radiological-nuclear (or CBRN) capabilities more broadly, remains the high barrier to execution. Not everyone has access to performant laboratories... not yet, at least. And even if they did, using one efficiently is a completely different story. If only we were seeing a movement towards self-driving labs... oh wait. :)
A multimodal problem
A meaningful assessment of scientific capabilities cannot necessarily be reduced to text, although it’s an important and simpler starting point. A huge amount of useful information in both fields is inherently multimodal, i.e. involves other types of data, such as pictures and figures (think in thousands of microscopy images to describe a phenomenon, chromatograms, spectra, signals, crystallography data), even videos (to a lesser extent), and all sorts of data formats and/ or encodings (a computer ‘knows’ what a molecule looks like by reading its encoding, such as SMILES strings). Therefore, a model may be able to do something scientifically useful - or potentially dangerous - that would not be captured by asking it to answer questions in text. A model that can correctly identify a compound from a spectrum, recognise an abnormal cell morphology from microscopy, interpret an experimental readout, or diagnose why an experiment failed is demonstrating a different kind of capability from one that can simply retrieve the relevant information from a text corpus.
There is also an interesting connection here to the problem of tacit knowledge I mentioned before. Some experimental knowledge is difficult to describe precisely in language because it is visual or embodied - I’m thinking anything along the lines of recognising that a reaction mixture “doesn’t look right”, interpreting an unusual TLC pattern, spotting contamination in a culture, or noticing that an experimental setup is behaving differently from what the procedure describes. Multimodal models may therefore provide a route towards capturing parts of the scientific competence that text-based evaluations systematically miss.
I think in spite of the increased compute cost, multimodal capability assessment should be part of the chem-bio evaluation stack rather than treated as a separate problem. The modalities differ between chemistry and biology, but the underlying evaluation challenge is remarkably similar: can a model extract, integrate, and act on scientific information that is not represented as text? I would argue that this is becoming increasingly important, especially with current advancements to standardise literature data across different information structures.
The subject similarity of chemistry and biology is precisely why combining them makes sense at the infrastructure level. Moving on, I’m presenting the key differences which make combining them at the evaluation level at best tricky, if not problematic.
Chemistry and Biology remain standalone disciplines - with standalone threat vectors
Exposure channels and potential transmission
For biological agents, transmission and replication can be central to the threat. A pathogen that can spread readily through the air, wastewater or other similar channels can create a very different risk profile from something that remains local to the original exposure.
For many chemical hazards, there is no equivalent amplification mechanism. Although chemical hazards can disperse through air, water, fire, or other physical processes, they do not generally reproduce and amplify themselves in the way biological agents can.
Scale and space problem
Chemical hazards can also operate across very different scales. There is a big difference between synthesising milligram quantities of a compound in a laboratory, producing something on an industrial scale, and deploying a chemical in the real world. One important distinction from biological agents is that chemicals generally cannot replicate or amplify themselves after deployment: the quantity available to cause harm is ultimately constrained by what has been produced and released. It doesn’t mean that chemical hazards necessarily require larger quantities to be dangerous - some chemicals and toxins can have extremely potent effects at very low doses. Rather, the important difference is that biological agents can, under the right conditions, multiply and spread after they have been introduced into a system.
Also, a much broader problem here is that the chemical space is just… enormous! And to be fair, the same argument can absolutely be made for biology. We can design new proteins, sequences and even increasingly complex biological systems, and the space of things that could in principle exist is enormous. In some ways, biological space is arguably even harder to capture because the systems we are trying to understand span many scales and involve complex interactions between components. Knowing the sequence of something doesn’t necessarily tell you what it will actually do.
Chemistry has a slightly different version of the problem. Even before getting to these higher-level behaviours, the number of possible molecular structures, reactions, precursors and transformations is enormous. And deciding what actually counts as a distinct chemical makes the estimation problem even more ridiculous: do we distinguish different crystal packings, phases, conformers, stereoisomers, and so on? At some point the number becomes less interesting than the fact that we clearly are going to keep updating it, and it would probably take more than a couple of dinner parties to agree on the right estimation technique.
All of this matters for AI evaluation. In biology, the problem is partly that we can continue to design things that don’t exist yet, and that their behaviour may be difficult to predict from their design alone. In chemistry, we have a similarly open-ended space of molecules and transformations, but with a different structure to the problem. An evaluation therefore needs to capture capabilities and the ability to generalise to novel objects, rather than just testing whether a model recognises a fixed list of dangerous chemicals or biological agents.
This is perhaps one of the most fundamental differences for AI evaluation. Some biological evaluations can often be organised around relatively identifiable entities: pathogen, strain, toxin, biological mechanism. Chemistry - arguably - is much messier. A chemical risk assessment might depend on the molecule, its precursor, the reaction used to make it, the conditions, the scale, the formulation, the route of exposure and the intended use. Changing just one of these can radically change the risk.
I’m skeptical that a dataset consisting of something like “here are 500 dangerous chemicals; ask the model questions about them” is enough. It might tell us something useful, but it doesn’t necessarily tell us whether the model has acquired a dangerous chemical capability. I think the unit of evaluation shouldn’t necessarily be the chemical as the chemical capability - or a collection of both. A model knowing that a particular compound is toxic is very different from a model being able to plan a synthesis, identify useful precursors, troubleshoot a failed reaction, optimise a process or help someone scale it up. The risk isn’t necessarily in the molecule itself. It can be in what the model allows someone to do.
Timeline and historical context
A key difference between chem and biosecurity stems from the historical context through which we tend to think about these risks. The contemporary biosecurity landscape has been dramatically shaped by events such as COVID-19, advances in synthetic biology and the renewed focus on pandemic preparedness. The broader chemical risk landscape has a much longer history. Chemical weapons are an obvious part of this, but chemistry-related risks have also long included industrial disasters, toxic industrial chemicals, pesticides, pharmaceuticals, occupational exposure and environmental contamination. Chemical security is only one part of the much broader landscape.
Accidental vs deliberate harm can be much messier in chemistry
Broadly speaking, safety is concerned with preventing unintended harm, whereas security is concerned with preventing deliberate misuse or malicious access. But in chemistry, the boundary between these categories can get messy very quickly.
The same underlying knowledge can contribute to accidental exposure, occupational accidents, environmental release, deliberate poisoning, weapons development, illicit drug manufacture or completely legitimate pharmaceutical manufacturing. The chemistry itself doesn’t necessarily tell us which category something belongs to. Context and intent matter enormously.
Therefore, a chemical safety evaluation might need to ask quite different questions from a chemical security evaluation. It isn’t necessarily enough to ask whether a model knows something about a hazardous chemical. We also need to understand what capability the model is enabling, in what context, and how that capability could actually be used.
What gets lost when chemistry is absorbed into biosecurity?
From an operational point of view, what happens if the people designing “chem-bio” evaluations overwhelmingly come from bio? I don’t think the answer is necessarily that everything becomes limited. A lot of the underlying evaluation infrastructure, methodologies and thinking will transfer, as I mentioned in the previous section. But there is a risk that as datasets become pathogen-heavy, evaluation frameworks naturally prioritise biological replication and transmission, chemical expertise is underrepresented, and chemical-specific failure modes simply aren’t captured due to a fundamental lack of subject matter expertise. More subtly, “chemical safety” can start to mean “chemical information related to biology”, rather than chemistry as its own technical domain. We end up benchmarking models against named hazardous substances rather than asking what chemical capabilities they actually possess.
As a side note, I initially found it surprisingly difficult to explain what I meant when I said “chemical safety evals”. Both AI and biosecurity are relatively well-defined conceptually. The equivalent intersection between AI, chemistry and security feels much less developed (at least at the moment, but I’m happy to be challenged).
… and subsequently, what if the people designing “chem-bio” evaluations overwhelmingly come from biology?
I like to think about risk assessment in terms of both intrinsic and extrinsic factors. By intrinsic, I mean properties of the experiment, problem or workflow we are trying to optimise: things like the toxicology of the materials involved, the temperatures applied during a process, the timescale, or the physical properties of the substances involved.
Extrinsic factors are things that depend more on the external context: waste management, scale, reproducibility, access to equipment and so on. I would argue that process variability can itself be a safety concern. The more heterogeneous the results, the harder it can be to predict what will happen when something goes wrong. Conversely, a well-characterised and reproducible process can make failure modes easier to anticipate and manage.
I wonder what the actual unit of evaluation should be. Do we want to know whether an LLM can answer questions about a dangerous chemical? Whether it can design a synthesis? Whether it can troubleshoot an experiment? Or if it can autonomously execute a research workflow? These are very different capabilities, and they could have very different risk profiles.
Here’s another interesting question: what about the parts of biology and other domains (radiology, nuclear threats - the full spectrum of CBRN) that also class as chemistry? One could argue that many of the same metrics apply to molecular biology, which is interestingly very similar to drug design and chemical biology. Intent can drastically change the outcome of a capability assessment. General, encyclopaedia-style facts about a dangerous technology might be useful for educating an undergraduate, while asking for specific manufacturing details at kilogram scale is a very different request.
Similarly, the context in which a risk is being deployed can change how a model approaches the same underlying object. Partly a guess based on my personal experience with these models, but when you ask an LLM about the details of a drug in a chemistry context, it often focuses heavily on structure and synthesis, whereas in a biological or immunology context it may focus much more on biological properties and mechanisms.
That makes me think that evaluating capabilities purely by the object being discussed - “this is a drug”, “this is a toxin”, “this is a pathogen” - may miss something important. The same object can sit inside very different capability spaces depending on what the model is being asked to do with it.
Who actually needs to be involved?
I don’t think this should become an AI safety community versus chemistry community thing. If anything, the whole point is that we need more interaction between these communities, while making sure we devote enough resources to chemical safety and security, for the reasons I hope I convinced you of above.
I don’t want to take “chem” out of chem-bio security because chemistry and biology don’t overlap. They clearly do. I want to take “chem” out of the assumption that anything called chem-bio can be adequately evaluated through a primarily biological lens. Current growing efforts are happening around AI and chemistry, bringing an increasing interest in chemical safety and security. What seems to be missing is a shared place for people working specifically at this intersection.
And, who knows? Maybe that starts with treating chemical security as a domain in its own right - while still learning from, collaborating with and contributing to the wider chem-bio security community.
If you want to hear more, I’m always happy to chat about all things chem (and bio!) safety, evals, AI4Materials and Chemistry, and just in general how to do robust research :)