A question shadows our AI era: What if this technology, whose triumphs might include prolonging human lives and conserving our planet, were to enable an attack involving chemical, biological, radiological, or nuclear material? Thankfully, such risks—lumped together as “CBRN”—are historically rare, and would-be attackers typically fail. Yet the risk is such that AI companies, governments, and safety researchers are working assiduously to mitigate it.
Most efforts take place out of sight, lest exposure offer help to a potential attacker. The downside is that public discourse on AI and CBRN is often limited, slanted, or heavy with speculation. With this challenge in mind, we sat down with five people who work on CBRN risks at Google DeepMind—Paige Kunkle, Adam Marsh, Ash Otter, Jeremy Ratcliff, and Victoria Langston—to better understand the issue.
This is what we learned.
1. What are CBRN risks?
CBRN covers a wide range of hazards, from the alleged poisoning of Russian opposition leader Alexei Navalny with a toxin found in South American frogs, to the recent drone attack on a nuclear power plant in the United Arab Emirates that risked a “very, very serious” radiological incident.
The AI community focuses on preventing threat actors from using the technology to aid such attacks. But there is also growing interest in using AI to boost society’s resilience to attacks, as well as to naturally occurring events (like the recent Ebola outbreak) or accidents, such as when scrap sellers inadvertently expose others to harmful radioactive materials found in medical or industrial devices.
2. Why address the four CBRN components together?
CBRN risks have much in common. The most likely threat actors share certain characteristics, such as a desire to outwit an adversary who outmatches them in more conventional weapons, or to cause widespread fear with plausible deniability. Policy responses, such as international treaties, export controls, and customer screening, look similar across domains, as do the first responders to any incident, such as the police and the military.
Having a single CBRN team also enables AI labs to manage the growing overlaps with other risks they work on, such as cybersecurity, disinformation, and conventional explosive attacks—many CBRN weapons, like a radiological dispersal device, or “dirty bomb”, require explosives to disperse the material.
The four CBRN domains do differ. Acquiring the material for a bioweapon may involve finding and swabbing a dead animal. A radiological weapon may require sourcing and prying open a medical device. These processes, and the uplift that AI may provide, look very different—which is why specialist expertise in each domain is essential.
3. Is the world worrying too much about CBRN risks, or not enough?
When asked, the public expresses great concern about CBRN risks, including from AI. But this rarely translates into sustained attention, political will, or funding at the scale required.
Why? CBRN incidents are rare and underappreciated. When the security services thwart an attack, they deliberately say little about it. Products that could make society more resilient, such as vaccines, prophylactics for nerve-agent attacks, and alternatives to radioactive materials in medical devices, remain underdeveloped, untested, or not adopted at scale. This can be due to a lack of commercial incentives—the market for vaccines is often small, temporary, and hard to predict. Or other obstacles—it is hard to reliably and safely test a nerve-agent prophylactic.
The AI community does devote significant attention to CBRN risks in safety evaluations, governance documents, and media articles. But this attention is skewed toward bioweapons and in particular toward the concern that lone-wolf attackers may use AI to engineer a virus.
There is some logic to this. The self-replicating nature of pathogens means that such bioweapons could cause millions of deaths. But the technical obstacles are high and there hasn’t been a publicly documented pathogen attack, fatal to humans, since the 2001 US anthrax attacks—when lethal spores were mailed to media outlets and senators, killing five people. Conversely, several actors have used chemical weapons during this time, including Russia in Ukraine, the Assad government in Syria, and ISIS. Biological toxins like ricin, or a conventional explosive attack, would also be a more tractable option for most terrorist groups than trying to engineer a virus.
Explosives, toxins, and radiological materials do not spread like a virus, but they can still impose significant harm. In 1987, four people died in Goiânia, Brazil, after they were inadvertently exposed to cesium-137, a highly radioactive material, by two scrap sellers who had pried open an old radiotherapy machine. The wider decontamination effort saw homes destroyed and soil torn up, while tourism collapsed and residents faced discrimination. Today, dangerous radioactive materials are dotted across the world, with some vulnerable to theft and misuse.
For AI labs, the challenge is to allocate resources across all four CBRN domains in a way that prioritizes the most concerning risks, but does not lose sight of the wide range of plausible scenarios and society’s generally low resilience to them.
4. Does AI’s scientific upside outweigh its CBRN risks?
Optimists argue that science’s benefits have far outweighed its costs, and that using AI to accelerate science will be similar. Smallpox is estimated to have killed more people in the 20th century alone than all of that century’s genocides and military conflicts combined. In 1980, it was officially eradicated, in part due to vaccines. Skeptics counter that AI might favor attackers over defenders, and that this imbalance may be particularly strong in domains like nuclear and radiological weapons.
The need to use AI to accelerate science is arguably strongest in biology, because nature poses so many risks that we need to respond to. Urbanization, rising temperatures and deforestation are causing humans to encroach on animals’ habitats, increasing the risk that a pathogen jumps species. Food security is at risk as we pack ever more genetically similar plants and animals into dense conditions. Even without novel outbreaks, common infectious diseases like influenza and pneumonia kill more than 1 million children every year.
AI could help tackle these risks, but some advances may be vulnerable to misuse. For example, researchers have developed AI models that can predict which variants could cause a common virus to escape the immunity humans have built up. Such models can also predict the parts of the virus that are unable to mutate without harming the virus’s survival. The latter could serve as targets for new drugs and vaccines, which AI might help create. But safety advocates worry that adversaries could intentionally design harmful variants of the virus.
For radiological and nuclear risks, observers worry that AI may disproportionately favor attackers because many risks are computational and information-based, such as using LLMs to parse the huge amounts of regulatory information published by nuclear facilities to extract insights to help with an attack. The defenses against such attacks still rely largely on physical security measures—like armed security personnel, multi-layered access controls, and reinforced concrete containment structures. There are defensive AI applications, like helping to detect unusual radiation signatures, but there is no AI-enabled patch for the physical consequences of a catastrophic breach or the detonation of a nuclear weapon. The nuclear security community also strongly opposes integrating AI into the command, control, and communications architectures they use for early warning and weapons authorization, due to concerns over reliability, compressed decision-making and more.
Enabling only the positive AI for CBRN applications and never the misuse sounds impossible. But governments, companies and scientists have long had to designate certain infrastructure, information and tools as “higher risk” and control access to them—for example via strong security and export controls—while still supporting beneficial downstream applications. AI labs will need to draw on this experience in the coming years.
5. AI models lack practical know-how. Does that reduce the risk of a CBRN catastrophe?
Yes, but the bottlenecks are bigger in biology than in other domains and may weaken over time.
A seasoned biologist knows how to culture a cell even if they can’t articulate how exactly they do it. Threat actors often lack this kind of tacit knowledge. In 1995, Aum Shinrikyo, the Japanese yoga-school-turned-death-cult, used sarin gas to kill 13 people and wound more than 6,000 in an attack on the Tokyo subway. Yet it could have been far worse. Prior to pivoting to chemical weapons, the group attempted biological attacks that all failed. One review found that although the group had amassed significant scientific expertise, they still made a multitude of missteps—from using the wrong bacteria strains to contaminating the fermentation process.
What if Aum Shinrikyo had had access to modern AI models? Although not a direct approximation, in 2025, the non-profit group Active Site ran a randomized controlled trial to determine whether novices could safely perform the kinds of wet-lab tasks needed to synthesize a virus from its genetic sequence. One group had internet access; a second group also had LLM access. Tasks included using pipettes to measure and transfer liquid, growing and maintaining a living cell, and combining DNA fragments. To the surprise of experts polled beforehand, the study found no statistically significant uplift for the group with LLM access.
Why? In biology, written instructions for many common tasks are publicly available. However, scientists learn to apply these protocols through hands-on training and practice—along with lots of failure. From this, they gain a range of skills, including manual dexterity and a sense of touch. Current AI models lack any such equivalent.
From a CBRN-risk perspective, the Active Site study is reassuring, but only partially.
Biology is notoriously complex, noisy and hard to predict—making real-world tinkering and testing paramount. This is less true in other domains. Aum Shinrikyo members were able to master the basics of sarin, a chemical weapon, from the scientific literature. For nuclear weapons, AI-enabled simulations could provide highly accurate predictions, making computational-only assistance a critical non-proliferation risk.
Participants in the Active Site trial were also novices, including at using AI, so the results may be a poor guide to the uplift that more skilled bad actors might gain from the technology—the 2001 US anthrax attacks were allegedly carried out by a microbiologist, Bruce Edwards Ivins.
AI’s tacit knowledge may also grow if models are exposed to a richer view of the full scientific process, such as from lab notebooks or videos of scientists performing experiments. Emerging deployment surfaces, such as smart “XR” glasses, could provide users with more direct error detection and troubleshooting support, weakening the tacit knowledge barrier.
Robots could bypass the need for human tacit knowledge altogether, while sparing attackers the risk of serious injury. Robots already perform specialized CBRN tasks, such as handling liquids or radioactive materials. General-purpose robotics foundation models could enable machines to operate a wider range of equipment. However, reliably automating human experts’ sophisticated sense of touch will be hard. The companies building automated labs are also subject to laws prohibiting CBRN weapons and are unlikely to prioritize the niche, high-risk workflows that threat actors would require.
For AI labs, the main takeaway is that tacit knowledge is a bottleneck, but it varies by domain and may weaken over time, so evaluating AI models’ performance is a top priority.
The AI-lab response to CBRN risks
6. How do AI companies decide which CBRN threats to prioritize?
The threat landscape is almost limitless and riddled with uncertainties. The number of historical examples is small, and those who understand the emerging threats best typically can’t speak openly about them. So AI companies must work closely with governments and experts to understand the motivations, capabilities and methods of the most likely threat actors.
This threat modeling requires mapping the intermediate steps that an actor would follow in a given scenario and determining the most important bottlenecks where AI might provide uplift over public information. For example, in the ideation phase, a threat actor might seek guidance on the feasibility of different attacks. In planning and preparation, they might look for practical blueprints or training. In production and weaponization, they might look to troubleshoot the synthesis, fabrication, or transport of the weapons—or garner tips to evade detection.
The characteristics of a threat actor shape the kind of uplift they need. A lone terrorist with limited expertise and resources may want clear instructions and readily available materials. A large militia group may want advice on how to target infrastructure or run training programs.
Some kinds of uplift are also more consequential than others. Is AI unlocking a novel capability or speeding up something that threat actors can already do? Does the uplift remove a critical bottleneck, or do harder challenges remain? Could prospective attackers access this dangerous information in other ways?
One challenge is that experts often disagree about adversaries’ capabilities and motivations. When it comes to nation-states, some argue that bioweapons hold little strategic rationale as they are hard to develop and control, and their use would lead to public opprobrium. Others counter that nation-states’ actions are contingent on what their adversaries do, and that North Korea and possibly others are reported to have active bioweapon programs.
Long-standing norms against weapons use can also change. In 2018, the Russian military-intelligence agency GRU allegedly conducted the first offensive use of a chemical weapon on Western European soil since World War II, in the attempted assassination of Sergei Skripal and his daughter with the Novichok nerve agent. Even if certain actors don’t wish to use CBRN weapons, they may seek to stockpile weapons or materials for leverage. These stockpiles are then at risk of being stolen or abandoned in times of strife, such as when war breaks out. In 2025, the IAEA estimated that Iran possessed more than 440 kg of 60% enriched uranium, before it had to withdraw its inspectors from the country.
Ultimately, AI labs need threat models that are stable enough to allow them to make progress on the risks, but flexible enough to respond to dynamic geopolitics and rapid AI progress.
7. How do AI companies evaluate if their models might help threat actors?
Leading AI developers typically set thresholds for what they consider risky CBRN capabilities. For example, Google DeepMind’s Frontier Safety Framework judges a model to have reached a “critical capability level” if, in reference scenarios, it could provide low- to medium-resourced actors with enough of a capability uplift to risk severe harm.
Labs and their external partners use a variety of evaluation methods to test whether AI models reach these thresholds—although most evaluations remain unpublished, to avoid inadvertently aiding bad actors.
The first category is automated evaluations, like Lab-Bench, which ask AI models and agents questions, for example about molecular-biology protocols, or have them carry out tasks, such as writing code to operate liquid-handling robots—in a way that other AI systems can judge. These evaluations offer breadth and speed, but focus heavily on biology, at the expense of other domains, like radiological and nuclear weapons. They also struggle to capture the messy nature of biology, provide a clear signal about whether a risky capability has been reached, and can understate the capabilities that a skilled human could elicit from the model or agent.
Expert-led evaluations can deliver a higher-confidence assessment by allowing experienced humans to probe the AI over multi-turn interactions and explore specific scenarios. This includes expert red teamers trying to adversarially extract knowledge out of the model. These evaluations can offer greater signal, but are complex to run, with outputs that may run to hundreds of pages, making them challenging to interpret and standardize.
A third category is real-world trials, such as the Active Site RCT. They evaluate models’ ability to instruct humans, and conceivably robots, in a real-world environment. Such evaluations can offer the richest insight, but they are also slow, expensive, and hard to design well. They may also not generalize beyond the specific groups and tasks studied.
Looking ahead, there are several ways to improve AI CBRN evaluations. Priorities include:
Linking different methods—for example, using fast automated evaluations to check if models are progressing on the limitations that slower real-world trials highlight.
Designing more sophisticated evaluations for fast-improving AI agents that pair frontier LLMs with specialized science models and third-party databases and tools.
Developing privacy-preserving methods to allow third parties to securely evaluate AI labs’ models, without either side leaking sensitive information.
8. How do AI labs mitigate CBRN risks from their models?
The Frontier Model Forum outlined four main categories of mitigations: 1) limiting the model’s underlying capabilities; 2) shaping the model’s ability to refuse risky requests; 3) real-time interventions to block risky outputs; and 4) access restrictions.
The first—limiting model capabilities—changes how a model is trained and the resulting knowledge contained in its weights. For example, some experts have proposed removing “risky information” relating to virology from models’ training data, similar to how AI labs use cryptographic techniques to exclude child sexual abuse imagery. However, for open-weight models, such mitigations risk being undone if data is publicly available and an attacker uses it to fine-tune the model. Training on some “riskier” CBRN data may also be necessary, both to advance beneficial dual-use capabilities and to teach the model how to judge the boundary between safe and unsafe queries.
The second category—behavioral-alignment mitigations—teaches an AI model how to judge this boundary between safe and unsafe queries. Post-training methods, such as safety fine-tuning and reinforcement learning, can teach the model to identify and reject queries with harmful intent. However, over-refusing benign and beneficial queries remains a major challenge.
The third category detects and intervenes against risky model usage. “Classifiers” are fine-tuned AI models that screen users’ inputs to an LLM, and its outputs, blocking anything deemed risky in real-time. “Jailbreaks” may seek to bypass such safety measures, for instance by separating a harmful request into small, harmless-looking pieces. However, classifiers can become wise to this by evaluating outputs and inputs together. Labs can also use “linear probes”, small AI models that analyze an LLM’s internal math for signs of harmful content. One difficulty is making these mitigations robust to increasingly sophisticated and automated jailbreaks, while still fast enough to run at scale.
Ultimately, AI labs need to invest in a “defense-in-depth” approach, with multiple mitigations working in concert. This includes screening customers and sequencing access to new models, as well as analyzing user logs to detect unforeseen threats.
One challenge is that most mitigations focus on leading, closed AI models. Safety assessments suggest that open-weight models may pose a greater risk, because they have fewer safety mitigations, and those that do exist are easier to remove. One response would be to create more plug-and-play mitigations for those developing or hosting open models.
9. How could AI companies help make society more resilient to CBRN risks?
AI companies can use their technology to help prevent, detect, and respond to CBRN risks. Publicly signaling a willingness to do so may also decrease the likelihood of certain attacks, for example if attackers believe that they are more likely to fail in their goals or be detected.
To prevent threat actors from getting their hands on materials, the CBRN community relies heavily on lists of pathogens, toxins, chemicals, and machinery that it controls access to via monitoring, export controls, and customer screening. These list-based approaches are starting to fray as AI and technologies like genome editing enable actors to design toxins, chemicals, or viruses with similar functionality to those that are controlled, but with different DNA sequences or precursor ingredients.
AI could help to offset these risks in different ways. In chemistry, AI models can work backwards from a target molecule and identify alternative precursors and reaction pathways, which may highlight new materials to monitor and control. Companies that sell DNA could use AI to better analyze a customer’s background and the logic for their purchase. Researchers hope to use AI to help predict the function of a requested DNA sequence, and whether it is likely to be harmful, irrespective of whether it resembles a known pathogen or toxin—a major technical challenge.
International treaties prohibit the development of many kinds of CBRN weapons, but verifying that member states are upholding these commitments is a recurring challenge. To do so, the Organisation for the Prohibition of Chemical Weapons must review vast amounts of evidence every year—from documents exceeding 2,000 pages to handwritten scrawls. A recent OPCW working group recommended training an on-premises, air-gapped LLM to make this data queryable to inspectors and staff.
Researchers could also use AI to design materials that are less vulnerable to being used in attacks, such as alternatives for the highly radioactive materials used in medicine or industry, or new molecules to add to common chemicals to inhibit a dangerous reaction. AI could also help identify when threat actors get their hands on illicit materials, for example by processing large amounts of data from radiation monitors to detect stolen materials.
Should they occur, practitioners could also use AI to detect, characterize, and attribute CBRN attacks and incidents. For chemical weapons, this may mean analyzing a “chemical fingerprint” to determine the substance behind it. For biosecurity, it may mean helping to scale metagenomic sequencing—an approach that sequences the genomes of all microorganisms in the sample to help detect novel or less common biological outbreaks.
AI could also help society respond to a CBRN incident, by accelerating vaccines, diagnostics, and other medical countermeasures, such as better treatments for radiation exposure. Researchers have already used protein structure predictions from AlphaFold to better understand tuberculosis and malaria transmission, and to map vaccine and drug targets for threats like mpox and Nipah. The hope is that AI will ultimately provide a generalizable platform that can be quickly deployed in the event of a new outbreak.
Beyond medical products, AI could provide first responders to any incident, such as police and military, with better information, advice, and training about what they’re dealing with. For example, if a nuclear attack occurs, AI could quickly identify critical information about the detonation; generate maps and simulations to help first responders identify safe routes for evacuation and rescue efforts; and expand the use of robots.
Crucially, all these efforts will require close partnership with governments and external experts, including to manage the many dual-use risks that these applications would raise.
Acknowledgements
Thank you to the following experts who let us interview them or shared feedback on the draft, as well as those who prefer to remain anonymous. Any mistakes belong to the authors.
Victoria Langston, Paige Kunkle, Ash Otter, Adam Marsh, Jeremy Ratcliff, James Stevenson, Mor Hazan Taege, Mathias Voges, Zachary Kaplan, Akhil Jalan, Ellena Reid, Anthony Payne, Eva Lu, Jennifer Beroshi, Kevin Klyman, Stephen Johnson, and Suzy Pickering.
AI Manipulation
The notion of AIs manipulating people is a plot twist in countless sci-fi thrillers. But is “manipulative AI” really possible? If so, what might it look like?




