Tyler Lewy
Essay
2026

The threat of novelty

Scientists have successfully created viable viruses using generative AI. But how different is this from nature? Should we be worried?

§Read

New research from the laboratory of Brian Hie at Stanford University and the Arc Institute reported the first computationally generated virus. The biological kind. The report by King et al. titled “Generative design of bacteriophages with genome language models” (opens in a new tab) in the journal Science was immediately picked up by the press.

Each headline implied the creation of something never seen before in nature. But what does the paper actually show?

To begin, we need to describe their model. The laboratory did not pluck this virus out of thin air. Instead, they modeled it on one of the best described organisms in biology, the ΦX174 phage. Originally isolated from the Parisian sewers in 1935, this virus infects E. coli, a common bacteria species. Critically, ΦX174 is unable to infect other bacterial species let alone humans. The phage is small, containing ~5400 bases coding for 11 genes. It was the first DNA genome ever sequenced and the first to be artificially synthesized. For 90 years, ΦX174 has been a highly tractable model for biologists, generating a deep well of knowledge to pull from.

The stars of the show are the Evo 1 and Evo 2 genome language models. Similar to conventional LLMs, these models were trained on a vast corpus of genomic data with a singular goal: given a DNA string, predict the next base. King et al. specifically omitted viruses which could infect plants, animals, and fungi in an effort to assuage safety fears. Beyond this, they trained Evo 1 and Evo 2 on an additional ~15,000 Microviridae genomes, essentially the ΦX174 extended family tree to focus the model on the target. When they ran the model they generated 1000s of potential sequences, filtered them based on both viability and diversity screens to 302, successfully assembled 285 in vitro, and of these 16 were viable phage. To this virologist it seemed miraculous they retrieved anything at all. However, as King et al. began characterizing these phages, my wonder sobered up.

King et al. highlight some truly impressive findings. Evo-Φ36, one of the 16 viable phages, incorporated a protein from a completely different phage, G4. Prior attempts to make this substitution directly were all failures. Cryo-electron microscopy confirmed the smaller G4 protein fit inside the viral particle and preserved critical structural qualities. However, when testing the fitness of the phages, Evo-Φ36 was one of the worst performers. This raised an important caveat: this method can generate something novel, but that doesn’t make it an efficient virus. Indeed, even their best virus, Evo-Φ69, still was unable to match WA10, a naturally occurring phage.

The true utility of their approach lay in the combination therapy experiments. Antibacterial resistance is a dire health problem that current pharmaceutical incentives are not well positioned to solve. Phages are a viable supplement to antibiotics. King et al. showed that a cocktail of their 16 designed phages overcame resistance in two E. coli lines far easier than a combination of known ΦX174 variants. But this experiment led back to the same thought I had throughout the paper: is this any worse than nature? The mutant viruses that broke through the bacterial resistance were chimeras of input viruses reorganized through a process called recombination, a process which phages commonly do in nature. Using Evo allowed a greater diversity of input material, but ultimately the resistance arose through a natural process.

But back to the looming threat. First, the phage generated in these experiments are extremely unlikely to affect a human host. Their data show that the Evo-Φ variants were unable to infect most other E. coli strains. Even if you could get it into a human cell, the machinery is too foreign for the other phage proteins to use. They simply speak two different languages. On top of this, King et al. included two design safeguards. Eukaryotic viral sequences, those that could potentially infect plants, animals, or fungi, were deliberately excluded from the model. Additionally they filtered out potential phages that did not share ≥ 60% similarity with the ΦX174 spike protein, the element which governs which cells the virus can attack. Ultimately these safeguards were probably not the binding constraint. They demonstrate that by training their model on Microviridae genomes and providing the first few base pairs, they were unlikely to get any viable products besides ΦX174-like phages.

This is not to say that this technology could not be applied to dual-use research, research which could be used by a malicious actor to enact harm. But much of what they show suggests that nature is just as good at making viruses. None of the generated phages were as good at killing E. coli as isolates that we have already found in the field. All but one of the generated phages were ~95% similar or more to known microviruses, on par with other ΦX174-like isolates. Lastly, they chose a perfect target: small, relatively simple lifecycle, and abundant sequence data already present. Designing more complex viruses is not a trivial next step, especially if taking into account the ability of these phage to target bacteria in living creatures. But, if this technology could be used eventually to target difficult to treat bacterial infections like tuberculosis or multidrug resistant Staphylococcus aureus, then it will be well worth it.

A companion perspective piece (opens in a new tab) by Inglesby & Hanke, researchers at the Center for Health Security at Johns Hopkins University School of Public Health, uses this opportunity to argue for stricter governmental regulations. They rightly point out that King et al. did not synthesize new viruses de novo, but oversell the observed diversity, stating they “diverged substantially from any natural genome” when the measured divergence was ~5% from known sequences. Inglesby & Hanke recommend some good policies including federal monitoring of genomic synthesis. However their recommendation to not pursue generation of eukaryotic viral genomes using tools like Evo is a step too far. Safeguards are already in place to research undescribed clinical and natural viral isolates under federal biosafety level protocols. While functionally, Evo can generate “novel” sequences far more rapidly, it is unclear how this differs from finding naturally occurring viruses in nature. Every virus is “novel” until it is found.

The issue is not the generation of undescribed pathogens, but rather that the sampling bottleneck of acquiring field and clinical samples is potentially weakened. Therefore, the optimal regulatory step should target the safe use of novel sequences, not the generation of sequences themselves, permitting valuable viral research to continue.

It is important for the scientific community to discuss the potential ramifications of novel technologies like those described in the paper. But we must update our priors according to the evidence, not to the headlines. This tool allows the exploration of more ideas more rapidly, but the viruses that were successful were nearly identical to natural viruses we already know. What King et al. accomplished was more akin to speeding up evolution in a computer. The speculation of using Evo to design a truly novel virus, one that we have no basis for already, is unfounded. Safeguards are in place already to deal with novel naturally occurring viruses and these should be treated no differently. Ultimately, this paper should leave one with a splash of optimism. There are many dangers of misused or unaligned AI, but Evo does not appear to elevate that risk.

The two papers this piece is about

King SH, Driscoll CL, Li DB, Guo D, Merchant AT, Brixi G, Wilkinson ME, Hie BL. Generative design of bacteriophages with genome language models. Science 2026;393(6811):eaec2657. doi:10.1126/science.aec2657 (opens in a new tab)

Inglesby TV, Hanke MS. AI-designed viral genomes. Science 2026;393(6811):563–564. doi:10.1126/science.aej8512 (opens in a new tab)

Written August 2026. These are my own views and do not represent the NIH, NIAID, or Rocky Mountain Laboratories.

← All writing