AlphaFold Database Expands to Cover 2,800+ Virus Species

The European Bioinformatics Institute (EMBL-EBI), part of the European Molecular Biology Laboratory, announced on September 24 that the AlphaFold database has added a new batch of three-dimensional structure predictions for viral protein complexes, covering more than 2,800 virus species. The data can be freely queried and downloaded through the database's newly launched Pandemic Preparedness Portal.

The release is the result of a multi-party collaboration. Besides EMBL-EBI, Google DeepMind and Nvidia, participants include the Coalition for Epidemic Preparedness Innovations (CEPI), Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow. The structure predictions were generated with DeepMind's AlphaFold2, accelerated by Nvidia's BioNeMo inference runtime, and the prediction pipeline has been open-sourced on GitHub.

What's in the New Dataset

The selection criteria prioritized virus families known to infect humans. EMBL's announcement singled out two categories: rhinoviruses, common causes of the common cold, and emerging threats such as mpox.

Unlike earlier single-protein predictions, this batch focuses on complexes — how multiple proteins bind together to function. Many steps in how a virus invades and replicates inside a cell depend on protein-to-protein binding. According to Nvidia's blog post, about 30% of the newly added protein interactions are entirely new, with no prior experimental or literature record.

"This dataset is an engine for hypothesis generation."

From March to September

Viewed on a timeline, this release is the same group's second step this year. On March 16, EMBL-EBI, DeepMind, Nvidia and Seoul National University first added complexes to the AlphaFold database: about 30 million were predicted in total, with 1.7 million high-confidence homodimers added directly to the database and roughly 18 million lower-confidence predictions released for download, covering 20 key species including humans plus the WHO's priority pathogens list.

Six months later, the list of partners grew to include pandemic-preparedness bodies like CEPI and several virology and bioinformatics institutions, and the scope narrowed from "key species" to "viruses," with the stated purpose now spelled out more explicitly: vaccine and drug development. The timing was deliberate too — EMBL's announcement notes that the data went live one day before the UN General Assembly's September 25 high-level meeting on pandemic prevention, preparedness and response.

Nvidia's blog also cited an estimate from the Center for Global Development: by 2050, the world faces roughly a 50% chance of experiencing another pandemic on the scale of COVID-19.

The Limits to Keep in Mind

EMBL was explicit in its announcement about the boundaries here. The structure predictions cannot be used to infer the consequences of gene mutations, cannot determine how a host and pathogen interact, and cannot be used to judge whether a virus's virulence or transmissibility has changed — any conclusions still need laboratory validation. Joe Grove, a professor of molecular virology at the University of Glasgow, put it bluntly:

"A protein complex structure alone doesn't tell us what happens when a virus mutates."

None of the announcements directly addressed whether making viral structure data public could pose biosecurity risks, nor did they explain whether any high-risk pathogens were deliberately excluded from the selection. EMBL's announcement likewise did not specify the data's licensing terms.

For structural biology and vaccine R&D teams in China, the dataset is available to download and use directly, with a low barrier to entry; the harder part comes later, in wet-lab validation — a step AI cannot do much to help with.

Sources: EMBL-EBI official announcement, Nvidia official blog, CocoLoop, Nature News; virus counts, the 30% new-interaction figure and the 260 million total predictions follow the figures given by EMBL and Nvidia.