Google DeepMind released SynthID Bio on September 30, a system that writes a watermark directly into the amino acid sequences of AI-designed proteins. A companion methods paper appeared simultaneously in Nature, and DeepMind has opened the code, in vitro experimental data, and model weights to the research community.
The tool answers a narrow question: given a protein sequence, can you verify it came from a trusted AI design pipeline? DeepMind frames it strictly as a provenance check — it does not and cannot judge whether a protein is dangerous.
How the watermark gets written in
SynthID Bio sits on top of ProteinMPNN, a widely used protein design tool. ProteinMPNN works by placing amino acids one position at a time along a protein backbone, and at many positions several chemically similar candidates are interchangeable — leucine, isoleucine, and valine, for instance, can often substitute for one another without breaking the structure.
SynthID Bio intervenes at exactly that step. Based on a watermark key and the amino acids already chosen earlier in the sequence, it proposes a candidate. The design pipeline only accepts the suggestion when doing so doesn't compromise protein function. Any single position looks unremarkable on its own, but across the full sequence, amino acids that match the key's preference show up more often than chance would predict.
Detection is likewise statistical. Whoever holds the key scans the full sequence, counts how many positions align with what the key would have suggested, and compares that count against a threshold. As Ars Technica's coverage notes, where that threshold is set directly shapes the trade-off between false positives and false negatives — the output is a probabilistic call, not a yes/no verdict.
What the experiments showed
The team used the watermarked pipeline to design binding proteins against three targets: the vascular endothelial growth factor VEGF-A, the receptor-binding domain of the SARS-CoV-2 spike protein, and the immune checkpoint protein PD-L1. In vitro testing showed the watermarked designs matched unwatermarked ones on hit rate and binding affinity, with no drop in sequence diversity. DeepMind says these are the first watermarked, biologically functional protein binders.
They applied the same approach to AlphaFold 3's structure-prediction output as well, where detection came close to perfect and held up against numerical noise and small coordinate shifts.
James Diggans, Twist Bioscience's vice president for policy and biosecurity, offered this assessment:
"AI is expanding what scientists can design, and DNA synthesis companies have an important role in helping that innovation scale responsibly."
Next to text watermarking
SynthID started on images before expanding to text, audio, and video; the text watermark was open-sourced in 2024. This site has previously covered OpenAI building Google's watermark into ChatGPT, and Anthropic embedding an invisible watermark in Claude's output. All of these schemes share the same weak point: paraphrasing. Rewrite a passage in different words and the statistical signal gets diluted.
Proteins have an equivalent of paraphrasing. The reporting lists several: very short proteins don't have enough positions to hold a reliable watermark; appending a natural sequence after a watermarked one — fusing on a fluorescent protein, for example — washes out the signal; and many protein design tools aren't built on ProteinMPNN or don't place amino acids one at a time at all, so the watermark has nowhere to attach. DeepMind itself acknowledges that robustness against deliberate tampering remains an open problem.
There's also the matter of key management. The whole system's trustworthiness hinges on who holds the key, how it's distributed, and what happens if it leaks — and there's no established industry arrangement for any of that yet.
Where this fits in the biosecurity chain
Today's biosecurity defenses mostly sit at the DNA synthesis stage: synthesis companies screen ordered sequences against databases of known dangerous sequences. A novel AI-designed protein may not resemble anything in those databases, so screening can miss it. A watermark offers a different clue — telling a synthesis company that a sequence came from a registered design pipeline.
It fills a provenance gap, not a harm-assessment one. A sequence without a watermark may simply have come from a different tool; a sequence with one isn't automatically safe. DeepMind's announcement is explicit that no single biosecurity measure can solve this problem on its own. For now, DeepMind is reaching out to partners through synthidbio@google.com, and how many synthesis companies and design tools end up adopting it remains to be seen.
Sources: Google DeepMind official blog, Nature methods paper, CocoLoop, Ars Technica. Target proteins, hit rates, and affinity comparisons follow the paper and official blog; limitations draw on Ars Technica's analysis of the method.