
Watermarking synthetic biology with SynthID Bio
A new technique embeds imperceptible digital signatures directly into biological code and 3D structures. The approach aims to strengthen biosecurity and prevent the pollution of public scientific databases without harming biological function.
Published by Jin · 2 min read · 1 OCT 2026
- VEGF-A, SARS-CoV-2 spike protein RBD, PD-L1
Generative artificial intelligence is transforming biology by predicting protein structures, designing new proteins, and developing bacteriophages, which are viruses that infect bacteria. While these tools accelerate research, they also introduce risks. Novel AI designs can bypass traditional DNA synthesis screening, and mislabeled 3D structures can mislead public databases.
Introducing SynthID Bio
SynthID Bio is a family of watermarking methods built to embed a verifiable signal directly into biological sequences and structures. For sequences, the technique subtly guides the choice of amino acids. For predicted three-dimensional structures, it adjusts atomic coordinates. Laboratory testing shows these modifications preserve normal biological function.
Researchers tested the approach on protein binders, which are molecules built to latch onto other proteins. Using the binder design method AlphaProteo alongside a watermarked version of ProteinMPNN for protein sequence generation, the team evaluated designs across three target proteins:
- VEGF-A
- SARS-CoV-2 spike protein RBD
- PD-L1
Wet-lab testing showed the watermarked designs matched the hit rate, binding affinity, and natural sequence diversity of unwatermarked versions. For protein folding, the method fine-tunes a small part of AlphaFold 3’s diffusion network, embedding the watermark directly into model weights to ensure predicted coordinates carry a detectable signature.
Strengthening biosecurity
Biosecurity relies on layered defenses where independent safety measures cover potential gaps. SynthID Bio provides an automated verification signal for DNA synthesis screening, helping providers confirm that unfamiliar sequences originated from trusted models with built-in safeguards. This helps screeners focus resources on sequences that warrant closer review.
The approach can also help maintain the integrity of public scientific databases such as the Protein Data Bank, UniProt, and GenBank by flagging synthetic entries. Future work includes improving robustness against tampering, combining the approach with provenance metadata, and expanding watermarking to more complex biological objects like the genome of a bacteriophage designed by the genomic model Evo 2.
Source — Original announcement ↗
Worth a read?
Comments · 0