Google DeepMind adapts SynthID watermarking to AI-designed protein sequences
Researchers have extended Google's SynthID-text watermarking system to ProteinMPNN, a widely used protein sequence design model, creating a method called SynthIDBio-sequence. The approach embeds a detectable signal during the autoregressive sampling of amino acids, using a random seed, scoring function and tournament sampling algorithm similar to the text version. Because protein design typically uses low-temperature sampling that limits randomness, the team also developed a distortionary variant with an added filter to preserve detectability.
GoKawiil's interpretation of the reporting above, not reported fact.
Watermarking AI-generated biological sequences could give researchers and regulators a way to trace synthetic proteins back to their origin, which may become important as generative protein design tools proliferate in biotech and pharmaceutical research. The tradeoff between preserving protein function and maintaining a detectable watermark signal highlighted here suggests that applying text-style watermarking to biological data is technically harder than for language, since low-entropy sampling common in protein design limits how much signal can be hidden without altering functional properties.
- SynthIDBio-sequence adapts DeepMind's SynthID-text watermarking to the ProteinMPNN protein design model.
- The method uses tournament sampling with a secret key, random seed generator and scoring function to embed a detectable signature.
- A distortionary variant with an added filter was introduced to cope with the low-entropy sampling typical of protein design.
Source: nature.com — Stutz, 2026-09-30
Published there as: “Function-preserving watermarking of AI-generated proteins”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.