Google’s DeepMind team adapted its SynthID watermarking technology to protein sequences by integrating a variant called SynthIDBio into the ProteinMPNN design pipeline. ProteinMPNN first generates a suitable backbone, then selects side‑chain amino acids sequentially; SynthIDBio proposes a candidate residue based on a secret key and the identities of previously chosen acids, and the proposal is accepted only if it preserves the protein’s predicted function. This allows watermark‑compatible residues to be inserted wherever functionally equivalent alternatives exist (e.g., leucine/isoleucine/valine), distributing the signal across the entire chain. Detection requires scanning the full sequence with the key and measuring how often the SynthIDBio‑suggested residues appear, treating the watermark as a statistical signature rather than a binary marker. Experiments showed that watermarked proteins designed to bind specific natural targets retained binding activity, indicating functional preservation. The method works reliably for proteins long enough to accumulate sufficient watermark residues; very short proteins may lack detectable signal, and fusion with non‑watermarked domains can dilute it. Security depends on protecting the cryptographic keys used for watermark generation and distribution, and the false‑positive/negative rates hinge on the chosen detection threshold. While not all protein‑design tools currently support the approach, the study demonstrates that AI‑generated proteins can be covertly marked without compromising their biochemical properties.
Read original
arstechnica/ai