Skip to content
Tech News
← Back to articles

Miniaturizing and modifying natural proteins with Raygun

read original more articles
Why This Matters

Raygun introduces a novel approach to protein design by representing proteins as probabilistic Gaussian distributions in a fixed-dimensional space, enabling faster and more flexible exploration of protein modifications. This advancement allows for more efficient and targeted modifications, including insertions and deletions, which are crucial for developing therapeutic proteins and expanding the capabilities of protein engineering. Its potential extension to RNA and genomic elements could further revolutionize biological sequence analysis and design.

Key Takeaways

Even though proteins evolve primarily residue by residue, natural selection does not evaluate them as such—it acts on holistic properties such as fold stability, binding affinity and catalytic activity, that emerge only at the level of the whole molecule. Raygun bridges this dichotomy by taking residue-level PLM embeddings and condensing them into a fixed-dimensional probabilistic space that allows reasoning across proteins of any size. Our representation of each protein as a Gaussian distribution, rooted in the central limit theorem rather than diffusion-based denoising, makes direct, single-shot sampling tractable, with generation roughly 100-fold faster than diffusion-based methods. More broadly, this principle of representing biological sequences of variable length in fixed-dimensional probabilistic spaces may extend to RNA families and genomic regulatory elements.

In protein design, the central challenge is staying on the manifold of feasible proteins while searching as broadly as possible across it. De novo methods4,5,48,49 start far from this manifold and attempt to reach it. Since they fix protein length at generation time, post-generation modifications are tractable only for substitutions, limiting the accessible design space. Raygun complements these approaches by starting on the manifold and exploring outwards from it, making indel-based exploration as tractable as substitutions and expanding the search radius available to the designer. The two strategies can be combined: when a de novo design requires structural optimization (shortening a loop, removing a redundant domain), Raygun enables targeted length modifications while preserving favourable features. This becomes critical when a design has the correct properties but the wrong length, such as a candidate that exceeds adeno-associated virus (AAV) packaging limits for gene therapy.

Our experimental results suggest that template-guided approaches can enable greater design novelty than de novo methods. Raygun shortened eGFP and mCherry to as little as 199 and 206 amino acids, shorter than 96% of fluorescent proteins in FPbase40, and 6 out of 8 tested variants exhibited fluorescence above background. The de novo GFP design by Hayes et al.2, using ESM-3, preserved the template length (229 amino acids), specified sequence and structure of 6 residues critical for chromophore formation, and constrained residues 58–71 for chromophore energetics. We imposed none of these constraints, yet most candidates preserved the chromophore spontaneously, and one carried a non-canonical chromophore sequence (Supplementary Note 3). Our designs and the ESM-3 GFP showed dim fluorescence and will require directed evolution to reach high fluorescence. Beyond the size reduction, the more notable finding was that a functional fluorescent protein could be obtained at all despite more than 40 coordinated indels and substitutions, given the well-known narrow fitness landscape of fluorescent proteins39. This suggests that Raygun’s edits respect the functional grammar of the protein, and that compact Raygun-generated fluorescent proteins could serve as scaffolds for new, more powerful biosensors.

For naturally evolved proteins, the relationship between sequence and function is also less straightforward than direct conservation suggests. Hie et al.50 have hypothesized that PLM-generated mutations, which follow evolutionary rules, should generally improve fitness, demonstrating this for specific antibodies. Our EGF results both support and complicate this view: using a sequence-focused approach without explicit structural optimization, we generated EGF variants with stronger EGFR binding than the native ligand, yet the key modifications occurred at peripheral positions that are presumably distant from the binding interface, indicating that function depends on global sequence context beyond direct binding-site residues.

A complementary lesson comes from the biotin ligase experiments. Moderate miniaturization preserved biotinylation activity, demonstrating Raygun-based optimization in a multi-domain setting. By contrast, extreme miniaturization—where Raygun autonomously identified and removed the DNA-binding domain to produce TurboID-11 (165 amino acids, approximately 50% reduction)—yielded a variant that was structurally stable and expressed in cells, but with weak biotinylation activity. Biotin ligation itself is a natural function that is well represented in PLM training data, but the high-strength activity of TurboID is an engineered, supra-evolutionary trait achieved through directed mutagenesis of BirA, and such supra-evolutionary functions may not be fully captured by evolutionary sequence statistics. Recovering them in aggressively miniaturized variants will likely require closing the loop with experimental feedback, such as directed evolution from Raygun candidates or function-specific assays to supplement PLM-based filters.

Future enhancements to Raygun could address scaling, multi-domain handling and directed domain manipulation. To assess the core conceptual contributions, we present results from a Raygun model trained on only 80,000 proteins from UniRef50, a data-efficient setup. Although this is already powerful, longer proteins challenged zero-shot reconstruction, suggesting limits to the expressiveness of a fixed-length representation; in such cases (for example, mTOR, with 2,549 amino acids) we recommend fine-tuning, and future work could explicitly account for multi-domain organization. Scaling to larger models and datasets continues to improve performance (Supplementary Note 1), and integrating more powerful or structure-infused PLMs such as ESM-32 and SaProt3 could further enhance Raygun. Directed manipulation of protein domains, including both removal and addition, offers another frontier whereby computational approaches can be guided by functional constraints.

Finally, Raygun, like other generative protein design tools, raises important biosafety considerations. We are signatories to the Responsible AI for Biodesign principles (https://responsiblebiodesign.ai/) and believe that computational tools must be developed and deployed with appropriate safeguards, including for concerns such as immunogenicity in therapeutic applications. Our validation results with computational functional prediction tools (such as, Pfam, ProTrek and CLEAN) suggest that they can serve as effective filters for identifying potentially concerning Raygun-generated sequences.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.