Skip to content
Tech News
← Back to articles

AI-redesigned starting points and outcomes enhance protein evolution

read original more articles
Why This Matters

This article highlights how AI-driven methods, such as ProteinMPNN, are revolutionizing protein engineering by enabling more precise and efficient design of protein sequences. These advancements can accelerate drug discovery, improve therapeutic development, and enhance our understanding of protein functions, ultimately benefiting both the tech industry and consumers. The integration of AI in protein evolution marks a significant step toward more innovative and personalized biomedical solutions.

Key Takeaways

General methods

Antibiotics from Gold Biotechnology were prepared in 1,000× stock solutions in water unless otherwise indicated: chloramphenicol (25 mg ml–1 in 70% ethanol), carbenicillin (50 mg ml−1), spectinomycin (50 mg ml−1), tetracycline (10 mg ml−1 in 50% ethanol) and kanamycin (25 mg ml–1). PCR amplifications were carried out using either Phusion U Green Multiplex PCR Master Mix (Thermo Fisher Scientific) or Q5 Hot Start High-Fidelity 2× Master Mix (New England BioLabs). DNA oligonucleotides were synthesized by Integrated DNA Technologies. Plasmids encoding synthetic, human codon-optimized ATXN2 cDNA sequences were obtained from GenScript. Plasmid constructs were assembled via Golden Gate or Gibson cloning protocols following previously described methods21,37. Cloning was performed in chemically competent E. coli Mach1 cells (Thermo Fisher Scientific) or NEB 5α cells (New England Biolabs). Plasmids from single colonies were amplified using the Illustra Templiphi 100 Amplification Kit (Cytiva) before sequencing via Sanger (Quintara Biosciences) or Nanopore (Quintara Biosciences or Plasmidsaurus). For bacterial experiments, plasmid DNA was purified using the QIAprep Spin Miniprep Kit (Qiagen), whereas plasmids for use in mammalian cells was isolated with the Plasmid Plus Midiprep Kit (Qiagen). All plasmids were eluted in nuclease-free water and quantified using a NanoDrop ONE UV-Vis spectrophotometer (Thermo Fisher Scientific). The list of plasmids and selection phages used in this study is available in Supplementary Table 1.

ProteinMPNN sequence design

Code and procedure for sequence design are available on GitHub (https://github.com/Nicholas-Krasnow/sequence-design-guide). In summary, AF2-predicted structures of each redesigned BoNT protease were used as input for ProteinMPNN, and residues outside specified thresholds for distance to substrate and evolutionary conservation were constrained from redesign. AF2 predictions were confirmed to agree with existing experimental structures while offering a model of regions unresolved by crystallography54,55,56. Sequences were generated in groups according to the constraint cut-offs used. Substrate distance constraints were 14 Å and 18 Å for BoNT/E, 10 Å, 14 Å and 18 Å for BoNT/F, and 18 Å for BoNT/X as measured in PyRosetta57. Conservation of each residue was determined as the residue frequency in a multiple sequence alignment of homologues identified from a database search of Uniref50 with the starting BoNT protease as the search query. Conservation constraints of 30% and 60% were used for all proteases; each residue that was at least as conserved as the cut-off and was the plurality residue in the alignment position was constrained from design. For BoNT/X, an additional criterion was applied to constrain residues predicted to lie within 14 Å of the belt domain in the holotoxin based on alignment to the BoNT/A holotoxin structure (Protein Data Bank ID 3BTA)58. During sequence generation, two sampling temperatures of 0.1 and 0.3 were tested. Cysteine was excluded from design to avoid oxidation. Structures of output sequences were predicted in AF2 and evaluated by calculating the predicted local distance difference test score and root mean square deviation to the input structure measured in PyMOL.

PROSS sequence design

PROSS variants of BoNT/E were generated using the PROSS webserver (https://pross.weizmann.ac.il/step/pross-terms/). The same AF2 input structure and the same distance constraints of 14 Å and 18 Å used for BoNT/E ProteinMPNN were used for PROSS. Because PROSS directly uses a multiple sequence alignment to sample putative stabilizing mutations, the conservation constraints used for ProteinMPNN were not applied. Folding energy was calculated with the ref2015 energy function. The webserver was run until 44 non-duplicate sequences were generated for each distance constraint group.

Small-scale protein expression and purification

Starter cultures were grown overnight from single colonies and used to inoculate 5 ml of medium supplemented with 50 µg ml−1 kanamycin per well in a 24-well plate (50 μl inoculum per well). Cultures were incubated at 37 °C with shaking until reaching mid-log phase, then were subjected to a 1-h cold shock on ice. Protein expression was induced with 1 mM IPTG (Gold Biotechnology) and cultures were incubated overnight at 16 °C.

Overnight cultures were harvested by centrifugation. Cell pellets were resuspended in 400 μl of B-PER reagent (Thermo Fisher Scientific) supplemented with 0.8 μl each of lysozyme and DNaseI and 16 μl of cOmplete protease inhibitor tablet solution prepared in 2 ml nuclease-free water (Roche). Lysis proceeded by incubation at room temperature for 15 min without freeze–thaw cycles. Lysates were transferred to microcentrifuge tubes and clarified by centrifugation at 20,000g for 5 min at 4 °C. To prepare affinity resin plates for purification (HisPur Cobalt Spin Plates, Thermo Fisher Scientific), storage solution was removed from the resin by centrifugation at 500g for 3 min. Wells were washed once with 400 μl ultrapure water and centrifuged again at 500g for 3 min. Plates were equilibrated by two washes with 400 μl wash buffer (20 mM HEPES and 200 mM NaCl, pH 7.3), each followed by centrifugation. Plates were chilled on ice.

Following lysis, 300 µl of each supernatant (cleared lysate) was transferred to the pre-chilled 96-well cobalt resin plate and incubated for 15 min on ice. Plates were centrifuged for 3 min at 500g at 4 °C, and flow-through was collected. The resin was washed four times with 400 μl of wash buffer containing 10 mM imidazole, each wash followed by centrifugation at 500g for 3 min at 4 °C. Bound proteins were eluted by incubating the resin with 200 μl of elution buffer (wash buffer + 250 mM imidazole, pH 7.3) for 1 min, followed by centrifugation at 500g for 3 min at 4 °C. Eluates were sealed and stored at 4 °C.

... continue reading