Skip to content
Tech News
← Back to articles

The Millennium Problems for Biology

read original more articles
Why This Matters

This proposal outlines an ambitious biological 'moonshot' challenge: reverse-translating proteins back into nucleic acid code, which would be revolutionary for proteomics, biotechnology, and synthetic biology by enabling sequencing of proteins the way we sequence DNA. Solving this could transform drug discovery, diagnostics, and our understanding of the proteome, areas where current technology lags far behind genomic sequencing capabilities.

Key Takeaways

Create an enzyme that can “reverse translate” an arbitrary peptide sequence into RNA or DNA.

Specifically, create a purified protein catalyst or fixed protein complex that processively reads an untagged polypeptide and synthesizes a covalent nucleic acid strand encoding its residue sequence under a preregistered codon convention, without a nucleic-acid template, preattached sequence barcode, residue-specific operator cycle, or database lookup. The resulting nucleic acid strand must be compatible with ordinary polymerases, ligases, and other similar enzymes, i.e., if nucleic acids other than RNA or DNA are used, they must be compatible with downstream amplification or sequencing reactions. For the challenge to be considered complete, at least 100 random peptide sequences of at least 50 amino acids each must be preregistered, synthesized, and pooled. It must then be shown that the sequences of these peptides can be inferred, without reference to a dictionary, by reverse translation and sequencing with at least 90% sequence accuracy. Moreover, the average read length must be at least 25 residues, and the average read quality score should be at least Q10.