GPN-Star is a phylogeny-informed genomic language modelling framework that consistently outperforms existing methods for predicting functional constraint and deleterious variants across the human genome. Notably, GPN-Star achieved particularly strong predictive performance for distal enhancer variants, suggesting that evolutionary constraint information derived from cross-species comparisons can provide valuable signals for modelling this historically challenging class of regulatory element60. Beyond predictive accuracy, it learns biologically meaningful representations spanning functional elements from enhancers to TFBS and their co-evolutionary dependencies—all without any supervision. Classical WGA-based models, such as PhastCons18 and PhyloP19, have been indispensable tools to biologists and clinicians for over two decades. Our results indicate that GPN-Star complements and extends these tools, offering improved performance across diverse species and alignments.
A central insight from our study is that the evolutionary timescale represented in the training data strongly influences the constraint learned by gLMs (Figs. 1c and 2i), corroborating and extending previous observations for classical phylogenetic methods7. Models trained on deeper timescales better capture coding and other highly constrained regions, whereas shallower timescales better inform rapidly evolving regulatory elements, including those constrained on primate-specific regulation8. The evolutionary timescale should therefore be considered an important design consideration for future gLMs.
Like classical conservation scores, GPN-Star predictions quantify evolutionary constraint at specific timescales rather than general variant pathogenicity. Appropriate timescales should therefore be chosen according to the biological question. For integrating or selecting among models in predictive or discovery workflows, we recommend data-driven approaches that learn appropriate weights and transformations from task-specific data, as we demonstrated with DeepRVAT in our RVAT analysis.
GPN-Star achieves state-of-the-art performance in functional constraint prediction with substantially smaller model and context sizes compared with existing single-sequence gLMs. This effectiveness probably stems from the explicit homology information provided by the alignment, which is not readily available to single-sequence models. Although we observed modest performance gains from increased model sizes, improvements from expanding context sizes were relatively small (Supplementary Fig. 18). A plausible explanation is that WGAs are constructed relative to a reference species and often consist of small, highly fragmented synteny blocks. As a result, when context size increases, nucleotides in a window may not be contiguous in the actual genomes of non-reference species, introducing misleading context information to the model. Exploring how to leverage larger context sizes with WGA data is a promising direction for future research. Advancement in model architecture design or alignment data structure is probably necessary.
A practical limitation is that GPN-Star requires WGAs during inference. To facilitate its use, we have released genome-wide predictions through public repositories. Also, the reliance on a fixed alignment format makes the model less suitable for analysing indels and structural variants. On the other hand, single-sequence gLMs offer greater flexibility and applicability to a broad range of tasks and can naturally leverage large context and process indels and structural variants. We envision that future work can bridge these paradigms potentially through flexible homology retrieval mechanisms, allowing models to dynamically incorporate evolutionary information without relying on rigid alignment structures.
Our results on pathogenicity prediction and complex trait heritability constitute a systematic evaluation of GPN-Star in the context of present knowledge of human genetic variation. However, the true value of the model depends on its capacity to enable biological discovery. We expect GPN-Star to hold great potential in human genetics applications by advancing our understanding of causal variants and human traits. Initial experiments demonstrate improved RVAT, and we expect even greater gains for non-coding analyses, in which variant prioritization remains challenging. GPN-Star may also improve functionally informed fine-mapping, polygenic risk prediction and phenotype prediction.
Beyond identifying evolutionary constraints and deleterious variants, an important future direction is to study human-specific adaptations. The present GPN-Star models focus on capturing evolutionary signals across species, which lack the resolution for studying selection in human populations. We anticipate this to be an important direction for future work and would require learning from even more localized evolutionary data, such as human population genomes and archaic human genomes.
It is remarkable how much can be learned from unlabelled DNA sequences alone. Evolutionary and functional genomics provide complementary views of genetic variation. Although evolutionary modelling outperformed present sequence-to-function approaches in our pathogenicity and heritability analyses, functional genomics remains essential for studying molecular phenotypes and tissue-specific regulation. Integrating these complementary sources of information into multimodal gLMs represents a promising direction.
Large-scale efforts to sequence and align genomes across the tree of life are accelerating21,61. GPN-Star is well positioned to take advantage of this growth, as it scales readily to new alignments and may gain further capacity from increasingly diverse training data. In this age of rapidly expanding genomic data, we expect GPN-Star to be a valuable tool for leveraging these data to advance our understanding of genetic variation.