Skip to content
Tech News
← Back to articles

Drug firms’ secret data supercharge AI protein models

read original more articles
Why This Matters

This story highlights how proprietary data from pharmaceutical companies can significantly enhance AI models for predicting protein structures, which is crucial for drug discovery. It underscores the importance of data sharing and collaboration in advancing AI capabilities in biotech. The development also reflects ongoing efforts to overcome data limitations faced by existing models like AlphaFold.

Key Takeaways

Artificial-intelligence-based models of protein structure could be improved by incorporating data from pharmaceutical companies.Credit: Miyako Nakamura/Getty

For drug discovery, protein-folding models such as AlphaFold have a data problem: there isn’t enough of it in public databases. Improving the performance of these artificial-intelligence-based tools will require extra data that provide examples of how proteins and drugs interact, some scientists argue.

Protein structures – locked away by the thousands in drug company vaults – offer one promising source. Today, a consortium of pharmaceutical companies reports that using such data to train AI models of protein folding improves model performance markedly.

The group used OpenFold3 — an open-source replication of AlphaFold 3 — to develop a new model trained on more than 20,000 proprietary protein structures. The system outperformed both comparable ones trained on public data alone and those trained on the siloed datasets of individual firms. The study, described in a blog post, has not been peer-reviewed, and the model is not publicly available.

“You add all this data, and you get a pretty big bump in performance,” says Mohammed AlQuraishi, a computational biologist at Columbia University in New York City, who was part of the effort.

AlphaFold is running out of data — so drug firms are building their own version

The findings, he says, strengthen the case for generating similar publicly available datasets to supercharge protein-folding AIs. One such project, called OpenBind and supported by up to £8 million (US$10.8 million) in UK government funding, released hundreds of new protein structures last month, with thousands more in the works.

An untapped vein

The Protein Data Bank (PDB), an open repository of more than 200,000 experimentally determined protein structures, was the bedrock of AlphaFold 2’s training data. It enabled the tool to predict protein structures with startling accuracy — a breakthrough recognized with the 2024 Nobel Prize in Chemistry.

The model’s successors, including AlphaFold 3, added the ability to predict how proteins will interact with other molecules, including potential drugs. But the PDB has relatively few examples of experimentally determined structures interacting with drug-like molecules — maybe just 10,000, says Paul Mortenson, vice-president for computational chemistry and informatics at Astex Pharmaceuticals in Cambridge, UK.

... continue reading