Pharma consortium trains AI protein model on 20,000 proprietary structures

This digest was compiled by AI from multiple sources — links to the originals are below.
A consortium of pharmaceutical companies trained an AI protein-folding model on more than 20,000 proprietary protein structures, markedly improving its performance. The model outperformed versions trained on public data alone or on individual firms' siloed datasets. The study was described in a blog post and has not been peer-reviewed.
Key Facts
- The consortium used OpenFold3, an open-source replication of AlphaFold 3, to train the new model on more than 20,000 proprietary protein structures.
- The model outperformed comparable systems trained on public data alone and those trained on individual firms' siloed datasets.
- The Protein Data Bank contains relatively few examples of experimentally determined structures interacting with drug-like molecules, perhaps just 10,000, according to Paul Mortenson of Astex Pharmaceuticals.
- The OpenBind project, supported by up to £8 million (US$10.8 million) in UK government funding, released hundreds of new protein structures last month.
Training Data Gap
The Protein Data Bank (PDB), an open repository of more than 200,000 experimentally determined protein structures, was the bedrock of AlphaFold 2's training data. AlphaFold 2's accuracy earned its developers the 2024 Nobel Prize in Chemistry. AlphaFold 3 added the ability to predict how proteins interact with other molecules, including potential drugs. The PDB has relatively few examples of experimentally determined structures interacting with drug-like molecules, maybe just 10,000, says Paul Mortenson, vice-president for computational chemistry and informatics at Astex Pharmaceuticals in Cambridge, UK.
Model Performance
The consortium used OpenFold3, an open-source replication of AlphaFold 3, to develop a new model trained on more than 20,000 proprietary protein structures. The system outperformed both comparable ones trained on public data alone and those trained on the siloed datasets of individual firms. The study, described in a blog post, has not been peer-reviewed, and the model is not publicly available. "You add all this data, and you get a pretty big bump in performance," says Mohammed AlQuraishi, a computational biologist at Columbia University in New York City, who was part of the effort.
Open Data Initiatives
The findings strengthen the case for generating similar publicly available datasets to supercharge protein-folding AIs, according to AlQuraishi. One such project, called OpenBind and supported by up to £8 million (US$10.8 million) in UK government funding, released hundreds of new protein structures last month, with thousands more in the works. Research has suggested that the accuracy of AlphaFold 3 and other co-folding models drops off a cliff when challenged to predict interactions between molecules highly dissimilar to those on which they were trained.