HomeSciencePharma Vault Data Boosts AI Protein Pr
SCIENCE

Pharma Vault Data Boosts AI Protein Predictions in New Study

A drugmaker consortium fine-tuned an open-source protein model on 20,167 private structures, sharply improving predictions over public baselines.

WHAT YOU NEED TO KNOW
  • The AI Structural Biology Network fine-tuned OpenFold3 on 20,167 proprietary protein-ligand structures from five pharmaceutical firms.
  • The pooled model accurately predicted more than half of 1,056 test structures, compared to one-third for public OpenFold3 and about 40% for Boltz-2.
  • While the Protein Data Bank contains over 200,000 structures, Astex estimates only around 10,000 involve drug-like molecules.

A pharmaceutical industry consortium trained an artificial-intelligence protein model on more than 20,000 proprietary molecular structures, significantly boosting prediction accuracy over tools trained solely on public data, Nature reported.

Drugmakers AbbVie and Astex Pharmaceuticals formed the AI Structural Biology (AISB) Network last year with several other companies to test pooled internal data. The group fine-tuned OpenFold3, an open-source replication of AlphaFold 3, using 20,167 structures showing proteins bound to potential drug molecules, known as ligands. Five companies supplied proprietary data gathered through X-ray crystallography and cryo-electron microscopy while keeping their underlying datasets private.

Public databases contain comparatively few examples of protein-drug interactions. The Protein Data Bank holds more than 200,000 experimentally determined protein structures—the foundation that earned AlphaFold 2's creators the 2024 Nobel Prize in Chemistry—yet features only about 10,000 structures bound to drug-like molecules, according to Paul Mortenson, vice-president for computational chemistry and informatics at Astex.

Test results showed marked accuracy gains when the model evaluated 1,056 reserved protein-ligand structures. The AISB system predicted more than half of those test structures with high accuracy, compared to one-third for the public version of OpenFold3 and roughly 40% for rival open-source model Boltz-2. John Karanicolas, head of computational drug discovery at AbbVie, noted that the model also outperformed systems trained only on isolated data from single companies.

Columbia University computational biologist Mohammed AlQuraishi said the results support efforts to build comparable public datasets for protein folding. One such initiative, the OpenBind project backed by up to £8 million ($10.8 million) in UK government funding, released hundreds of new protein structures last month and is preparing thousands more. The AISB team detailed its findings in a blog post, with plans to submit a paper to a peer-reviewed journal, though the model itself is not publicly available.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →