HomeScienceAI Models Uncover Long-Standing Errors
SCIENCE

AI Models Uncover Long-Standing Errors in Scientific Databases

Researchers are using AI agents to audit scientific literature, identifying calculation errors and decades-old mistakes in reference databases.

WHAT YOU NEED TO KNOW
  • Sebastian Pios found an AI model exposed errors in a 75-year-old reference database for molecular boiling points.
  • An automated audit by SAI Labs reproduced more than 80% of claims from only eight of 92 evaluated 2026 ICML papers.
  • A study led by Stanford's James Zou found objective errors in NeurIPS papers rose 55% between 2021 and 2025.

Artificial intelligence models are uncovering long-standing errors in chemistry databases and academic research papers, Nature reported.

Sebastian Pios, a theoretical chemist at Zhejiang Lab in Hangzhou, China, discovered the discrepancies while using an AI system to predict molecular boiling points. When the AI produced numbers that contradicted a 75-year-old reference database, Pios checked the original literature and found the reference entries were incorrect rather than the model. The AI system also identified a typo in an older paper and incorrect values from century-old boiling-point measurements.

Other researchers are deploying automated tools across broader scientific literature. On July 22, SAI Labs, a research-review firm based in Delaware, published an evaluation of 168 papers selected for oral presentation at the 2026 International Conference on Machine Learning (ICML). AI agents extracted central claims, downloaded accompanying resources, and attempted to rerun experiments. Among 92 papers with at least five testable claims, the agents reproduced more than two claims for 34 papers, and successfully repeated over 80% of claims for eight papers.

Odd Erik Gundersen, a computer scientist at the Norwegian University of Science and Technology in Trondheim, warned that AI fact-checking tools make mistakes and still require human oversight.

James Zou, a computer scientist at Stanford University, highlighted the scale of automated audits compared to manual checking. In a preprint study, Zou and his colleagues used an automated tool to scan papers from the NeurIPS conference for errors. The tool found that objective errors per paper rose from 3.8 in 2021 to 5.9 in 2025, marking a 55% increase.

The NeurIPS audit examined objective errors in formulae, calculations, and figures, while deliberately omitting subjective assessments of novelty and data interpretation. Federico Bianchi, a machine-learning scientist at Together AI and co-author of the study, said the team intentionally left judgments about novelty and significance to human reviewers.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →