Artificial intelligence models are uncovering long-standing errors in chemistry databases and academic research papers, Nature reported.
Sebastian Pios, a theoretical chemist at Zhejiang Lab in Hangzhou, China, discovered the discrepancies while using an AI system to predict molecular boiling points. When the AI produced numbers that contradicted a 75-year-old reference database, Pios checked the original literature and found the reference entries were incorrect rather than the model. The AI system also identified a typo in an older paper and incorrect values from century-old boiling-point measurements.
Other researchers are deploying automated tools across broader scientific literature. On July 22, SAI Labs, a research-review firm based in Delaware, published an evaluation of 168 papers selected for oral presentation at the 2026 International Conference on Machine Learning (ICML). AI agents extracted central claims, downloaded accompanying resources, and attempted to rerun experiments. Among 92 papers with at least five testable claims, the agents reproduced more than two claims for 34 papers, and successfully repeated over 80% of claims for eight papers.
Odd Erik Gundersen, a computer scientist at the Norwegian University of Science and Technology in Trondheim, warned that AI fact-checking tools make mistakes and still require human oversight.
James Zou, a computer scientist at Stanford University, highlighted the scale of automated audits compared to manual checking. In a preprint study, Zou and his colleagues used an automated tool to scan papers from the NeurIPS conference for errors. The tool found that objective errors per paper rose from 3.8 in 2021 to 5.9 in 2025, marking a 55% increase.
The NeurIPS audit examined objective errors in formulae, calculations, and figures, while deliberately omitting subjective assessments of novelty and data interpretation. Federico Bianchi, a machine-learning scientist at Together AI and co-author of the study, said the team intentionally left judgments about novelty and significance to human reviewers.
