Ai2 has open-sourced AstaBrief 8B, a language model trained to generate cited research reports from scientific literature, according to an announcement published on Hugging Face. The model powers Fast mode in Asta, the institute's agentic research system, where it runs alongside a slower Thinking mode powered by Claude.
Across the Asta pipeline, Fast mode averages 51.1 seconds to produce a finished report, compared with 178.5 seconds for Thinking mode. Ai2 achieved the speed increase by training the model to write the entire report in a single pass from retrieved literature excerpts, removing the intermediate snippet clustering and summarisation steps used in Thinking mode.
Training and Filtering
Ai2 developed AstaBrief from the base model Qwen3-8B using supervised fine-tuning and direct preference optimization rather than reinforcement learning. The team extracted 90,000 research queries from Asta user logs and filtered them for privacy, language, and relevance. From those queries, Ai2 generated 47,000 fine-tuning examples through its ScholarQA pipeline using proprietary models including Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1.
For the preference optimization stage, Ai2 compiled roughly 6,000 paired report examples. Competing reports were generated with o3, o4-mini, DeepSeek-V3, and DeepSeek-R1. Two judge models, GPT-4.1 and DeepSeek-R1, evaluated each pair, and Ai2 retained only examples where both judges agreed. The team tested four statistical filters to remove weak training examples: output-to-input token ratio, citation relevance, citation density, and citation diversity. Filtering out synthetic reports with low citation density produced the strongest performance gains.
Evaluation and Deployment
Ai2 tested the model against SQABench-CS2, a benchmark of 200 computer science research questions, measuring rubric score, answer precision, citation precision, and citation recall. In a separate 14-question study with three human researchers, two preferred AstaBrief over competing systems on citation accuracy metrics, while DR Tulu won on overall preference. Ai2 completed most training and evaluation work in 2025.
Early platform data from 374 Asta users showed that 29.1 percent used Fast mode across two or more days, averaging 3.67 report threads per user. Twenty-three percent of users who tried Fast mode stayed on it without returning to Thinking mode, and Fast mode received an 84.2 percent positive feedback rating compared with 85.2 percent for Thinking mode.
Ai2 released an example workflow alongside the model weights so research institutions can run AstaBrief locally on proprietary PDF collections behind institutional firewalls. The organisation is also developing future scientific models under the NSF OMAI initiative and planning updates to its Olmo architecture.
