The Allen Institute for AI (Ai2) open-sourced AstaBrief 8B on October 2, a model built specifically to turn research questions and retrieved literature snippets into cited research reports. It's a fine-tune of Alibaba's Qwen3-8B, with weights and training data posted to Hugging Face; Ai2 also published a GitHub example showing how to adapt the pipeline to a user's own PDFs.
AstaBrief now powers the "Fast mode" in Asta, Ai2's research-agent platform. The original Thinking mode, powered by Claude, took an average of 178.5 seconds per report; switching to AstaBrief cut that to 51.1 seconds, roughly 3.5x faster. Ai2 explained the motivation on its blog:
"We wanted to help scientists generate cited reports faster, with a model they could download and run themselves."
Skipping reinforcement learning, just two rounds of data
Training happened in two stages. In the supervised fine-tuning stage, Ai2 pulled 90,000 questions from real Asta user queries and used Claude 3.5/3.7 Sonnet, o3, o4-mini, and GPT-4.1 to generate target answers, ending up with about 47,000 usable samples. The preference-optimization (DPO) stage used a different set of questions: one side was reports written by the Claude-powered ScholarQA pipeline, the other side was versions from o3, o4-mini, DeepSeek-V3, or DeepSeek-R1. Two judge models, GPT-4.1 and DeepSeek-R1, compared the pairs, and only pairs where both judges agreed were kept, leaving about 6,000 pairs. Ai2 says the two judges agreed with human judgment 95% of the time.
The team skipped reinforcement learning, citing instability and cost. The pipeline was also simplified: the whole report is generated in one pass rather than section by section, cutting out the compute-heavy steps of snippet summarization and clustering.
Among several data-filtering methods tried, the one with the clearest effect was simple: stripping out synthetic samples with low citation density. Ai2's takeaway is that a direct signal — whether the output is actually grounded — beats more elaborate, layered filtering.
Benchmarked against closed-source pipelines
The main evaluation used SQABench-CS2, 200 computer-science questions, scored on four dimensions: content coverage, paragraph relevance, whether citations support the claims, and whether every claim is cited. In LLM-as-judge comparisons, AstaBrief traded wins and losses with both the Claude-powered pipeline and DR Tulu, Ai2's earlier open research model. Human evaluation was small-scale: three researchers each reviewed 14 to 15 questions, and two of them ranked AstaBrief's citation accuracy first. A second benchmark, DeepScholarBench, drew 63 questions from recent arXiv papers.
Usage data comes from 374 Asta users: 29.1% used Fast mode on more than two days, averaging 3.67 report conversations per person; 23% switched to using Fast mode exclusively afterward, and 18% toggled back and forth between the two modes.
Among comparable tools, both OpenAI's and Google's deep-research features run on closed-source large models — users can't get the weights or plug the pipeline into their own literature library. AstaBrief's trade-off is to shrink the model to 8B in exchange for being downloadable and modifiable. Rough estimate: the half-precision weights are about 16GB, small enough to load on a single 24GB-VRAM GPU, lowering the bar considerably for university labs to adopt it. The base model is Qwen3, but the training questions were mostly English-language computer science, and Ai2 hasn't published data on how well it handles Chinese-language literature.
Evaluation stops at 2025
Ai2 flagged its own limitation: these evaluations were done in 2025 against the frontier models of that time, and haven't been rerun against today's models. So the claim of "trading blows with the Claude pipeline" refers to the Claude 3.7 generation — Anthropic has since moved on to Opus 5.5.
Ai2 also pointed out a blind spot the citation metrics can't catch: a model can cite the right paper but overstate what it says — for instance, turning a finding from one specific sample into a general rule, or rewriting a descriptive result as an action recommendation. The four existing metrics can't measure that kind of drift, and Ai2 says the next step is specifically testing whether a model stays within the original claim's scope and evidential strength.
Sources: Ai2 official blog, CocoLoop; generation time, training sample size, and usage percentages are as disclosed by Ai2; VRAM requirements are an editorial estimate.