AstaBrief 8B: How AllenAI Trained a Small Open Model to Generate Cited Scientific Reports The Allen Institute for AI (Ai2) released AstaBrief 8B on October 2, 2026, an open-weights model built on Qwen3-8B that generates fully cited scientific reports from a research question and retrieved literature excerpts in a single forward pass, now serving as "Fast mode" in Ai2's Asta platform. Ai2 trained it with supervised fine-tuning on 47,000 filtered examples followed by DPO on roughly 6,000 judge-agreed preference pairs, reporting that filtering out low-citation-density reports improved grounding more than any other single intervention. The model averages 51.1 seconds per report versus 178.5 seconds for the prior Claude-powered pipeline, and its weights, data, and workflows are available on Hugging Face under Apache 2.0. The Allen Institute for AI Ai2 released AstaBrief 8B https://huggingface.co/allenai/AstaBrief 8B on October 2, 2026 — an open-weights model that takes a research question and a set of retrieved literature excerpts and produces a fully cited scientific report in a single forward pass. The model is now live as "Fast mode" inside Asta, Ai2's agentic platform for scientific work, and the weights, training data, and example workflows are all available on Hugging Face under an Apache 2.0 license. The release is interesting not just because of what the model does, but because of what Ai2 learned while building it: the most important lever for improving citation quality in a fine-tuned report-generation model turned out to be a simple data-filtering heuristic, not a more sophisticated training algorithm. Asta's report-generation feature previously ran entirely on a Claude-powered "Thinking mode" pipeline. That pipeline summarizes retrieved snippets, clusters them by theme, and writes the report section by section — a multi-step process that averages 178.5 seconds per report. AstaBrief replaces that pipeline for users who want speed. Given the same research question and retrieved excerpts, it writes the full report in one pass, averaging 51.1 seconds — roughly 3.5 times faster. The model is built on Qwen3-8B https://huggingface.co/Qwen/Qwen3-8B and fine-tuned specifically for the scientific report-generation task. Because the weights are open and the model is small enough to run on a single GPU, institutions can deploy it behind their own firewall — useful when research questions involve sensitive or unpublished work that cannot be sent to a third-party API. Ai2 considered reinforcement learning for training AstaBrief, citing its own earlier DR Tulu https://huggingface.co/allenai/DrTulu work as evidence that RL can improve long-form report generation in open-weights models. The team ultimately chose a simpler two-stage recipe: supervised fine-tuning SFT followed by direct preference optimization DPO . The reasoning was practical. RL training is unstable and expensive; SFT plus DPO is cheaper, easier to debug, and faster to iterate on. For a model that needs to be shipped as a production feature, those properties matter. SFT stage: Ai2 started with real user queries submitted through the ScholarQA framework that underpins Asta. After filtering for quality, relevance, and privacy — removing bot traffic, very short queries, non-English requests, and prompts containing personal information — the team had a pool of 90,000 research-focused queries. Full-report targets were generated using a mix of proprietary systems: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. Quality filtering reduced this to 47,000 usable training examples. DPO stage: Preference pairs came from a separate query subset not used during SFT. Each pair consisted of one report from the ScholarQA pipeline typically Claude 3.5 or 3.7 Sonnet against a report from o3, o4-mini, DeepSeek-V3, or DeepSeek-R1. Two judge models — GPT-4.1 and DeepSeek-R1 — independently picked a winner for each pair. Only pairs where both judges agreed were kept, producing approximately 6,000 final examples. Ai2 reports that the judges agreed with human preferences 95% of the time. DPO training ran on 8×H100 GPUs using Ai2's open-instruct framework, starting from the AstaBrief-8B-SFT checkpoint. Ai2 tested four statistics-based filters on the synthetic training reports before using them for SFT: The strongest quality gains came from filtering out reports with low citation density. Reports with large stretches of unsupported text — even if they were otherwise well-written — degraded the model's ability to produce grounded output. Removing them improved both answer precision and citation quality in the resulting model more than any other single intervention. The broader lesson Ai2 draws from this is that scientific specialization in a fine-tuned model is not primarily a function of adding more scientific text to pretraining. The composition and quality of post-training data — specifically, whether the training targets model the behavior you actually want — matters more. Ai2's primary evaluation target was SQABench-CS2, a set of 200 user-written computer science research questions. The model is scored on four metrics: rubric score necessary content coverage , answer precision paragraph relevance , citation precision citation support , and citation recall claims support . On the ScholarQA-CS2 test set, AstaBrief 8B averaged 87 across tracked metrics, compared with 83.7 for the SFT-only checkpoint and 77.3 for base Qwen3-8B. In LLM-judged pairwise comparisons against the Claude-powered Asta pipeline, AstaBrief 8B won 55% of the time on the development split and 72% on the test split. A separate human study with three scientific researchers found that two of the three preferred AstaBrief over the other systems specifically on citation accuracy — the metric that the citation density filtering was designed to improve. Ai2 notes that most training and evaluation was completed in 2025 and that it has not rerun the full evaluation against current frontier models, so direct comparisons to the latest Claude or GPT releases should be treated with caution. Among 374 Asta users who tried Fast mode, 29.1% used it on two or more days. Users generated an average of 3.67 report threads. Twenty-three percent never switched back to Thinking mode for future threads; another 18% alternated between modes. Positive feedback ran at 84.2% for Fast mode versus 85.2% for Thinking mode — a gap small enough that Ai2 considers the quality roughly comparable for most use cases. Ai2 describes AstaBrief as one experiment in a longer line of work running from ScholarQA https://github.com/allenai/ai2-scholarqa-lib and DR Tulu toward future versions of OLMo. Stated next steps include more fine-grained preference learning, stronger RAG-plus-RL approaches, multi-turn and multi-tool capabilities, additional scientific data sources, and query decomposition. The release also fits into Ai2's broader NSF OMAI initiative — a U.S. national effort to build fully open AI infrastructure for scientific discovery. AstaBrief is positioned as a practical demonstration that a small, open, specialized model can close most of the quality gap with a proprietary pipeline while running faster and at lower cost. For practitioners building retrieval-augmented generation pipelines for scientific or technical domains, the citation density finding is the most transferable takeaway: if your training data contains large stretches of unsupported claims, filtering those examples out before fine-tuning is likely to improve grounding more than switching to a more complex training algorithm. The model weights, SFT and DPO training datasets, and an example local-report workflow are available at Ai2's Hugging Face organization https://huggingface.co/allenai .