Ai2 says its system can train larger sparse models with less overhead. Its trillion-parameter result measures system performance, not a completed model.
By [RuntimeWire Staff](https://runtimewire.com/author/runtimewire-staff)
· Published
Primary source: [Hugging Face Newsroom](https://huggingface.co/blog/allenai/olmocore3)
Why it matters #
Ai2's release makes MoE training infrastructure part of its open-model strategy, but the trillion-parameter headline is a system benchmark, not a trained model. The distinction sets a useful bar for judging both its technical claims and the access the code actually provides.
The Allen Institute for AI (Ai2) released Olmo-core 3 on October 1st, opening a redesigned training system for mixture-of-experts models and reporting benchmarks up to 1.2 trillion total parameters. The result is a systems test, not evidence that Ai2 has trained a trillion-parameter model to useful quality.
Ai2 was founded in 2014 by Microsoft co-founder Paul Allen as a Seattle-based nonprofit research institute. Olmo-core 3 extends Allen's founding commitment to openly shared AI research. Ai2 says the new framework will underpin the next generation of Olmo, which it plans to build as a mixture-of-experts model. In a Hugging Face announcement, Ai2 presents the framework, technical report and interactive walkthrough alongside the public code.
The release targets a practical constraint in sparse models. A mixture-of-experts, or MoE, routes each token to a subset of a much larger pool of specialized model components. That can keep the amount of computation per token lower than a dense model of similar total capacity. But training still requires keeping model weights and optimizer state in memory, while moving tokens to the right experts across GPUs adds communication overhead. Those costs can eat into the savings from activating only part of the model.
Ai2 says its new stack addresses that overhead by changing how the system distributes work. Its earlier MoE implementation used fully sharded data parallelism, gathering and resharing weights for each small batch. Olmo-core 3 instead uses a distributed data parallel approach that keeps experts on GPUs and routes data to them. Ai2 combines that design with expert and pipeline parallelism, a distributed optimizer, GPU-resident routing and grouped matrix computations.
The benchmark is the claim
In one comparison, Ai2 increased the expert pool from eight to 128 while selecting four experts for each token. Ai2 says active parameters stayed near 3.2 billion as total capacity grew from 4.6 billion to 47 billion, with training throughput falling by less than 5%.
A separate preliminary test compared the new system with Ai2's earlier implementation on eight NVIDIA B300 GPUs. For a 47-billion-parameter MoE, Ai2 reports 52,000 tokens per second per GPU, against 19,400 with the previous stack, or about 2.7 times the throughput. A four-GPU test of MXFP8, a lower-precision number format, showed about 21% higher throughput than the BF16 baseline and peak active memory declining from 103 GiB to 95 GiB. These are Ai2's own results; the announcement does not establish independent reproduction or disclose the benchmarks' compute cost, energy use or duration.
The largest configuration needs careful reading. Ai2 reports testing a 1.2-trillion-parameter setup across 512 B300 GPUs, with 58.36 billion parameters active per token and a peak observed throughput of 858 TFLOP/s per GPU. Ai2 says it used random routing to measure system performance, rather than model quality. A separate short-capacity test reached 2.38 trillion parameters using DeepEP v2, which the institute explicitly describes as a configuration test rather than a full training run. Neither figure should be read as a trained, evaluated model available to use.
Scaling expert count while keeping active computation roughly stable tests whether routing and communication can keep pace as a model's total capacity expands. In Ai2's framing, infrastructure is valuable because a model's architecture is only part of the cost: data movement, memory and coordination determine whether sparse computation pays off in practice.
An open stack for the next Olmo
Olmo-core 3 follows earlier Ai2 work on sparse models, including OlmoE, which used 64 routed experts. The institute says the more recent Olmo 3 model used a dense architecture; its next Olmo generation is planned as an MoE. The new training system therefore serves both as a public research tool and as groundwork for Ai2's own next model.
The project extends Allen's founding premise into a layer that often remains harder to inspect than model weights. Ai2 argues that open model development is more useful when researchers can inspect the training infrastructure and decisions behind a model alongside its final parameters. The Olmo-core repository makes the framework available to researchers and developers, and the interactive walkthrough illustrates how data, experts and model layers are split across GPUs.
Code alone does not remove the hardware barrier for smaller labs. Ai2's largest reported benchmark used 512 high-end GPUs, and the announcement provides no cost figure that would establish how accessible a comparable run is in practice. The nearer-term contribution is a more inspectable implementation and a set of reported performance results that other teams can scrutinize and attempt to reproduce. For Ai2, the immediate payoff is also internal: it has released the infrastructure it says will support its next Olmo model before that model arrives.