- Reflection AI introduced Beam, a 501-billion-parameter sparse mixture-of-experts model with 23 billion active parameters. <sup>[1]</sup>
- The startup says Beam achieves scores comparable to Z.ai’s GLM-5.2 on advanced reasoning benchmarks while using three to four times less inference compute. <sup>[1]</sup><sup>[2]</sup>
- Reflection says its reinforcement-learning run generated more than 100 million rollouts on 10,500 Nvidia GB300 GPUs over four weeks. <sup>[1]</sup>
- Beam remains in final red-teaming and evaluation, with its weights and technical documentation scheduled for release later in October. <sup>[1]</sup><sup>[2]</sup>
- Independent performance testing is not yet available because the model weights have not been publicly released. <sup>[2]</sup>
Reflection AI on October 5 introduced Beam, a 501-billion-parameter open-weight model aimed at coding, reasoning and agentic workloads. The Nvidia-backed startup said Beam is competitive with leading Chinese open models, including Z.ai’s GLM-5.2, while requiring substantially less inference compute. [1][2]
Beam is a sparse mixture-of-experts system with 23 billion active parameters. Reflection said it pretrained the model on 23.8 trillion tokens and ran a large-scale reinforcement-learning program that generated more than 100 million rollouts on 10,500 Nvidia GB300 GPUs over four weeks. [1]
The model is not yet broadly available. Reflection said Beam is undergoing final red-teaming and evaluations, with the weights, technical report, model card and developer artifacts scheduled for release later this month under an Apache 2.0 license. [1]
The Model #
Beam is a text-only model designed for software development, reasoning and tool-using agents. Reflection describes it as a sparse mixture-of-experts architecture with 501 billion total parameters and 23 billion activated for each token. [1][2]
The company said Beam was pretrained on 23.8 trillion tokens from curated web material and proprietary licensed datasets. It also said the model has a 1 million-token context window. [1][2]
Reflection’s training program included nearly one million reinforcement-learning environments covering software engineering, terminal use, competitive coding, STEM, web search and other tasks. [1]
Performance Claims #
Reflection said Beam achieves scores comparable to Z.ai’s GLM-5.2 on advanced reasoning benchmarks and is approaching Alibaba’s Qwen 3.8-Max on coding and agentic tasks. The company said Beam’s inference efficiency is three to four times better than GLM-5.2 and more than four times better than leading Western open models. [1][2]
Those results remain company-reported. TechCrunch noted that Reflection’s performance claims had not been independently verified as of the launch because the public weights and full evaluation materials were not yet available. [2]
Reflection acknowledged that Moonshot AI’s Kimi K3 remains ahead of Beam on raw capability, while presenting efficiency as the model’s principal advantage. [1]
The Nvidia Connection #
Nvidia is an investor in Reflection, and the company used Nvidia GB300 hardware for Beam’s reported reinforcement-learning run. The launch places an Nvidia-backed U.S. model directly into a market increasingly shaped by open-weight systems from Chinese companies such as Z.ai, Alibaba and Moonshot AI. [1][2]
The startup has been developing an “AI factory” strategy aimed at enterprises and governments that want to train or customize models on their own data and computing infrastructure. Axios previously reported that Reflection had been securing additional Nvidia server capacity through agreements with Nebius and SpaceX. [3]
Availability #
A limited group of users can access an early version of Beam through a waitlist while Reflection completes red-teaming and evaluations. The company said it will publish the model weights, technical report, model card, inference code and other developer materials later in October. [1]
Reflection said it also plans distribution through hyperscalers and neocloud providers, along with integrations for open-source libraries and model-serving tools. Until the weights and evaluation materials are released, developers and outside researchers cannot independently reproduce the company’s benchmark claims. [1][2]
Companies mentioned #
Further sources #
The stories that matter, in one email. Free — unsubscribe anytime.