October 5, 20261 min read
We are introducing Beam, Reflection’s first open-weight model. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.
Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.
Together, these efforts produced competitive open-weight performance with frontier inference compute efficiency.
Beam is undergoing final red-teaming and evaluations. You can sign up here for early access to the model. We will release the weights, technical report, model card, and developer artifacts later this month.
Model Capability #
We trained Beam with a particular focus on coding and agentic performance. Beam advances the Western open-weight frontier and is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks. Where frontier open models like Kimi K3 remain ahead on raw capability, Beam's advantage is efficiency at inference time.
The below figure shows Beam's performance across a range of coding, agentic, reasoning, and STEM benchmarks. NR denotes scores that have not been reported.
Beam pairs coding and agentic capabilities with highly efficient reasoning. On advanced reasoning benchmarks, it achieves scores comparable to GLM-5.2 while using 3–4× less inference compute. Efficiency gains are even more pronounced when comparing to models in the 2T+ parameter family like Qwen 3.8-Max, which require significantly more inference compute per token.
These results translate into more intelligence per token, delivering strong model capabilities at lower cost, making Beam a powerful workhorse model for enterprise coding and agentic workloads.
High-Compute Reinforcement Learning
We made high-compute reinforcement learning a central scaling axis for Beam, investing in RL science, data, and infrastructure to turn more compute into stronger capabilities. Scaling RL enables more extensive exploration of problem-solving strategies, while longer rollouts support multi-step reasoning, tool use, and adaptation to environment feedback.
To scale reinforcement learning, we deployed 10.5K NVIDIA GB300 GPUs for four weeks generating more than 100 million rollouts with a maximum context length of 256K tokens. Training and grading used approximately 1.3 billion sandboxes. To sustain a run of this magnitude, we sourced one million high-quality coding, agentic, and STEM environments. We believe this is one of the largest scale RL runs conducted by any open lab to date. Across our evaluation suite, capabilities continued to improve as we increased RL compute, with no sign of a plateau.
We trained Beam with asynchronous policy gradients. At scale, policy staleness becomes a major source of instability for these methods. Long running rollouts have tokens that are generated by multiple model checkpoints, with earlier tokens becoming increasingly stale relative to the current policy. Numerical mismatch between training and inference engines further compounds this challenge.
We developed new algorithms to maintain stable learning under these conditions while systematically reducing training–inference mismatch throughout our pipeline. These advances enable fully asynchronous RL at scale that remains stable, even when learning from interactions generated more than a day earlier.
Learning to reason efficiently
We trained Beam with a controllable length penalty that rewards successful solutions while discouraging unnecessary tokens. Early in RL, performance improved even as completion lengths fell: the model learned to solve tasks more effectively with less reasoning. Later, as Beam developed stronger agentic capabilities, completion lengths grew again, but those additional tokens supported further gains in performance. Throughout training, RL improved the tradeoff between capability and token usage.
Users can control this tradeoff through Beam’s reasoning effort parameter: lower settings favor shorter responses, while higher settings allow longer reasoning to improve performance on demanding tasks. This gives users the flexibility to match reasoning effort to their task and compute budget.
How behavior generalizes with RL
We designed Beam’s RL training to develop reasoning and agentic capabilities that generalize beyond its training tasks. During a phase of training on reasoning, software engineering, and terminal tasks, we saw consistent gains in browsing despite the absence of browsing tasks from the RL mixture. This transfer suggests that Beam was learning broader agentic capabilities that generalize across domains. When given web access, it organically learned to search for and query other large language models, and to use OCR APIs to read documents.
The demos below showcase Beam applying these capabilities across research, application development, gameplay, and machine learning workflows. The examples range from building a live NYC subway dashboard using public data to creating interactive applications and preparing model fine-tuning notebooks. Although Beam is text-only, it can work with information from other modalities when represented as text. In another out of distribution domain, Beam also created a fine-tuning notebook for the latest and smallest Gemma-4 model on a Text2SQL task.
Together, these demos illustrate the breadth of tasks Beam can tackle by combining reasoning, coding, and tool use. Each example includes the initial request and resulting output, along with relevant setup and user iterations.
Scaling reinforcement learning environments
Frontier-scale reinforcement learning requires a large volume of difficult, high-quality tasks. We built a pool of nearly one million environments, primarily through synthetic data pipelines, supplemented by proprietary vendor data and open-source sources.
We relied heavily on an iterative curation process. First, we synthesized or sourced environments across a broad set of domains including software engineering, terminal use, competitive coding, STEM, web search, tool use, and general knowledge work. Second, we heavily filtered tasks for difficulty (ensuring they were neither consistently solvable, nor impossible for the model) and quality (e.g., not underspecified, misleading, guessable, hackable, or otherwise broken or noisy). Third, we tested the tasks through RL, which allowed us to identify further quality or difficulty issues and inform the next iteration of sourcing and filtering.
Throughout Beam’s development, we found that compromises in data quality led to capability plateaus and other training issues. Systematic improvements to task quality were essential to sustaining capability gains throughout the run, which ended with no sign of saturation.
Frontier RL infrastructure
High-compute agentic RL requires generating rollouts, executing tools, evaluating outcomes, and updating the model at scale. We built an asynchronous platform that lets these processes run independently while coordinating the flow of experience and model updates.
During Beam’s training, we sustained an average of 110K concurrent rollouts. Seven capabilities made this practical:
Fully asynchronous execution: Agents generate rollouts while the trainer learns and publishes new model versions. Each token is tagged with the version that produced it, allowing the training algorithm to account for policy staleness as completed rollouts flow into training.
Flexible compute allocation: We adjusted the balance between inference and training as the workload evolved, operating at inference-to-training GPU ratios from 3.9:1 to 5.4:1. We also resized the trainer across five GPU mesh configurations within the same training lineage without losing training state.
Fast model updates: New weights reached the inference fleet in a median of approximately 12 seconds. Hierarchical distribution transfers weights across racks over RoCE, then shares them locally over NVLink. Compared with every replica pulling weights directly, this reduced cross-rack traffic by 75% and made fleet-wide adoption of new weights 2.2× faster.
Resilience to inference failures: During the run, 71 inference incidents were handled without terminating the training job. Inference capacity recovered in a median of eight minutes, with lost capacity accounting for just 0.02% of elapsed serving GPU-minutes.
Environments at scale: We supported up to 170K concurrent sandboxes during the run. Across the platform, we processed more than one billion sandbox creation requests, spanning over 20 clusters, two clouds, and four regions. 90% of new sandboxes were ready in under 10 seconds.
Efficient Trainer Packing: Dynamic packing kept training batches 99.99% full on average, holding per-GPU trainer throughput within 1.5% as mean rollout length grew almost 70%.
Observability and reward integrity: Per-token records enabled numerical consistency checks between training and inference at every step. Independent judges re-screened passing solutions for verifier exploits, while replayable records made rewards and their use in training inspectable.
Together, these capabilities enabled us to train on longer interactions and more demanding environments while maintaining throughput, recovering from failures, and checking the integrity of the learning process.
Pretraining a Foundation for Reasoning #
Reinforcement learning builds on top of a robust base model. To facilitate reasoning, we ensured Beam’s foundation had rich knowledge in coding domains, innate agentic capabilities that could be amplified, and stable MoE optimization dynamics.
While developing Beam, we pretrained a series of iteratively bigger models to establish and verify our scaling recipe. Making their performance predictable required carefully designing model tiers, curating diverse in-house code and web validation sets, and aggressively decontaminating all training data against them. The scaling held; the final Beam Base matches its predicted performance, and also matches or outperforms accessible similar-sized open-source base models.
Stable and balanced MoE optimization
The Beam architecture and optimization recipe emphasizes a numerically healthy foundation for sustained downstream RL. It combines interleaved local and global attention, fine-grained routed experts, a controlled residual stream, and multiple forms of load balancing to ensure stable, balanced expert utilization and healthy signal propagation through residuals.
For expert utilization, we built on auxiliary-loss-free load balancing [(DeepSeek-AI et al., 2024)](https://arxiv.org/abs/2412.19437v2), introducing cosine decay of expert-bias updates to reduce routing perturbations later in training. Sequence-level balancing further encourages balanced expert utilization on data outside the pretraining distribution, preparing the model for the changing distribution of downstream RL. As a result, the final pretrained base has almost-perfect uniform utilization, ensuring all experts can be used for learned reasoning.
For the residual stream, we developed a depth-based scaling approach that counteracts activation growth as sublayer outputs accumulate, helping keep residual norms stable as model depth increases. Combined with SandwichNorm, elementwise attention gating, and FP32 residual accumulation – which reduces rounding error when adding small updates to the stream – this recipe controls activation growth and outliers, thereby ensuring healthy signal propagation throughout all layers of Beam. This stability persists throughout pretraining, reinforcement learning, and alignment.
Quality-centric data curation
Beam was pretrained on 23.8 trillion diverse high-quality tokens from the web, public sources, and proprietary licensed datasets. Our data pipeline was designed to give Beam a foundation for downstream agentic coding: source code, technical explanations, and mathematical and scientific knowledge, preserved through every stage of curation. We train on almost all publicly accessible and unrestrictively-licensed code and code documentation on the web.
We trained our own quality classifiers for web, code, and STEM content, divided data into fine-grained quality tiers, and weighed training toward stronger material. After extensive scientific iteration, we optimized both precision and recall of data curation significantly beyond conventional web filters used in state-of-the-art OSS data frameworks. On one hand, about 95% of raw Internet tokens are eliminated through parsing, deduplication, and curation. On the other hand, we found that conventional techniques would have missed roughly 1.8 trillion high-quality tokens we retain, including 87% of our curated web-code tokens.
Code modeling requires its own curation for the highest performance. For each language, we applied individually tuned filters, removed low-quality autogenerated and unlearnable content, and trained classifiers to identify corrupted content or code that may hurt training stability. Overrepresented languages and file types are rebalanced to further broaden exposure.
We also developed a high-throughput pipeline for processing PDF artifacts, to ensure the base Beam has knowledge across a wide range of STEM topics. It integrates a vision-language OCR model with in-house quality classifiers and artifact detectors that catch malformed reconstructions, distributed across thousands of GPUs to process petabytes of technical data.
We repeated code and technical content multiple times to increase the model’s exposure over the course of its training horizon. This required careful attention to fuzzy deduplication, packing algorithms, and the science of overtraining, to ensure every repeated data source helps rather than hurts generalization.
Frontier-grade pretraining infrastructure
Beam was pretrained end-to-end in under four weeks on a cluster of 6,144 NVIDIA GB300 NVL72 GPUs. To achieve the reliability, performance, and development velocity required to train Beam at scale, we built nearly the entire infrastructure stack in-house. This included a novel topology-aware, Kubernetes-based scheduler across our clusters; an internal node lifecycle system with continuous health monitoring and alerts; and a silent data corruption (SDC) detection system capable of semi-autonomous rewinds and restarts. Together, these systems gave us the performance and operational control required to train our own frontier model efficiently and with greater stability.
As a result of extensive investments in the training recipe stability, infrastructure, and data quality, the overall pretraining run finished with an extremely smooth trajectory. We executed nine semi-automatic rewinds throughout, attributed either to non-deterministic gradient norm spikes or to suspected SDCs. In addition, the run’s goodput (the share of wall-clock time spent on training steps retained in the final model) reached 92.3% towards the end thanks to improvements in checkpointing, fault detection, and node health management.
Building a strong prior for RL in Midtraining
Our midtraining stage was designed specifically as a foundation for high-compute RL, developing the knowledge, reasoning, and tool-use capabilities that allow Beam to learn from more demanding tasks.
We built multi-stage data curation pipelines that capture the structure and complexity of real-world tasks while expanding coverage of capabilities that are difficult to learn from raw data alone. Starting from carefully selected real-world examples, these pipelines transform, combine, and extend material into training data designed to teach specific capabilities. This includes long, reasoning-rich documents that expose the model to extended chains of logic.
Midtraining also extends Beam’s effective context length to 1M tokens. We combine structured code repositories, long-horizon tasks, and high-quality long-form documents to teach the model to identify, retain, and connect relevant information across long sequences.
Together, these capabilities give RL a stronger starting point, enabling Beam to explore more complex solutions, work through longer interactions, and learn from tasks that would otherwise be out of reach.
Safety and Alignment #
Our safety and alignment work involved training a second model from our pretrained checkpoint using a separate SFT and RL pipeline on data specifically targeting the principles by which the model should abide. We merged the capabilities of the two teachers – a large-scale RL teacher and a dedicated safety and alignment teacher – via multi-teacher on-policy distillation (MOPD).
We organized Beam's safety and alignment principles into three tiers:
(1) Rules that Beam should not break: following our safety policies and maintaining its identity as an AI agent.
(2) Qualities that Beam should consistently satisfy, such as: using the context it is given, making accurate claims, acknowledging uncertainty, and transparently following the user’s request.
(3) The default style for how Beam should interact: direct, thorough, efficient, and proactive in anticipating what the user may need next.
We designed the RL environments in the alignment stage to incentivize Beam’s adherence to these principles. Many of these environments involved non-verifiable rewards judged by a generative reward model, yet we were able to use them to predictably shape the model’s behavior. Forecasting how rewards would affect different behaviors before running RL, we were able to predict RL gains better than the Best-of-N ceiling alone (r = 0.79 compared to r=0.46 with BoN ceiling). This allowed us to iterate on rubrics and reward design, squash behaviors like hallucinations and excessive formatting, and improve Beam’s overall interaction quality.
Our safety training used deliberative alignment (Guan et al., 2024) techniques to incorporate our safety policy directly into the model's reasoning. The dataset was built adversarially and iteratively: in each round, we trained a model, generated prompts that elicited harmful or over-refusing behavior from it, and folded the successful attacks back into SFT mixture for the next round. For safety RL, we similarly sourced single-turn, multi-turn, jailbreak, and agentic scenarios in which a simulated adversary pressures a tool-using model to take unsafe actions, to simultaneously reduce over-refusals and harmful compliance.
We will publish the results of our safety evaluations in our model technical report and will open-source safety evaluations we developed and used internally to create a shared, inspectable standard that the open ecosystem can test against and contribute to.
The Path Ahead #
This preview shows what Beam can do today. We are making this early version of Beam available to a select group of users; you can sign up for the waitlist here.
We want Beam to be widely accessible and easy to build on. This month, we will release the weights under an Apache 2.0 license, along with documentation and the full stack for running, evaluating, and fine-tuning the model. We will be launching Beam with an ecosystem of distribution partners, as well as integration with a broad range of open source libraries and harnesses, so developers can use Beam across existing open-source workflows.
Beam is the first model in a series, and the first demonstration of the open intelligence our team is committed to building. We are already training what comes next, with the goal of bringing the open frontier closer to the frontier of intelligence with every release.