Reflection publishes Beam benchmark scores ahead of its weight release Reflection AI published benchmark scores for Beam, its first open-weight model, showing 80.1 on Terminal-Bench 2.1 versus Z.ai's GLM 5.2 at 81.0 and 65.5 on SWE-bench Pro v1 versus GLM 5.2's 62.1, ahead of the model's planned weight release under an Apache 2.0 license. Co-founder Misha Laskin introduced Beam on October 5th as a 501-billion-parameter mixture-of-experts model with 23 billion active parameters per token, trained on 23.8 trillion tokens and a four-week reinforcement-learning run on 10,500 Nvidia GB300 GPUs generating more than 100 million rollouts. Reflection claims Beam's reasoning scores are comparable to GLM 5.2 at three to four times less inference compute, but its published estimate excludes prompt prefill, context-dependent attention and serving overhead, leaving outside testing to establish real deployment costs. Reflection publishes Beam benchmark scores ahead of its weight release Reflection AI's first scorecard puts Beam near or ahead of some open models on selected coding tests and behind others; its compute comparison excludes several costs of serving a model. By Ryan Merket https://runtimewire.com/author/ryan-merket · Published · Updated Primary source: Financial Times https://www.ft.com/content/353be303-3271-43f0-aefb-69b6ed7a0a6f Why it matters Beam's scorecard gives developers an initial basis for comparing Reflection's first open-weight model on coding and reasoning tasks. Its uneven results and compute estimates, which omit parts of the serving workload, leave outside testing to establish how the claims translate to real deployments. On October 5th, Misha Laskin https://twitter.com/MishaLaskin?ref=runtimewire introduced Beam https://reflection.ai/blog/introducing-beam?ref=runtimewire , the first open-weight model from Reflection AI https://reflection.ai/?ref=runtimewire , the startup he co-founded with fellow former Google DeepMind researcher Ioannis Antonoglou. Reflection is pitching the model as a US-built alternative to Chinese systems that have become leading options for developers and organizations seeking customizable AI. Laskin left research work at DeepMind, where he worked on Gemini and reinforcement learning https://runtimewire.com/models/huggingface/ydiffraction-reinforcement-learning-7897455e5c11c599 , to build Reflection with Antonoglou, an early DeepMind engineer who helped create AlphaGo. Before DeepMind, Laskin completed a theoretical physics PhD at the University of Chicago, worked as a postdoctoral scholar at UC Berkeley, and founded a Y Combinator-backed startup, according to his personal site https://mishalaskin.github.io/?ref=runtimewire . Reflection's original 2024 launch focused on autonomous coding; it has since made open models and user control central to its pitch. In its account of its open-model thesis https://reflection.ai/blog/frontier-open-intelligence?ref=runtimewire , Reflection argues that developers, companies and governments should be able to inspect, adapt and operate AI systems themselves. Beam is intended for coding, reasoning and agentic tasks, where a model can use tools to complete multi-step work. Laskin told the Financial Times https://www.ft.com/content/353be303-3271-43f0-aefb-69b6ed7a0a6f?ref=runtimewire that the company sees demand from organizations that want sovereign AI systems and do not want to rely on Chinese models. A benchmark claim with a narrower edge Beam is a 501-billion-parameter mixture-of-experts model, with 23 billion parameters active for each token, according to Reflection's technical announcement https://reflection.ai/blog/introducing-beam?ref=runtimewire . The company says it trained the model on 23.8 trillion tokens and used 10,500 Nvidia GB300 GPUs for a four-week reinforcement-learning run that generated more than 100 million rollouts. The company-published results are mixed. Reflection's scorecard shows Beam at 80.1 on Terminal-Bench 2.1, just below Z.ai's GLM 5.2 https://runtimewire.com/models/z-ai/glm-5.2 at 81.0. On SWE-bench Pro v1, Beam scores 65.5 against GLM 5.2's 62.1. It trails several listed models on some other tasks, including Kimi K3 https://runtimewire.com/models/moonshotai/kimi-k3 and Qwen 3.8-Max on Terminal-Bench. These are company-published comparisons, not a single overall ranking. Reflection says Beam's reasoning scores are comparable to GLM 5.2 while requiring three to four times less inference compute. Its published estimate uses benchmark data from Artificial Analysis and DataCurve, but excludes prompt prefill, context-dependent attention and serving overhead. That makes it a model-compute comparison, not a measured estimate of what a customer will pay to run the system. At launch, Reflection offered early access by sign-up and said it planned to release Beam's weights under an Apache 2.0 license, together with a technical report and developer tools, later in October. Those materials will let outside developers examine the system and test the performance claims against their own workloads. A costly bet on control Open-weight models accounted for 56 percent of tokens processed through Vercel's AI Gateway in August, up from 7 percent in December, the Financial Times reported https://www.ft.com/content/353be303-3271-43f0-aefb-69b6ed7a0a6f?ref=runtimewire . That figure describes Vercel's gateway traffic, not the entire AI market. It also points to the commercial appeal of a capable model that customers can run themselves. The FT reports that Reflection has raised more than $4 billion since its 2024 launch and reached a $25 billion valuation. Semafor reported https://www.semafor.com/article/10/05/2026/reflection-ai-unveils-an-open-source-answer-to-chinese-labs?ref=runtimewire that Reflection closed its latest funding round at a $25 billion pre-money valuation in June, naming Nvidia, Sequoia and Citigroup among its investors. With that scale of financing, Beam needs to give investors, customers and government partners reason to believe Reflection can turn its open-model ambition into a durable business. The strategy extends beyond model downloads. The FT reports that Reflection has partnered with the Pentagon and the US Department of Energy, and has agreements to develop models for US allies including South Korea. Semafor says Reflection is also pursuing government and business demand for systems that customers can control, and Laskin told the outlet the company is training a more powerful follow-on model. Semafor https://www.semafor.com/article/10/05/2026/reflection-ai-unveils-an-open-source-answer-to-chinese-labs?ref=runtimewire Chinese model builders are also connecting open weights to commercial use. RuntimeWire reported in August that Moonshot AI was seeking a share of cloud-service revenue tied to its Kimi K3 model https://runtimewire.com/article/moonshot-kimi-k3-us-cloud-revenue-sharing . Reflection frames its case around infrastructure and sovereignty: organizations should be able to run and customize models on their own terms, with the option to host them instead of calling a hosted API. Before Reflection releases its next model, Beam needs to show that an open US system can compete on real coding and agent workloads and that its efficiency claims hold up when organizations calculate deployment costs. Third-party use of the weights, once released, will test those claims against workloads beyond Reflection's scorecard.