Reflection AI's first model is out, and the pitch isn't that it beats China's best open models outright. It's that it matches one of them while burning far less compute than any other American or European rival.
Reflection AI unveiled Beam on October 5, the Nvidia-backed startup's first frontier open-weight model and the concrete follow-through on a plan the company had been signaling for months. According to TechCrunch, Beam scores comparably on advanced reasoning benchmarks to GLM-5.2, the flagship open model that Beijing-based Z.ai released in July. Reflection isn't claiming it beat GLM-5.2. It's claiming it matched it, then did so for a fraction of the compute.
That efficiency claim is the real headline. Reflection says Beam is three to four times more efficient than rival open models built by Western companies, measured in token cost and inference-time compute for coding and agentic tasks. The company credits that gain to high-compute reinforcement learning during training, a method meant to make the model reason in fewer steps rather than simply throwing more parameters at the problem. Weights are due out later this month, Reflection told TechCrunch, which means the benchmark claims are not yet independently verifiable against a public checkpoint.
Frankly, the compute angle matters more to the business than the benchmark score does. Every point of reasoning quality Beam can hold while using a quarter of the tokens is a point of margin for whoever runs it. Founders and infrastructure buyers evaluating model economics have spent 2026 watching DeepSeek and Qwen undercut Western labs on price per token. A model that claims Chinese-tier reasoning at Chinese-tier cost, but built and hosted in the US, is exactly the gap Reflection is trying to fill.
Reflection was founded in March 2024 by Misha Laskin, now its CEO, and Ioannis Antonoglou, its president and CTO, who spent twelve years at Google DeepMind and co-created AlphaGo. The company has raised roughly $4.7 billion in total. Its Series A priced the company at around $545 million. By October 2025, a $2 billion round led by Nvidia's $800 million check pushed that to $8 billion. As of June 2026, Bloomberg and Fortune both put Reflection's valuation at $25 billion, a roughly 45x climb in a little over a year. Other backers include Sequoia Capital, Lightspeed Venture Partners, DST Global, former Google CEO Eric Schmidt, and Donald Trump Jr.'s 1789 Capital.
Reflection AI readies a US open-weight model to challenge DeepSeek and Qwen Axios's October 4 scoop names no model, date, or benchmark yet, but says Nvidia-backed Reflection, now valued at $25 billion, has over $7 billion in SpaceX and Nebius compute already committed. - open weight AI model challenges DeepSeek Qwen - Nvidia backed Reflection AI open source model
Nvidia doesn't need another model to sell chips. It needs the market for open-weight AI to stay competitive enough that enterprises keep buying GPUs to run models themselves rather than renting inference from a handful of closed API providers. A credible American open-weight lab serves that goal directly, and Reflection explicitly frames itself as the US answer to DeepSeek and Qwen rather than a challenger to closed labs like OpenAI or Anthropic.
The company has also locked down compute to back the ambition. It struck a multi-year deal with SpaceX worth up to $6.3 billion for access to Nvidia GB300 chips at SpaceX's Colossus 2 facility starting in July 2026, and committed more than $1 billion to European cloud provider Nebius through 2029 for the same chip family. That's an unusual amount of infrastructure for a company that, until this week, had shipped exactly zero public models.
Z.ai's GLM-5.2, the model Reflection is benchmarking against, has been out since July and already has real-world adoption to its name, which gives Beam's comparison some teeth. It's not matching a lab demo. It's matching a model developers are actually using. Whether Beam's efficiency claims hold up once outside researchers can run their own tests is the open question, and it won't get answered until the weights actually land this month. Until then, Reflection has a benchmark chart and a very large bet riding on it.
Also read: ReleasePad Wants Your AI Coding Assistant To Write Your Release Notes Too • C.H. Robinson buys RXO for $5.8 billion to bet big on AI freight matching • System One Models Explained - What Are They and Why Should Founders Care
This article is posted in AI News, check it out for more related stories.
Join the discussion #
Open in the community → Almost there. Sign in and your reply posts straight away.
Baseten built the fastest GLM-5.2 API on earth and the playbook tells you where inference is heading
Baseten is serving Zhipu AI's GLM-5.2 at 593.7 tokens per second, roughly 12.8 times faster than the next-fastest provider. The optimization stack , NVFP4 quantization on NVIDIA Blackwell, prefill-decode disaggregation via NVIDIA Dynamo, and multi-token prediction , is a preview of how the inference compute race gets won, and why deployment... - fastest GLM-5.2 API provider - inference speed tokens per second