Reflection Beam: 501B Parameters, 23B Active, 3–4x Less Compute Claim

October 5, 2026 • news
Open WeightsMixture of ExpertsNVIDIA

Reflection AI, a two-year-old Brooklyn startup founded by two former Google DeepMind researchers, introduced Beam on Monday, its first frontier open-weight model. The announcement confirms earlier Axios reporting and positions Beam as a Western answer to open-weight models from DeepSeek, Qwen, and Z.ai.

Beam is a text-only mixture-of-experts model with 501 billion total parameters and 23 billion active parameters during inference. Reflection reports pretraining on 23.8 trillion tokens and a 1 million-token context window. The company describes the model as trained with high-compute reinforcement learning to be effective at reasoning, coding, and agentic tasks at "a fraction of the token cost and inference time compute" of rivals.

By contrast, Z.ai's GLM-5.2 contains roughly 744 billion total parameters with 40 billion active parameters. Beam's smaller active path is central to Reflection's claim that the model uses "3-4x less inference compute" while matching GLM-5.2 on advanced reasoning benchmarks and outperforming leading Western open models. Those performance figures have not been independently verified; Reflection calls Beam a "workhorse model" for enterprises, the public sector, and developers. The company plans to release weights and full technical details later this month, with distribution through hyperscalers and neoclouds and integrations across open-source libraries.

Model Developer Total Parameters Active Parameters Modality
Beam Reflection AI 501 billion 23 billion Text-only
GLM-5.2 Z.ai 744 billion 40 billion Not specified
Inkling Thinking Machines Lab Not disclosed Not disclosed Multimodal

Reflection is positioning Beam against closed labs such as OpenAI and Anthropic, Chinese open models, and Western rivals including Meta, Mistral, and Cohere. TechCrunch AI notes that Beam's most direct U.S. rival might be Inkling, the open model from Mira Murati's Thinking Machines Lab released in July. In Reflection's own benchmarks, Beam outscores Inkling on four coding tests where both models report results. The comparison has a structural caveat: Inkling is multimodal, while Beam is text-only.

The startup has raised about $4.7 billion from investors including Nvidia, Sequoia Capital, and Lightspeed Venture Partners, according to PitchBook. Its last round valued Reflection at a $25 billion pre-money valuation. The company has also been locking up compute, signing deals this summer worth more than $7 billion with SpaceX and Nebius to secure access to Nvidia GB300 chips through 2029.

Reflection is targeting enterprises and sovereign nations with an "AI factory" strategy, in which institutions train open-weight models on their own proprietary data. Jensen Huang, whose company backs Reflection, has long promoted the AI factory concept. That open ecosystem push would also benefit Nvidia, whose GPUs would power those AI factories. Axios reported that hedge funds and trading firms are among those eager to build such systems. Reflection has begun testing a sovereign AI factory partnership with Shinsegae Group in South Korea.

This infrastructure bet aligns with a broader shift in AI gains toward systems work: infrastructure rewrites not model weights drive 2026 ai gain. For enterprises evaluating Beam, the hardware requirement question will turn on active parameters rather than the 501 billion total.

AI Mastery analysis

Beam's design is less about raw scale than about the ratio between stored parameters and active inference cost. A 501B-parameter model that activates only 23B parameters can retain a large trained capacity while minimizing per-token compute, which matters for self-hosted and sovereign deployments. Beam's 23 billion active parameters are substantially lighter than GLM-5.2's 40 billion, but compute savings also depend on routing efficiency, memory bandwidth, and serving software.

The unresolved question is whether a text-only, heavily optimized model can hold up against multimodal open-weights such as Inkling outside Reflection's chosen coding benchmarks. Reflection's comparison with Inkling is narrow: multimodal models often route capacity to non-text modalities that Beam does not support. Until independent evaluations appear, the "3-4x less inference compute" figure should be treated as a vendor-reported architectural target, not a production baseline. The open-weight ecosystem is bifurcating between broad multimodal models and narrower, compute-sparse reasoning engines, and Beam will now test how far a compute-sparse design can go against more general rivals.

Sources

Frequently asked questions

How many active parameters does Reflection Beam use during inference?

Beam has 501 billion total parameters but activates only 23 billion during inference. Reflection says the sparse mixture-of-experts design is what enables lower compute per token.

Does Reflection Beam match Z.ai GLM-5.2 on reasoning benchmarks?

Reflection claims Beam scores on par with GLM-5.2 on advanced reasoning benchmarks while using 3-4x less inference compute. TechCrunch notes these performance claims have not been independently verified.

What is Reflection Beam's context window?

Reflection reports that Beam was pretrained on 23.8 trillion tokens and supports a 1 million-token context window. It is text-only and trained with high-compute reinforcement learning for reasoning, coding, and agentic tasks.

When will Reflection release Beam weights?

Reflection says it will release Beam's weights and full technical details later this month, with distribution through hyperscalers and neoclouds. Open-source library integrations are planned at launch.

Free interactive tools for the decisions this piece raises.

Related Reading