Reflection Beam: 501B Parameters, 23B Active, 3–4x Less Compute Claim
In this article
Reflection AI, a two-year-old Brooklyn startup founded by two former Google DeepMind researchers, introduced Beam on Monday, its first frontier open-weight model. The announcement confirms earlier Axios reporting and positions Beam as a Western answer to open-weight models from DeepSeek, Qwen, and Z.ai.
Beam is a text-only mixture-of-experts model with 501 billion total parameters and 23 billion active parameters during inference. Reflection reports pretraining on 23.8 trillion tokens and a 1 million-token context window. The company describes the model as trained with high-compute reinforcement learning to be effective at reasoning, coding, and agentic tasks at "a fraction of the token cost and inference time compute" of rivals.
By contrast, Z.ai's GLM-5.2 contains roughly 744 billion total parameters with 40 billion active parameters. Beam's smaller active path is central to Reflection's claim that the model uses "3-4x less inference compute" while matching GLM-5.2 on advanced reasoning benchmarks and outperforming leading Western open models. Those performance figures have not been independently verified; Reflection calls Beam a "workhorse model" for enterprises, the public sector, and developers. The company plans to release weights and full technical details later this month, with distribution through hyperscalers and neoclouds and integrations across open-source libraries.
| Model | Developer | Total Parameters | Active Parameters | Modality |
|---|---|---|---|---|
| Beam | Reflection AI | 501 billion | 23 billion | Text-only |
| GLM-5.2 | Z.ai | 744 billion | 40 billion | Not specified |
| Inkling | Thinking Machines Lab | Not disclosed | Not disclosed | Multimodal |
Reflection is positioning Beam against closed labs such as OpenAI and Anthropic, Chinese open models, and Western rivals including Meta, Mistral, and Cohere. TechCrunch AI notes that Beam's most direct U.S. rival might be Inkling, the open model from Mira Murati's Thinking Machines Lab released in July. In Reflection's own benchmarks, Beam outscores Inkling on four coding tests where both models report results. The comparison has a structural caveat: Inkling is multimodal, while Beam is text-only.
The startup has raised about $4.7 billion from investors including Nvidia, Sequoia Capital, and Lightspeed Venture Partners, according to PitchBook. Its last round valued Reflection at a $25 billion pre-money valuation. The company has also been locking up compute, signing deals this summer worth more than $7 billion with SpaceX and Nebius to secure access to Nvidia GB300 chips through 2029.
Reflection is targeting enterprises and sovereign nations with an "AI factory" strategy, in which institutions train open-weight models on their own proprietary data. Jensen Huang, whose company backs Reflection, has long promoted the AI factory concept. That open ecosystem push would also benefit Nvidia, whose GPUs would power those AI factories. Axios reported that hedge funds and trading firms are among those eager to build such systems. Reflection has begun testing a sovereign AI factory partnership with Shinsegae Group in South Korea.
This infrastructure bet aligns with a broader shift in AI gains toward systems work: infrastructure rewrites not model weights drive 2026 ai gain. For enterprises evaluating Beam, the hardware requirement question will turn on active parameters rather than the 501 billion total.
AI Mastery analysis
Beam's design is less about raw scale than about the ratio between stored parameters and active inference cost. A 501B-parameter model that activates only 23B parameters can retain a large trained capacity while minimizing per-token compute, which matters for self-hosted and sovereign deployments. Beam's 23 billion active parameters are substantially lighter than GLM-5.2's 40 billion, but compute savings also depend on routing efficiency, memory bandwidth, and serving software.
The unresolved question is whether a text-only, heavily optimized model can hold up against multimodal open-weights such as Inkling outside Reflection's chosen coding benchmarks. Reflection's comparison with Inkling is narrow: multimodal models often route capacity to non-text modalities that Beam does not support. Until independent evaluations appear, the "3-4x less inference compute" figure should be treated as a vendor-reported architectural target, not a production baseline. The open-weight ecosystem is bifurcating between broad multimodal models and narrower, compute-sparse reasoning engines, and Beam will now test how far a compute-sparse design can go against more general rivals.
Sources
Frequently asked questions
How many active parameters does Reflection Beam use during inference?
Beam has 501 billion total parameters but activates only 23 billion during inference. Reflection says the sparse mixture-of-experts design is what enables lower compute per token.
Does Reflection Beam match Z.ai GLM-5.2 on reasoning benchmarks?
Reflection claims Beam scores on par with GLM-5.2 on advanced reasoning benchmarks while using 3-4x less inference compute. TechCrunch notes these performance claims have not been independently verified.
What is Reflection Beam's context window?
Reflection reports that Beam was pretrained on 23.8 trillion tokens and supports a 1 million-token context window. It is text-only and trained with high-compute reinforcement learning for reasoning, coding, and agentic tasks.
When will Reflection release Beam weights?
Reflection says it will release Beam's weights and full technical details later this month, with distribution through hyperscalers and neoclouds. Open-source library integrations are planned at launch.
Related Reading

NVIDIA Nemotron 3.5 Lightning: 30B MoE with 3B Active Parameters
NVIDIA ships Nemotron 3.5 Lightning, a 30B MoE with 3B active parameters and 1M-token context, plus NeMo Switchyard for per-step agent routing.
Kolibri Activates Just 3.46B of Its 78.1B Parameters
Aleph Alpha’s bilingual MoE offers a 1,048,576-token context and an FP8 checkpoint designed to run on one high-memory GPU.
Jina-ocr-v1: 3.4B MoE Parses 2.57 Pages/sec on One A100
Jina AI's 3.4B MoE document parser posts 2.57 pages/sec on a single A100 — highest of 14 systems benchmarked — via lossless speculative decoding.