Reflection AI has unveiled Beam, a 501-billion-parameter open-weight model that activates only about 23 billion parameters for each token it generates. The design is a direct pitch to enterprise buyers: frontier-level coding and reasoning at a fraction of the usual running cost.
The company says it pretrained the text-only model on 23.8 trillion tokens, then ran large-scale reinforcement learning across more than 100 million rollouts. It claims Beam matches a leading Chinese open model on advanced reasoning while using three to four times less inference compute, and reports strong scores on standard coding benchmarks. Those figures are company-reported and await independent testing.
The architecture is the story. Beam uses a sparse mixture-of-experts design, routing each request to a subset of specialised components. Total parameter count therefore overstates the true cost of serving the model — a distinction technology executives increasingly use to compare AI suppliers on cost per task rather than headline size.
Reflection, founded by former Google DeepMind researchers and backed by Nvidia among others, plans to release the weights under an Apache 2.0 licence later this month after final red-team testing. Until then, developers cannot self-host, quantise or fine-tune it, and early access runs through a waitlist.
For corporate technology chiefs, the strategic question is whether open weights reduce dependence on a small number of closed API providers. A credible Western open model would give governments and regulated industries a sovereign option they currently source largely from Chinese labs.
The caveat is timing. Rivals’ newer models still lead on some raw-capability measures, and Beam’s efficiency advantage will only be proven once outsiders can run it on their own workloads and their own hardware bills.



