AI · 7 October 2026
Reflection's Beam becomes the most capable open-weight model built outside China
Reflection has released Beam, its first open-weight model. The mixture-of-experts system activates just 23 billion of its 501 billion parameters per token and aims to match GLM 5.2 on coding and reasoning while using three to four times less compute. The strategy against Chinese open-weight rivals like Deepseek and Qwen is efficiency over raw performance. The article Reflection's Beam becomes the most capable open-weight model built outside China appeared first on The Decoder .
What happened
Reflection has released Beam, its first open-weight language model, positioning it as the most capable open-weight system built outside China. Beam is a mixture-of-experts model with 501 billion total parameters, of which only 23 billion are activated per token — a design intended to deliver strong reasoning and coding performance while keeping inference costs far lower than comparably sized dense models.
According to The Decoder, Reflection says Beam is built to match the coding and reasoning benchmarks of GLM 5.2, a leading Chinese open-weight model, while using three to four times less compute to do so. The company is framing its approach explicitly as efficiency-first, rather than chasing the raw benchmark scores pursued by rivals such as DeepSeek and Qwen.
Why it matters
Beam's release underscores a widening split in open-weight AI strategy: where several Chinese labs have competed primarily on headline benchmark performance, Reflection is betting that compute efficiency — not just capability — will determine which open models get adopted at scale. A mixture-of-experts architecture that activates a small fraction of its parameters per token lowers the cost of running large models in production, which matters directly to any organisation weighing whether to self-host AI rather than rely on proprietary APIs.
For technology and transformation leaders, the significance lies less in Beam's specific benchmark scores and more in what it signals about the open-weight model market: efficiency is becoming a competitive axis in its own right, and the field of serious open alternatives to closed, proprietary models is no longer dominated by a single region.
By the numbers
- 501 billion total parameters in Beam's mixture-of-experts architecture
- 23 billion parameters activated per token during inference
- 3 to 4 times less compute claimed versus GLM 5.2 for comparable performance
The Renascence take
Coverage of new model releases tends to fixate on benchmark parity — who matches whom on coding or reasoning tests. The more consequential story in Beam is architectural: efficiency claims like this are really a bet on who gets to run frontier-grade AI in-house, at a cost that makes continuous, embedded use viable rather than a budget-gated experiment.
Most organisations evaluating open-weight models still ask "how smart is it?" before asking "what does it cost to run this, every day, for every employee or customer interaction?" That second question is where efficiency-first releases like Beam actually compete. A customer-obsessed operator should treat compute efficiency as a service-design input, not an engineering footnote — because the models that are cheap enough to run everywhere are the ones that will quietly end up shaping everyday service and support experiences, regardless of which one tops a benchmark table this month.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.
