Changes confirmed medium confidence

Reflection AI Previews 501B Beam Model Before Open-Weight Release

The sparse mixture-of-experts model is in limited early access while weights, a technical report, model card and developer artifacts remain promised for later in October.

Edited by Tyronne Panaino

Reflection AI introduced Beam on October 5 as its first open-weight model, but the model is not yet broadly available as a downloadable release. The company's preview says an early version is open to a select group while final red-teaming and evaluations continue. Reflection plans to publish the weights, technical report, model card and developer artifacts later in October.

Beam is a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active. Reflection positions it for coding, reasoning and agentic work. The distinction between preview and release matters: readers can assess the architecture and the company's reported results now, but the artifacts needed for independent inspection, deployment and safety review remain a future checkpoint.

A limited preview comes before the open-weight package

Reflection says it will release Beam's weights under an Apache 2.0 license, together with documentation and a stack for running, evaluating and fine-tuning the model. It also plans integrations with distribution partners, open-source libraries and agent harnesses. None of those future availability statements should be read as proof that the full package was public when this article was released.

For developers, the practical test begins when the promised files arrive. Weight access alone will not answer questions about supported inference formats, memory requirements, quantization, context handling, tool interfaces or reproducibility. A model card and technical report should also clarify the evaluation settings, safety findings and intended deployment boundaries that the preview leaves incomplete.

Reflection reports training at substantial scale

The company says Beam was pretrained on 23.8 trillion tokens drawn from web material and proprietary licensed datasets. It also reports a reinforcement-learning run that generated more than 100 million rollouts on 10,500 NVIDIA GB300 GPUs over four weeks. Those figures describe Reflection's own training record and were not independently audited in the fetched evidence.

The preview describes a large asynchronous reinforcement-learning platform in which agents generate rollouts while the trainer learns and publishes new model versions. Reflection says the system tracked model versions at token level, moved new weights through its inference fleet and supported large numbers of concurrent sandboxes. These infrastructure claims help explain how the lab says it scaled agent training, but they do not establish the cost, efficiency or reliability another operator would achieve.

Benchmark tables are evidence, not a verdict

Reflection publishes results across coding, terminal, reasoning, search and tool-use tests and compares Beam with several other open models. It also presents inference-efficiency estimates based on active parameter counts and generated tokens. The company explicitly notes that these estimates exclude prompt prefill, context-dependent attention work and serving overhead, making them approximate compute comparisons rather than measured end-to-end inference cost.

That methodological limit is important. Benchmark tables can indicate where a model deserves further testing, but they cannot establish production latency, total hardware demand, task reliability or operating cost across different serving stacks. Independent evaluators will need the weights, exact harnesses and evaluation settings before the comparison can be reproduced on equal terms.

Safety evidence is still part of the release checkpoint

Reflection says Beam is undergoing final red-teaming and evaluations. The company describes a separate safety-and-alignment training pipeline and says it plans to publish safety results in the technical report, along with internally developed safety evaluations. Until those materials are public, readers cannot inspect the full safety methodology, results, failure categories or release mitigations from this source alone.

The next verifiable milestone is therefore not another benchmark claim. It is publication of the weights, license, model card, technical report, developer stack and safety results promised for October, followed by independent testing. That evidence will show whether the announced open-weight package is complete and whether outside teams can reproduce the reported capabilities and constraints.

Status

Confirmed. Reflection AI has announced Beam and opened a limited early-access preview. Internal confidence is medium because model performance, training scale and safety methods are reported by the developer, while the weights and supporting evidence package were not yet public in the fetched source.

Sources

Update note: Last reviewed 2026-10-08. We will revise this post when Reflection publishes the promised weights, model card, technical report, developer artifacts and safety results.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.

More Changes coverage