The Spectrum Dispatch News

technology

Reflection Releases Beam, 501B Open-Weight Model for Coding and Reasoning

Beam is a sparse mixture-of-experts model with 23 billion active parameters, trained with large-scale reinforcement learning on 10.5K GPUs for four weeks.

Reflection Releases Beam, 501B Open-Weight Model for Coding and Reasoning

Reflection has introduced Beam, its first open-weight model, a sparse Mixture-of-Experts architecture with 501 billion total parameters and 23 billion active parameters. According to the company, Beam is designed for coding, reasoning, and agentic workloads and will be released later this month, with early access available via signup.

Reflection Releases Beam, 501B Open-Weight Model for Coding and Reasoning

Beam’s development involved major investments in both pretraining and reinforcement learning. The model was pretrained on 23.8 trillion diverse, curated, high-quality tokens from web and proprietary licensed datasets. Reflection conducted what it describes as one of the largest scale RL runs by any open lab, deploying 10.5K NVIDIA GB300 GPUs for four weeks to generate over 100 million rollouts with a maximum context length of 256K tokens. The company sourced approximately one million high-quality coding, agentic, and STEM environments for training.

On coding and agentic tasks, Beam is competitive with larger open models like GLM 5.2 and approaches Qwen 3.8-Max performance. The model’s key advantage is inference time efficiency—it achieves scores comparable to GLM-5.2 while using 3–4× less inference compute. Compared to models in the 2 trillion+ parameter family like Qwen 3.8-Max, efficiency gains are even more pronounced.

Reflection developed algorithms to maintain stable learning during high-compute RL training, addressing challenges like policy staleness in asynchronous policy gradient methods. The company also implemented a controllable length penalty during training that rewards successful solutions while discouraging unnecessary tokens. Users can adjust this reasoning effort parameter to balance response length against compute budget.

During RL training, Beam demonstrated capability transfer beyond its training tasks. The model showed consistent gains in web browsing despite the absence of browsing tasks from the RL training mixture. When given web access, it organically learned to search and query other large language models, and to use OCR APIs to read documents. Throughout the RL run, capabilities continued improving with increased compute, with no sign of saturation.

Key facts

  • Beam has 501 billion total parameters with 23 billion active parameters
  • Trained on 23.8 trillion tokens from diverse, curated sources
  • RL training used 10.5K NVIDIA GB300 GPUs for four weeks, generating over 100 million rollouts
  • Achieves comparable performance to GLM-5.2 while using 3–4× less inference compute
  • One million high-quality environments sourced for RL training
  • Model weights and technical report to be released later this month

Sources

← All posts