| ▲ | htrp 4 hours ago | |||||||||||||||||||
> Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads. > Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training. Early access, no weights no tech details, just a sign up here for info | ||||||||||||||||||||
| ▲ | wronglebowski 4 hours ago | parent | next [-] | |||||||||||||||||||
I'm all for more open models, but talk is cheap and this is a rather pointless announcement without anything backing it up. Publish your weights and HF repo or shut up IMO. | ||||||||||||||||||||
| ▲ | zelphirkalt 4 hours ago | parent | prev | next [-] | |||||||||||||||||||
And also a "proprietary data set" hahaha... Probably just means they don't want to show it, and it is data, that either they shouldn't have, or that there is nothing special about their training data and it is just meant to sound like there is some secret ingredient, while there is none. | ||||||||||||||||||||
| ||||||||||||||||||||
| ▲ | Loquebantur 3 hours ago | parent | prev [-] | |||||||||||||||||||
> Beam is undergoing final red-teaming and evaluations. You can sign up here for early access to the model. > We will release the weights, technical report, model card, and developer artifacts later this month. | ||||||||||||||||||||