Beam: Reflection's 501B open-weight model

(reflection.ai)

149 points | by Philpax 2 hours ago ago

38 comments

  • Ariarule an hour ago ago

    Always glad to see more open-weight models, but this caption on the 2nd demo image had me do a double-take: "Land or Water Generalization Experiment: We recreated the viral X puzzle by asking Beam to create a fixed 180×90 grid for longitudes -179° to 179° and latitudes -89° to 89°, with 16,200 points. This puzzle is a few days old, so could not appear in the training data, thus testing the model’s generalization. Beam gets 95.5% coverage right, putting us between Opus 5 (92.5%) and Fable 5 (97.8%), which shows how well it generalizes to novel new tasks."

    Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...

    • charlieyu1 5 minutes ago ago

      I don't think age of the puzzle even matters, all models have search capacities these days

    • extr an hour ago ago

      Yeah I remember when the original post about this came out. Def not recent. Though I think their point survives in that they didn't exactly RL on this.

  • htrp 2 hours ago ago

    > Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.

    > Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.

    Early access, no weights no tech details, just a sign up here for info

    • wronglebowski an hour ago ago

      I'm all for more open models, but talk is cheap and this is a rather pointless announcement without anything backing it up. Publish your weights and HF repo or shut up IMO.

    • Loquebantur 9 minutes ago ago

      > Beam is undergoing final red-teaming and evaluations. You can sign up here for early access to the model.

      > We will release the weights, technical report, model card, and developer artifacts later this month.

    • zelphirkalt an hour ago ago

      And also a "proprietary data set" hahaha... Probably just means they don't want to show it, and it is data, that either they shouldn't have, or that there is nothing special about their training data and it is just meant to sound like there is some secret ingredient, while there is none.

      • janalsncm 6 minutes ago ago

        Not sharing the data is pretty standard because 1) it tends to get the lawyers involved and 2) good data is critical for getting good results.

        Imo you can get better results with great data and generic modeling techniques than with incredible modeling techniques and crappy data. Because if you have crappy data, you won’t even know if your model is good because your evals will also be bad.

        This is why Anthropic is throwing a fit about the Chinese distillation “attacks”. Clean reasoning traces are gold.

  • NorwegianDude an hour ago ago

    Bigger and still worse than existing free Chinese models that are smaller? Open weight models are nice, but at this point it seems western models are very far behind Chinese ones, despite Chinese companies publishing a lot of their findings. I hope we get more open models and more providers, as being stuck with a model from China or US with no competition is risky.

    Google does do a great job with Gemma models. It's one of the few language models actually good at language. OpenAI's top closed models can't even write norwegian correctly.

    • mirekrusin 16 minutes ago ago

      It takes time / few iterations to get it right (and it's moving target), but yes, expensive trial, my personal feeling is that they went a bit too high, at the same time who knows, maybe good move – as they're saying RL didn't plateau. It feels like they had something like $100M budget for it?

  • onlyrealcuzzo an hour ago ago

    This appears to be larger than DeepSeek v4.1 Flash, more expensive to run, and worse on every measured metric.

    Am I missing something?

    • swiftcoder 28 minutes ago ago

      > Am I missing something?

      It's pretty clear from their framing ("Beam advances the Western open-weight frontier") that one of their main selling points is not being a Chinese lab.

      I can't imagine that mattering to many individuals, but I guess someone out there has a government contract that forbids the use of foreign models

      • htrp 26 minutes ago ago

        Reflection raised on the idea of creating the "American Deepseek Project"

      • mirekrusin 23 minutes ago ago

        Multiple independent approaches are cool and all but fully open source model training (datasets, pipeline, checkpoints) should be taking advantage of being open and share runs/budget between different entities.

    • dotancohen an hour ago ago

      We're still at the stage where every new entrant is welcome in my opinion. Doesn't need to be record-breaking upon initial release.

      • halJordan 18 minutes ago ago

        I disagree. Sure let them play and see if they can improve. But this model has more compute and more training data than the predecessors it fails to surpass. That only means their training regime is inferior if their predecessors did so much more with so much less. That inferiority should not be encouraged.

    • jstummbillig 44 minutes ago ago

      Apparently there is more to making good models than copying everything on the internet.

    • Centigonal 30 minutes ago ago

      new entrant in this weight class, US lab.

    • martini333 an hour ago ago

      Beam goes brrrr

  • drubs an hour ago ago

    I remember being in the room with pretraining day 1 to help monitor the training job launch. Watching this model train from day 1 has been an amazing experience!

  • hypfer 30 minutes ago ago

    Someone should name their next model "Workhorse" just for SEO reasons.

    It's interesting how the industry converged to this very term, given that very less work is being done by horses since quite a while.

  • aeetes an hour ago ago

    the performance chart puts the better open source models behind the fold making it seem like it outperforms them... but it doesn't! all for open source models but this announcement is misleading

    • brumbelow an hour ago ago

      Yes. All the link made me realize is that I should checkout Deepseek 4.1 flash

  • zopper an hour ago ago

    Open model that is not yet open or widely accessible via API. Primarily comparing to non-SOTA models like Inkling and GLM 5.2. Included comparison to GLM 5.3 and DeepSeek V4.1 Flash in the table, but not in the charts (I assume they would make them look bad). Also no results from AA Index or Arena.

  • keeganpoppen an hour ago ago

    very curious to see more about what kinds of hardware you can run this on and the perf. characteristics… on the face of it, it seems like optimizing for inference speed might(?) be good for running on smaller hardware, but i suppose it could be the other way around and it is actually much resource-hungrier for the number of parameters, etc. …

  • 38 minutes ago ago
    [deleted]
  • wg0 an hour ago ago

    Suppose I inherited a data center spanning several hundred acres full of GPUs and free electricity.

    Where do I get the data?

    I mean, this many models. They have to start somewhere.

    • petu an hour ago ago

      I guess public datasets on HuggingFace and some shadow libraries content is enough to start.

      e.g. fineweb dataset is 50TB https://huggingface.co/datasets/HuggingFaceFW/fineweb

    • lucrbvi an hour ago ago

      There are a lot of open-research on pre-training, post-training and RL data mixtures and sourcing.

      I recommend checking papers from Datalogy, Nvidia Nemotron, Ai2 (Ollmo, Tulu, ...) and the recent model from Aleph Alpha if you want to learn more.

    • altcognito an hour ago ago

      If you ask a model, they will generally tell you where to get data. Modern frontier models have the large advantage of having tens if not hundreds of millions of users providing use cases to train against to improve their responses.

    • ttul 38 minutes ago ago
    • konfusinomicon an hour ago ago

      forget the data....sell it and go live your life!

    • Hamuko an hour ago ago

      Get data from Claude. That's what the Chinese (allegedly) do.

      • kbwal7 44 minutes ago ago

        Note that this sort of distillation is NOT for pre-training data (which is tens of trillions of tokens). I think the allegations against Chinese companies by Anthropic is more so that they distill SFT data (which is good for post-training, but you still need a strong base model)

  • sharktheone an hour ago ago

    Am I the only one who thought of the BEAM VM after the first word of thee title?