Remix.run Logo
▲ halJordan a day ago

I disagree. Sure let them play and see if they can improve. But this model has more compute and more training data than the predecessors it fails to surpass. That only means their training regime is inferior if their predecessors did so much more with so much less. That inferiority should not be encouraged.

▲kelnos a day ago | parent | next [-]

You don't just magically do better than everyone else on every metric on your first go at something. Doing worse than others and refining is how pretty much everything works.

▲halJordan a day ago | parent [-]

I think i captured that in my first (second?) sentence

▲janalsncm a day ago | parent | prev | next [-]

The reality is they trained a model and it looks worse on benchmarks than Qwen or GLM. I don’t see how sharing the weights hurts anyone? Even when Llama 4 came out and it was a dumpster fire, it didn’t affect me personally.

> That only means their training regime is inferior if their predecessors did so much more with so much less

Hard to imagine how that wouldn’t be the case. They probably missed the boat on distilling Claude (or their lawyers said no), they probably didn’t hire an army of math PhDs to write reasoning traces, they don’t have millions of DAUs in a coding agent to train from, and they probably have less money, less experience, fewer top tier researchers, and fewer resources for experiments. They are an underdog without a doubt.

None of that means they shouldn’t release their model.

▲halJordan a day ago | parent [-]

Them releasing the weights doesn't hurt anyone. It's the peanut gallery clamoring to put them onto the same pedestal as actual tier 1 companies simply because they aren't named openai or anthropic that is hurtful.

▲thinkcontext a day ago | parent | prev [-]

I wonder if that's an indication that they are not distilling which limits how good they can get.

▲halJordan a day ago | parent [-]

Openai, grok, and Anthropic aren't distilling. Theyre just second class. It's not a big deal, we just shouldn't be lauding them for being second class.

▲girvo 20 hours ago | parent | next [-]

> Openai, grok, and Anthropic aren't distilling

Says who? We know Grok does at the least. They admitted it openly.

▲thinkcontext a day ago | parent | prev [-]

Musk said under oath that they use distillation for Grok.

▲disgruntledphd2 18 hours ago | parent [-]

And it still sucks. My apologies to the Cursor team but that's just very very poor performance.

Alternative explanation is that the Chinese have far more technical talent than anyone else, along with the infra and capital to build out these models.

My money is on the latter explanation, tbh.