Remix.run Logo
pvillano 3 hours ago

Imagine yourself the CEO of a big AI company. It takes about a month to develop and train a model, so you release a new model every month. A startup says they can 10x your efficiency. What does that get you? You can't release a new model every three days. You can't 10x R&D either. You definitely can't tell investors that you are growing at the same rate, but selling off assets and cancelling purchasing contracts. So you just never improve efficiency enough to use less energy than the previous model version.

I don't believe this is actually happening.

impossiblefork 31 minutes ago | parent | next [-]

So you can make a huge internal model that you can then distill from?

alex_duf 2 hours ago | parent | prev | next [-]

A 10x reduction in pre-training means a 10x faster feedback loop. I'm pretty sure any lab would sign up for that. You can start experimenting on different approaches much more aggressively.

enzyme1234 3 hours ago | parent | prev | next [-]

the obvious answer is that you can try more experiments over the same period of time, so you find more improvements per month, and the rate of improvement of the models you release increases

awestroke 3 hours ago | parent | prev | next [-]

> What does that get you?

Cheaper model training runs? Ability to scale training to larger model sizes without extending training time?

pvillano 3 hours ago | parent [-]

Yes, but a 10x larger model is only marginally better, and 10x cheaper training runs is only useful if you can find a use for 10x as many.

aaronblohowiak 3 hours ago | parent [-]

more experiments.

BoorishBears 3 hours ago | parent | prev [-]

This seems weirdly pessemistic: frontier labs have much stronger pretraining than most open weights models

And reading the release it feels very obvious this is also a ton of aligning their data mix with coding and science: we don't know that this model doesn't have terrible world knowledge or is ruined for anything related to subjective preference

They also repeatedly mention knowledge almost as if they saw that skepticism coming, but then limit knowledge to topics where more understanding of how code/scientific writing looks would produce the same graph as having actual world knowledge maintained.

That's not nefarious (they literally build coding models), but it also means the resulting model isn't necessarily competitive with a frontier model in a broader way.

This feels like the inverse approach to what Thinking Machines did with Inkling (trying to train as "un-spikey" a base model as possible)