| ▲ | nvme0n1p1 4 days ago |
| That doesn't really answer the question. What they do internally for training the next model is a separate issue. I'm talking about the models they offer publicly. Per the article, companies are dropping OpenAI+Anthropic (partly) because of costs. If distilling is so simple and easy, why doesn't OpenAI take this "quick shortcut" and serve a self-distilled model externally, so they can charge reasonable prices and stop bleeding customers? Surely they can at least match the Chinese labs' efficiency, right? Wouldn't more customers and less opex look good for the IPO? |
|
| ▲ | linkregister 4 days ago | parent [-] |
| Anthropic, OpenAI, GDM, and Meta spend more on training than other labs by an order of magnitude. If they felt safe reducing this spend they would. These labs fear getting outcompeted. |
| |
| ▲ | nvme0n1p1 4 days ago | parent [-] | | Again, I am talking about inference, not training. Please read. | | |
| ▲ | linkregister 4 days ago | parent | next [-] | | I indeed failed to understand your point. Isn't that what they already do with Claude Haiku, GPT-5.6-Terra, etc? | | |
| ▲ | nvme0n1p1 3 days ago | parent [-] | | That's the goal, but those smaller models don't match the price:performance of leading open-weight models, which is why companies are switching away (as explained in TFA). The open weight labs figured out some secret sauce that (so far) big name labs are unable to replicate, so instead of competing, they're going on the defensive with claims of distillation attacks. |
| |
| ▲ | linkregister 4 days ago | parent | prev [-] | | How do you think they fund training? This is just as asinine as insisting that drug manufacturers only price medications based on production costs. | | |
| ▲ | nvme0n1p1 4 days ago | parent [-] | | Less opex = more profits. They can use that money to fund training. I don't see how that could possibly be a bad thing. |
|
|
|