| ▲ | gymbeaux 3 hours ago | |||||||||||||||||||||||||||||||||||||
What would happen to Nvidia, Anthropic, OpenAI, if tomorrow someone released an open weights model on HuggingFace that matched performance and accuracy of Opus 5 running locally on an RTX 5070? That won’t happen tomorrow, but it will likely happen someday… what’s the plan beyond “don’t be the one holding the bags?” | ||||||||||||||||||||||||||||||||||||||
| ▲ | jkahrs595 17 minutes ago | parent | next [-] | |||||||||||||||||||||||||||||||||||||
Workloads will inflate just as they have been. Remember when llm assisted development used to be good only for a function, then a whole file, then a handful of files, then a code base, then a full stack, etc etc etc. People will claim to have “enough” even though they already have the equivalent of last years capabilities locally. | ||||||||||||||||||||||||||||||||||||||
| ▲ | jimbo808 an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
There’s no reason to assume frontier-level intelligence eventually collapses all the way onto a midrange consumer GPU. In fact, there are quite a few reasons not to assume that (information-theoretic constraints, etc). | ||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||
| ▲ | notatoad 13 minutes ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
probably not all that much... the market would dip, just like every time a new open weights model gets announced. but hundreds of millions of people aren't going to immediately self-hosting their own models. the biggest winner in that scenario would be ai providers, who suddenly have a capable model that they can serve much more efficiently. and the incumbents have a whole lot of compute. wouldn't anthropic and openAI just start offering that open weights model at prices that nobody else could compete with? | ||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||
| ▲ | fooker an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
> on an RTX 5070 RTX 5070 prices go up ~N times. Nvidia makes more money because it's easier to make these things than it's to make a GB300. | ||||||||||||||||||||||||||||||||||||||
| ▲ | ColdStream 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
Those companies will be quick to copy the tech, inference cost would plummet and there is a greater chance that these companies could make it to solvency. At least in the short term. Long term it might not be so great as consume hardware catches up. | ||||||||||||||||||||||||||||||||||||||
| ▲ | lisplist an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
If you could run Opus 5 on a 5070 then the labs must have achieved RSI at that point | ||||||||||||||||||||||||||||||||||||||
| ▲ | martinald 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
Nothing would really change IMO? 99% of users don't have anything like a RTX5070 (mobile especially). Even if it did, it still doesn't make much economic sense running a model locally vs on a datacentre. For example, I managed to just about squeeze a Q2 quant of Qwen 3.7 27b on my 9070XT. I get around 60tps decode (slightly faster prefill). _but_ it uses 300W of power to do so. At UK electricity rates of 30c/kWh this works out at something like 42c/MTok. I can get far far better models on openrouter cheaper than that, plus I'm not horrendously constrained on context length. | ||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||
| ▲ | nl 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
It seems very very unlikely that an Opus 5 matching local model that runs on a 5070 will be released within the next 5 years (I don't want to say "ever"). If it does happen then NVidia will sell a lot of 5070s though! | ||||||||||||||||||||||||||||||||||||||
| ▲ | milkshakes 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
inference is the cheap part; training is expensive. what compute infrastructure would train this mythical magic model? | ||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||
| ▲ | drivebyhooting 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
Inference time scaling means whoever had the most compute has the highest intelligence model. | ||||||||||||||||||||||||||||||||||||||
| ▲ | d_sem 3 hours ago | parent | prev [-] | |||||||||||||||||||||||||||||||||||||
I guess I'd like to understand the technical reasoning on how you think an how an Opus 5 could over time fit on an RTX 5070. | ||||||||||||||||||||||||||||||||||||||