| ▲ | kingleopold a day ago | ||||||||||||||||
not to miss, future models will be more compute hungry too. Current hardware prices are still goin up and no it's not cheaper to run your AI for like %99 of the people because of lots of costs, it's not just hardware. | |||||||||||||||||
| ▲ | lardosaurusrex a day ago | parent | next [-] | ||||||||||||||||
I think this fails to take into account how many people are fine with "fast enough" vs "fastest". I've seen people happily use AI that takes several minutes to generate text or edit an image because to them they already aren't using their computer when they tell it to start; they just grab their phone and walk away and come back only to check in on it. I feel like people here and on other technology discussions -- although it's worse here -- don't seem to parse what being the minority means. They know they're one of the few to have access to such incredible hardware -- whether it be rented or purchased for way too much cash -- but they only see their own kin; their own ilk. They only compare themselves to the best. The reality is that nobody expects data centre speed nor power in their own home and are satisfied to just go "haha its thinking" and let their computer quietly tick in the background as opposed to paying outragious prices for subscriptions or hardware. | |||||||||||||||||
| |||||||||||||||||
| ▲ | ericd a day ago | parent | prev [-] | ||||||||||||||||
The per token costs plummet with more concurrents. A box that can do 100 tps at request depth 1 might be able to do 3000 tps at request depth 64. Less per thread, but massively more per GPU/joule/etc. That’s the economy of scale of running in a DC rather than locally that they were referring to. | |||||||||||||||||