| ▲ | akmarinov 4 hours ago | ||||||||||||||||
The author put in the numbers, but maybe you didn’t read them. 45 t/s a second is perfectly respectable especially with no limits and 24/7 uptime with very little power draw on the Studio. Luna is at around 100 t/s for comparison, but it’s a worse model than 5.3 Flash | |||||||||||||||||
| ▲ | mike_hearn 2 minutes ago | parent | next [-] | ||||||||||||||||
I think it's been pretty much proven by now that there are no cases where local inferencing is better than remote inferencing, unless absolute privacy is a hard requirement. The efficiencies that come with datacenter scale and hw can't be beaten. | |||||||||||||||||
| ▲ | sho 4 hours ago | parent | prev [-] | ||||||||||||||||
The joke is that macs are famously slow at prompt prefill and you are not getting anything back in 3 seconds, or probably even 30. Once they get generating, it can be acceptable, but the TTFT is horrendous. There's a ton of well-understood things Apple can and hopefully will do to massively accelerate every stage of this pipeline and hopefully they're hard at work implementing most of them for m7. | |||||||||||||||||
| |||||||||||||||||