| ▲ | hadlock a day ago | |||||||
the important thing is that Qwen 3.7 27B will run unlimited jobs on my consumer grade laptop at 60 tokens/second for free, forever, in about 1-2 years | ||||||||
| ▲ | villish a day ago | parent | next [-] | |||||||
Thats only important if running it locally is critical for privacy reasons or just as a hobby. Time has a cost in business. If a model needs 30 million tokens to achieve a similar result as another that can do it in 10 million, that 60 tokens per second will take a long time. | ||||||||
| ||||||||
| ▲ | criley2 14 hours ago | parent | prev [-] | |||||||
It's not free. You're paying electricity and you're ignoring the cost of the hardware. Even on electricity alone, there are cloud providers who may beat your laptop on price per million tokens. Qwen 3.8 flash is interesting in this space. Not to say that there aren't other benefits of running models locally, I loaded Qwen 3.8 27B 6bit MLX just yesterday. | ||||||||