| ▲ | devy 5 hours ago | ||||||||||||||||
Exactly, it burns the tokens 3000x faster, which means the budget ($$$$$$) runs out so faster it will stop super quick, not able to perform long-duration work. At 27B parameter size, the intelligence is not able to accomplish work within a short amount time. Consequently, it become not usable. | |||||||||||||||||
| ▲ | gerdesj 4 hours ago | parent | next [-] | ||||||||||||||||
I (we) run Qwen3.8-27B-FP8 on a DGX Spark box - that's roughly £4000 of hardware. I did benchmark it in various ways and it runs quite well but it is a quantised jobbie and 1.5k t/s is also rather faster than anything I can possibly hope to achieve. To run that model at those sorts of speeds is going to need some serious investment and you are going to have to pay for it. | |||||||||||||||||
| |||||||||||||||||
| ▲ | a012 2 hours ago | parent | prev [-] | ||||||||||||||||
Unusable is too stretch IMO, you can still use it in tiny tasks that’ll respond almost instantly | |||||||||||||||||