| ▲ | OtherShrezzing 2 hours ago | |
> A model you can run on a loptop is simply not going to work as well as it's needed for programming The models you can run on a high-spec laptop today are approximately where frontier models were 12-18mo ago (albeit at a lower tok/s rate). If you scan back through hn comments from that era, you’ll find plenty of people saying “this is powerful enough to massively increase my productivity”. | ||
| ▲ | anon373839 an hour ago | parent [-] | |
> albeit at a lower tok/s rate Not always! I get 80-100 tok/s from Qwen 3.6 35B-A3B on a MacBook Pro thanks to MTP. With long contexts that dips to around 50-60. However, prefill is much slower than API models. So it becomes really, really, really critical to not have cache misses. | ||