| ▲ | syntaxing 2 hours ago | |||||||||||||
I am so excited for Qwen 3.8 27B. It’s a shame how slow prefill (~3-400) is on a strix halo but it’s such a good model for agentic tasks. | ||||||||||||||
| ▲ | colingauvin an hour ago | parent | next [-] | |||||||||||||
Prefill is survivable if you cache well. But what kills me is the context. Qwen 27 needs a ton of room for KV Cache. I guess not an issue on a 128 GB Halo or Spark, but if you are running of consumer/prosumer GPUs it's miserable to be compacting every 120k tokens. | ||||||||||||||
| ▲ | tarr11 an hour ago | parent | prev | next [-] | |||||||||||||
What type of agentic tasks are you using it for (eg how complex)? | ||||||||||||||
| ||||||||||||||
| ▲ | LoganDark an hour ago | parent | prev | next [-] | |||||||||||||
I find that 35B-A3B is much easier to run on my M4 Max (both prefill and generation) | ||||||||||||||
| ||||||||||||||
| ▲ | CamperBob2 an hour ago | parent | prev [-] | |||||||||||||
How are you running it on a Strix Halo? The weights aren't out yet, are they? | ||||||||||||||
| ||||||||||||||