| ▲ | purpleflame1257 3 hours ago | ||||||||||||||||
There's a real hole here at Q3. A critical breakpoint here is sub 16-GB cards, which covers the 5080, 5070 Ti, 5060ti, and several other cards from this generation and the last. It would be instructive to see where the quality knee is. | |||||||||||||||||
| ▲ | civvv 3 hours ago | parent | next [-] | ||||||||||||||||
Running Q3 on my AMD RX 9070XT. 32k context and 32/TPS. Apart from the context window preventing it from doing any large tasks, this thing is seriously powerful. I could probably push it to 64k context. Local open models are the future, and I am definitely getting a more powerful card. Very fun! | |||||||||||||||||
| |||||||||||||||||
| ▲ | dofm 2 hours ago | parent | prev | next [-] | ||||||||||||||||
There is an interesting new dynamic 3 bit quantisation I have been meaning to test: https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF Luke of Luke’s Dev Lab on YouTube had a look at it. It seems to outperform the typical 3-bit quantisation but whether it outperforms the new Unsloth dynamic I don’t know. | |||||||||||||||||
| ▲ | selectodude an hour ago | parent | prev | next [-] | ||||||||||||||||
I have a 5080, three OpenAI Pro token resets, and I’m on paternity leave. Astra seems pretty clever. Maybe I’ll give it a task. | |||||||||||||||||
| ▲ | jadbox 3 hours ago | parent | prev [-] | ||||||||||||||||
Q3 XL and Q3 XS are the two I'm trying to decide on | |||||||||||||||||
| |||||||||||||||||