Remix.run Logo
purpleflame1257 3 hours ago

There's a real hole here at Q3. A critical breakpoint here is sub 16-GB cards, which covers the 5080, 5070 Ti, 5060ti, and several other cards from this generation and the last. It would be instructive to see where the quality knee is.

civvv 3 hours ago | parent | next [-]

Running Q3 on my AMD RX 9070XT. 32k context and 32/TPS. Apart from the context window preventing it from doing any large tasks, this thing is seriously powerful. I could probably push it to 64k context. Local open models are the future, and I am definitely getting a more powerful card. Very fun!

Forgeties79 2 hours ago | parent | next [-]

What are you offloading to ram (or even CPU)? I’m using a 9080 (not XT) and having trouble with context/token rates

slim 3 hours ago | parent | prev [-]

Running Q3 on 5060ti with 64k context. It runs great

dofm 2 hours ago | parent | prev | next [-]

There is an interesting new dynamic 3 bit quantisation I have been meaning to test:

https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

Luke of Luke’s Dev Lab on YouTube had a look at it. It seems to outperform the typical 3-bit quantisation but whether it outperforms the new Unsloth dynamic I don’t know.

selectodude an hour ago | parent | prev | next [-]

I have a 5080, three OpenAI Pro token resets, and I’m on paternity leave. Astra seems pretty clever. Maybe I’ll give it a task.

jadbox 3 hours ago | parent | prev [-]

Q3 XL and Q3 XS are the two I'm trying to decide on

dofm 2 hours ago | parent [-]

You might want to test this new dynamic GGUF:

https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

(I don’t know much about it, just saw a YouTube video about it last night)

kennywinker an hour ago | parent [-]

Another one to try:

https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash-GGUF

Runs the 3bit model faster than the 2bit one runs on my old-ass card. Can’t vouch for its intelligence yet, but i suspect whatever loss in smarts it takes is made up for by the extra resolution.