| ▲ | civvv 2 hours ago | |
Running Q3 on my AMD RX 9070XT. 32k context and 32/TPS. Apart from the context window preventing it from doing any large tasks, this thing is seriously powerful. I could probably push it to 64k context. Local open models are the future, and I am definitely getting a more powerful card. Very fun! | ||
| ▲ | slim 2 hours ago | parent | next [-] | |
Running Q3 on 5060ti with 64k context. It runs great | ||
| ▲ | Forgeties79 an hour ago | parent | prev [-] | |
What are you offloading to ram (or even CPU)? I’m using a 9080 (not XT) and having trouble with context/token rates | ||