What are you offloading to ram (or even CPU)? I’m using a 9080 (not XT) and having trouble with context/token rates