| ▲ | Lwerewolf 11 hours ago | ||||||||||||||||
Deleted earlier, didn't see you post, pasting here: /* Just started testing with the gguf (with gpu offload, m5 max 128gb), q4_k_m, running seemingly well. Speed is initially slightly faster than antirez/ds4 - decode tok/s in the 30s, prefill ~400-ish. Expected, given the slightly smaller size. Looks to be working fine, but too early to tell. Definitely likes to "think". */ Anyways, guessing that discord might be focusing on the nvfp4 stuff. I've noticed spelling mistakes in the thinking traces, tool calls have been fine so far. | |||||||||||||||||
| ▲ | Lwerewolf 7 hours ago | parent [-] | ||||||||||||||||
Well, just ran said gguf on the GeneralsX codebase with a pretty open-ended "Explain this codebase to me, and the general game loop." prompt, and...
...repeating forever.Since they mentioned that they're working on new quants, guess I'll wait. From earlier tests on work stuff, it's definitely capable. | |||||||||||||||||
| |||||||||||||||||