| ▲ | embedding-shape 11 hours ago | |||||||||||||||||||||||||
Update3, summarized from RTX6kPRO Discord: > The BF16 checkpoint doesn't exhibit any of the complaints [...] with our current quants, the models tend to choose the wrong logits sometimes [...] why we'll need a requant [...] We're not aware of any bugs in any runtimes themselves [...] We have two remaining things that we're trying to tackle: some people are reporting thinking being too hard to trigger (^^) , and others are saying it thinks too much. We've seen much more of the latter internally | ||||||||||||||||||||||||||
| ▲ | Lwerewolf 11 hours ago | parent [-] | |||||||||||||||||||||||||
Deleted earlier, didn't see you post, pasting here: /* Just started testing with the gguf (with gpu offload, m5 max 128gb), q4_k_m, running seemingly well. Speed is initially slightly faster than antirez/ds4 - decode tok/s in the 30s, prefill ~400-ish. Expected, given the slightly smaller size. Looks to be working fine, but too early to tell. Definitely likes to "think". */ Anyways, guessing that discord might be focusing on the nvfp4 stuff. I've noticed spelling mistakes in the thinking traces, tool calls have been fine so far. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||