| ▲ | karimf 3 hours ago | |
Practically ~20GB with KV cache > We quantize weights to ~4-bit, bringing the LM under 20 GB. We validated minimal to no degradation on agentic tasks under compression. https://www.reddit.com/r/LocalLLaMA/comments/1vkgsum/introdu... | ||