| ▲ | DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression(zartbot.github.io) | |||||||||||||||||||
| 127 points by mfiguiere 2 days ago | 10 comments | ||||||||||||||||||||
| ▲ | arikrahman 2 days ago | parent | next [-] | |||||||||||||||||||
I am very impressed with the KV Cache Compression work as well as the prefix cacheing making queries converge on practically free. | ||||||||||||||||||||
| ▲ | mmastrac 2 days ago | parent | prev | next [-] | |||||||||||||||||||
I've been working with an automatic incremental context compactor enabled and it's been surprisingly helpful. It was particularly effective with DS41f - I think I was running at an effective session length of 5M, with the model running around 300k-400k and it was holding on both speed and intelligence. TBH I also ran the 400tok/s preview and that was just nuts. I just let the thing compact over and over over the course of a day attacking a couple of tough problems | ||||||||||||||||||||
| ||||||||||||||||||||
| ▲ | vivzkestrel 2 days ago | parent | prev | next [-] | |||||||||||||||||||
404 on the blog page? https://zartbot.github.io/blog/ | ||||||||||||||||||||
| ||||||||||||||||||||
| ▲ | smy20011 2 days ago | parent | prev [-] | |||||||||||||||||||
Removed | ||||||||||||||||||||
| ||||||||||||||||||||