Remix.run Logo
▲ sharpshadow 5 hours ago

Response to DeepSeek’s technical paper and competition.

▲LPisGood 5 hours ago | parent | next [-]

Which paper are you referring to?

▲wg0 5 hours ago | parent | prev [-]

What's that in summary?

▲Wheen 4 hours ago | parent [-]

Not the person you're replying to, but judging by the emphasis on the cost of cached input tokens in the OP article, I'd guess it has to do with DeepSeek v4.1's KV cache efficiency. It uses <1000 bytes per token, so they're able to get 1M token context in under a GB.

Edit: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...