Remix.run Logo
onlyrealcuzzo an hour ago

> Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers).

Can anyone working at one of the main US labs (Google, OpenAI, Anthropic) comment on WTF they haven't even tried MLA - despite the obvious massive advantages?

I know enough to know they aren't completely incompetent. So there must be a quite good reason.

But it remains a mystery to me.

DeepSeek's MLA is like almost 2 years old at this time. They've got thousands of people working on this stuff. They clearly have the ability to at least try it...

aabdi 18 minutes ago | parent | next [-]

They already are?

There’s a measurable performance tradeoff versus gqa so there’s reluctance.

For the most part though the new deepseek v4 tech is hca and mhc and people are still catching on like with moe and rl. Wait for 6 12 months, minimum time for next pre train.

ronsor an hour ago | parent | prev [-]

Are they not?

The big US labs are opaque and don't publish much of any technical details anymore. We don't know what they are or aren't doing, honestly.