| ▲ | reissbaker 4 hours ago | |
I think it's pretty hard to hold that worldview: Anthropic couldn't ship a reasoning model until they copied DeepSeek R1's homework, and they've all copied DS-style super-sparse MoEs at this point too. | ||
| ▲ | yorwba 2 hours ago | parent | next [-] | |
With slightly different cherry-picking, you could equally well claim that DeepSeek couldn't ship a reasoning model until they copied the idea from OpenAI's o1-preview, and they also copied MoEs from Google Brain/Jagellonian University https://arxiv.org/abs/1701.06538 way back in 2017, too! But ultimately these were ideas floating around in the air, if one group hadn't done the experiment, someone else would have. | ||
| ▲ | chorizo 4 hours ago | parent | prev [-] | |
That’s a really good point. Folks really need to read the papers coming out of these Chinese labs. Every paper from the DeepSeek team has been a step change. | ||