Remix.run Logo
stymaar an hour ago

> Somehow I expected inference engines are generic LLM runtimes that can execute any weight.

In fact, this was close to be true until last year: almost every open model except DeepSeek had a very similar architecture that was pretty close to the GPT-2 one with very few variations on top (and sometimes an MoE architecture, which itself was a few year old at that point).

But a year ago there's been a cambrian explosion, first in attention mechanism but also in a bunch of other directions, mostly coming from China, and now there's a very massive diversity today's space.

anuj0456 42 minutes ago | parent [-]

yes, that is correct. after chinca came into picture the advancement in this field sky rockted