| ▲ | stymaar an hour ago | |
> Somehow I expected inference engines are generic LLM runtimes that can execute any weight. In fact, this was close to be true until last year: almost every open model except DeepSeek had a very similar architecture that was pretty close to the GPT-2 one with very few variations on top (and sometimes an MoE architecture, which itself was a few year old at that point). But a year ago there's been a cambrian explosion, first in attention mechanism but also in a bunch of other directions, mostly coming from China, and now there's a very massive diversity today's space. | ||
| ▲ | anuj0456 42 minutes ago | parent [-] | |
yes, that is correct. after chinca came into picture the advancement in this field sky rockted | ||