| ▲ | msdz an hour ago | |
Congratulations on the launch, it looks like an impressive product and tool! Q: From my (very, very limited!) understanding, I’m under the impression that part of the “inference engine inertia” is model- or at least architecture-specific code for most, if not each new open-weight model coming out. Assuming I got that right, do you plan on supporting everything vLLM/llama.cpp can do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models? | ||
| ▲ | anerli 39 minutes ago | parent [-] | |
Yeah, generally being able to focus on specific architectures lets you optimize better for those. However models of the same family (for example Qwen 3.5/3.6/ some 3.8 models) share the same architecture, so you only need to optimize once and new models can use the same kernels. There's also shared algorithms and kernels that can be optimized once and used across different families, so it's a bit nuanced. We plan to support any model architecture that we believe is somewhere along or close to the pareto frontier. There's some model families that are outdated or more niche that we don't necessarily want to put our focus into. | ||