Remix.run Logo
▲ simoiacos 3 hours ago

Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.

I'm also looking into expanding the protocol and the engine to support various steering techniques.

https://github.com/simoneiacomino/xenolith

▲aziis98 27 minutes ago | parent | next [-]

Just tried this on my Intel Ultra 7 255H, I also only have an iGPU. This does ~22tps! Love this.

I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR.

On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.

I'm pretty sure 2027 will be a very interesting year for local models and inference.

▲simoiacos 2 minutes ago | parent [-]

Please open a PR! I was too conservative with the supported devices.

If your GPU supports XMX we could also explore using it to improve the prefill kernel, but I don't have the hardware to test it myself.

▲ilaksh an hour ago | parent | prev [-]

I wish someone would add Intel support to ds4. And also improve AMD support.

Maybe Intel and AMD should help them with that.

▲simoiacos an hour ago | parent [-]

Yeah I see the value but I built Xenolith to target smaller models.

I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone's specific needs. Do you have a specific machine/spec in mind?

▲ilaksh an hour ago | parent [-]

The recent Intel GPU/AI cards. Really the same type of models as ds4