| ▲ | aziis98 an hour ago | |
Just tried this on my Intel Ultra 7 255H, I also only have an iGPU. This does ~22tps! Love this. I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR. On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement. I'm pretty sure 2027 will be a very interesting year for local models and inference. | ||
| ▲ | simoiacos 41 minutes ago | parent [-] | |
Please open a PR! I was too conservative with the supported devices. If your GPU supports XMX we could also explore using it to improve the prefill kernel, but I don't have the hardware to test it myself. | ||