| ▲ | jorgeleo 11 hours ago | |
Same thing for me. on an M5 max, Qwen 3.6 35b gives me between 150 and 200 tps using splash as inference engine. More than enough for guided code sessions, at 100% privacy. And i can use obliverated models if i am trying to harden my own app, something i cannot do with cloud providers. | ||
| ▲ | Phemist 5 hours ago | parent [-] | |
Cool did not know about Splash. Seems interesting! https://github.com/incoai/splash/issues/38 Looks like an issue exists to convert model weights for ornith1.5 as this is a magical process atm. | ||