| ▲ | rao-v 2 hours ago | |
Native dflash support on day 1 helps a lot! High quality speculative decoding speeds up a lot of agentic work. | ||
| ▲ | cmrdporcupine 13 minutes ago | parent [-] | |
You're right. I'm getting ~33tok/sec w/ dflash on it, using my personal home-built-for-Spark inference engine (not vLLM or llama.cpp based) That's pretty respectable. Still working on optimizing and cleaning up before I push it. | ||