| ▲ | DFlash 2: Keep Drafting Parallel(inco.ai) | |
| 18 points by mike-the-brain an hour ago | 3 comments | ||
| ▲ | hypfer 14 minutes ago | parent | next [-] | |
Amazing tech > An agent writes in an afternoon what a chatbot writes in a month But can you just.. not. Your tech is so good, it speaks for itself. Don't ruin that. | ||
| ▲ | adefa 17 minutes ago | parent | prev | next [-] | |
I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark. | ||
| ▲ | verdverm an hour ago | parent | prev [-] | |
vllm PR for DFlash2: https://github.com/vllm-project/vllm/pull/52816 | ||