| ▲ | bitexploder an hour ago | |
I absolutely do not want a public provider having any of my data for reverse engineering work. That is a hard pass from me. Also, there exists a $750 GPU (V100) that can run 4-bit 27B quant at >90 t/s. And I find it far from useless. It is not the most capable model, but when you just need to offload and rip through assembly and you have chores batched up, it's pretty good. I use Qwen Flash Next at a 3-bit quantization, point it at disassembly with goals, put it in a harness with auto-compaction and a loop, and let it rip. Sometimes I wake up, and it’s just hilariously off. Other times, it completely accomplished the goal. I have one Qwen Flash Next 3.8 running right now, and 2x27B on a 4bit quant as workers, and they stay busy. This was not possible with local models on this level of hardware even two months ago. I have Qwen Flash Next at >100 t/s. Things have never been better for local models. | ||