| |
| ▲ | bitexploder a day ago | parent | next [-] | | If you are really invested and have some system RAM you could get a 3-4 bit quant Qwen 3.5 35B-A3B running. There are builds that do expert caching, keeping the hot experts in cache. For something like disassembly, you're looking at being able to fit, if you have, say, 11 to 12 GB of VRAM, you could get at least three hot experts. For pure disassembly tests, I would say that would be pretty fast. A 4-bit quant is pretty decent and maintains most of the smarts of the larger quants. Depending on the GPU I would expect a decent token rate. It is medium strength local model, but if you harness and ground it well I expect it can reconstruct C code for you. The quality of your disassembler will matter here. If you have a lot of system RAM you could technically run Qwen Flash Next. On a 4080 with 16GB of RAM and 128GB of DDR5 I get ~35-40 t/s. And it is very capable. | |
| ▲ | pizza234 a day ago | parent | prev | next [-] | | I've been doing this type of work, and the answer is "yes and no". For autonomous work, even Qwen3.8-Flash-Next stumbles, although it does work to an extent. Qwen3.8-27b is useless. They're also slow, even on consumer systems with 24/32 GB VRAM. For generic help, I haven't tried, but I definitely wouldn't want a model that misleads me or takes a very long time to answer while I'm focused. Frontier models do this type of work without problems, both much faster and much more precisely, which makes local LLMs a waste of time and/or money. | | |
| ▲ | bitexploder a day ago | parent [-] | | I absolutely do not want a public provider having any of my data for reverse engineering work. That is a hard pass from me. Also, there exists a $750 GPU (V100) that can run 4-bit 27B quant at >90 t/s. And I find it far from useless. It is not the most capable model, but when you just need to offload and rip through assembly and you have chores batched up, it's pretty good. I use Qwen Flash Next at a 3-bit quantization, point it at disassembly with goals, put it in a harness with auto-compaction and a loop, and let it rip. Sometimes I wake up, and it’s just hilariously off. Other times, it completely accomplished the goal. I have one Qwen Flash Next 3.8 running right now, and 2x27B on a 4bit quant as workers, and they stay busy. This was not possible with local models on this level of hardware even two months ago. I have Qwen Flash Next at >100 t/s. Things have never been better for local models. | | |
| ▲ | pizza234 8 hours ago | parent [-] | | Parent's use case is for (abandoned) DOS programs, which hardly qualify as "personal data". Having said that, sure, if one has no options, everything left is "pretty good". I've left Qwen38-27b to reverse a tiny DOS program (few hundred bytes), and after more than an hour it was still struggling with debugger traces, misinterpreting basic DOS calls, and had produced no finished analysis. Possibly after a few hours it may have succeeded (surely with mistakes to find and correct), but then it'd look like a monkey at a typewriter more than else. Qwen3.8-Flash-Next is another level for sure, and it's a significant milestone for local LLMs IMO, since it can run on midrange GPUs, as long as there is a relatively large amount of system RAM (still not cheap). It's quite fast, although it also need to be taken into account that it's just moderately intelligent - if you observe the CoT while reversing, you'll find that struggles, performing many unproductive actions as well. |
|
| |
| ▲ | capnjngl a day ago | parent | prev [-] | | I don't know about that specific use case, but llmfit is your friend here https://github.com/AlexsJones/llmfit |
|