| ▲ | glimshe 3 hours ago | |||||||
Are there local models that can run on 8-12GB GPUs that can help reverse engineer retro software (DOS games and applications)? | ||||||||
| ▲ | bitexploder 11 minutes ago | parent | next [-] | |||||||
If you are really invested and have some system RAM you could get a 3-4 bit quant Qwen 3.5 35B-A3B running. There are builds that do expert caching, keeping the hot experts in cache. For something like disassembly, you're looking at being able to fit, if you have, say, 11 to 12 GB of VRAM, you could get at least three hot experts. For pure disassembly tests, I would say that would be pretty fast. A 4-bit quant is pretty decent and maintains most of the smarts of the larger quants. Depending on the GPU I would expect a decent token rate. It is medium strength local model, but if you harness and ground it well I expect it can reconstruct C code for you. The quality of your disassembler will matter here. If you have a lot of system RAM you could technically run Qwen Flash Next. On a 4080 with 16GB of RAM and 128GB of DDR5 I get ~35-40 t/s. And it is very capable. | ||||||||
| ▲ | pizza234 2 hours ago | parent | prev | next [-] | |||||||
I've been doing this type of work, and the answer is "yes and no". For autonomous work, even Qwen3.8-Flash-Next stumbles, although it does work to an extent. Qwen3.8-27b is useless. They're also slow, even on consumer systems with 24/32 GB VRAM. For generic help, I haven't tried, but I definitely wouldn't want a model that misleads me or takes a very long time to answer while I'm focused. Frontier models do this type of work without problems, both much faster and much more precisely, which makes local LLMs a waste of time and/or money. | ||||||||
| ||||||||
| ▲ | capnjngl 2 hours ago | parent | prev [-] | |||||||
I don't know about that specific use case, but llmfit is your friend here https://github.com/AlexsJones/llmfit | ||||||||