Remix.run Logo
magic_hamster 2 days ago

To be honest, running Deepseek v4 flash 0731 is enough for most what I need, and I like its responses way more. It's crazy that I can run this in a Q8 quantization in a home setup. It feels and performs like a frontier model.

The only issue with relying on local models is when you need them to prompt other models, and you might need to offload or switch models constantly which adds significant overhead.

But when it all works, its truly awe inspiring.