| ▲ | fluoridation 5 hours ago | |||||||||||||||||||||||||||||||
Interesting, I had assumed it'd be too large to fit. What quant and context size are you running? | ||||||||||||||||||||||||||||||||
| ▲ | tarruda 4 hours ago | parent [-] | |||||||||||||||||||||||||||||||
IQ3_XXS (~3.2 BPW). For me this is an option because my Mac studio is only used for serving LLMs, so I can afford to dedicate most of its RAM to this. I can run with 256k context and only uses ~117G, with the remaining (up to 125G which I can allocate to VRAM) being used for prompt caching and context checkpoints. I'm making my own quants, though the Vision-Exp version is outdated and won't work on llama.cpp master branch (I built it before llama added support): - https://huggingface.co/tarruda/DeepSeek-V4-Flash-0731-GGUF - https://huggingface.co/tarruda/DeepSeek-V4-Flash-Vision-Exp-... For the Vision-exp version, I also ran perplexity + KLD against the original MXFP4. Seems quite OK: https://huggingface.co/tarruda/DeepSeek-V4-Flash-Vision-Exp-... | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||