| ▲ | petu 3 hours ago | |
> and it needs 10 times more RAM. More like 3-6. Qwen 27B full quality is FP16. So 54GB. In practice most inference providers would serve FP8, so 27GB. DeepSeek V4 Flash in full quality is mostly FP4. ~167GB official release. So Deepseek has 140GB model size overhead... which is shared between 100s of users single inference node serves, so not even a gigabyte of VRAM per user. Memory required for 200K of context per user: | ||