| ▲ | Tepix 2 hours ago | |
Unfortunately, the MI300X is an OAM module. The MI350P is the one you want: It's a PCIe card, but it has less memory: 144GB. Luckily, DeepSeek V4 Flash will run in 144GB too because it's 256 MoE exports are native MXFP4 quantized. | ||
| ▲ | WhitneyLand an hour ago | parent [-] | |
How do you figure that? When they just loaded the weights alone, it was taking 156GB in vLLM. After warm-up and adding a KV cache pool, it took over 200GB. And this implementation is already cutting down the 1M token context window you would normally get. | ||