| ▲ | SparkyMcUnicorn an hour ago | |||||||
Yeah, I would appreciate if someone could make sense of the pricing differences between these models. How can a provider run DSv4F at lower cost than a 27B dense or 35B A3B model? Does it come down to utilization and/or specific model tricks and efficiencies (attention, kv cache, etc.)? DeepInfra prices: Qwen 3.6 27B: $0.32 in / $3.20 out Gemma 3 27B: $0.08 in / $0.16 out DeepSeek V4 Flash 0731: $0.08 in / $0.18 out Qwen 3.6 35B A3B: $0.10 in / $0.95 out https://openrouter.ai/qwen/qwen3.6-27b https://openrouter.ai/google/gemma-3-27b-it | ||||||||
| ▲ | mordae 32 minutes ago | parent [-] | |||||||
DeepSeek V4 Flash is natively FP4 MoE with very compact KV cache. Say 8 GB/s. Qwen 27B is about 60 GB/s at full FP16 precision. | ||||||||
| ||||||||