| ▲ | cjg007 2 hours ago | |
I've been using v4-flash without vision for this months — it's my go-to for code tasks. Now with vision, I'm wondering: if this model can do everything the text-only version does (plus see images), why keep the text-only one around? Is it just cost/latency? Or is there something text-only does better? | ||
| ▲ | bel8 an hour ago | parent [-] | |
Yes, it adds vision to the already capable text-only LLM according to DS: > This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. | ||