Yes, it adds vision to the already capable text-only LLM according to DS:
> This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge.
https://api-docs.deepseek.com/news/news260821/