| ▲ | v9v 9 hours ago | |||||||
Interesting. Wasn't Deepseek's founder saying that they had explicitly decided not to focus on multimodal models at all and were going text-only because they believed it was enough to achieve AGI? | ||||||||
| ▲ | johndough 8 hours ago | parent | next [-] | |||||||
It was explicitly said that they are pursuing multimodal support. A quote from the meeting transcript: https://github.com/demo-zexuan/liang-wenfeng-investor-meetin...
Earlier, the following was said, which might match more what you had in mind.
It is difficult to tell who said what, since the speaker ids are missing. | ||||||||
| ||||||||
| ▲ | swiftcoder 7 hours ago | parent | prev | next [-] | |||||||
Worth noting that deepseek has had a separate vision-capable model for some time, which also powers their chat interface's vision mode | ||||||||
| ▲ | dakolli 8 hours ago | parent | prev [-] | |||||||
I think you're thinking of Dario saying this about image generation. | ||||||||