| ▲ | volf_ 12 hours ago | |||||||||||||
GLM 5.3 and all previous models don't have a vision encoder and can only accept text. Ox-Alpha can accept video and images, so unless Z-ai added a pretty good vision encoder for this model, I don't think so. My money is on Moonshot and this being Kimi K3.5. The measured tps and latency is in-line with K3's tps and latency from Moonshot. MiniMax M3.5 is also possible (but the MiniiMax provider is a lot more performant than the lab behind ox-alpha, so less likely). | ||||||||||||||
| ▲ | nylonstrung 11 hours ago | parent | next [-] | |||||||||||||
It would be stranger to me that Kimi switched to GLM's tokenizer than that GLM added multimodal like Kimi and Deepseek both did recently | ||||||||||||||
| ▲ | minimaxir 2 hours ago | parent | prev | next [-] | |||||||||||||
The other tell from the provider angle is capacity. Whoever is hosting Ox Alpha has a lot of capacity which narrows down a lot of the Chinese companies. | ||||||||||||||
| ▲ | Bolwin 12 hours ago | parent | prev | next [-] | |||||||||||||
Glm had made vision models in the past. Look up GLM 5v. The only question now is if it's 5.3v, 5.4/5.5 or a dedicated flash/vision model | ||||||||||||||
| ||||||||||||||
| ▲ | Almondsetat 12 hours ago | parent | prev [-] | |||||||||||||
DeepSeek literally just came out with the vision-enabled version of Flash v4 which was purely text based. Why would GLM not be able to do the same thing? | ||||||||||||||
| ||||||||||||||