Remix.run Logo
▲ SubiculumCode 4 hours ago

Are any of these multimodal yet? I'd love to try asking a model with calibrated probabilities to answer question like, "do these shapes match?". Sure, you can ask a LLM....

▲necubi 3 hours ago | parent | next [-]

Cloudflare’s clef is multimodal (https://blog.cloudflare.com/clef-decision-models/)

(Disclaimer, I work at Cloudflare, but not on models)

▲sauhsoj 3 hours ago | parent | prev [-]

Strands Decider can take vision in. How does it go with that question?

▲Zopieux an hour ago | parent [-]

This is not advertised on their page, did you make this up?

I believe image classification/analysis by deciders (not just OCR, not everything is about text) is still lacking.

Cloudflare's Clef had fair results on my test, but it's larger and slower. Wondering about Strands.