Remix.run Logo
Show HN: VisionLaya: Jev with Vision capabilities(huggingface.co)
1 points by someguy101010 8 hours ago

This model makes calibrated, typed decisions about an image plus optional text. It answers choice, score and noul (yes/no probability) questions in one forward pass, with no text generation.

It adds image input to Laya by replacing Laya's ModernBERT encoder with SmolVLM-256M-Instruct. Laya's predict(state, questions) API, proper-scoring-rule training and temperature calibration are unchanged.

check out the live demo at https://huggingface.co/spaces/thaitea/laya-vision-demo

and the source

- https://huggingface.co/thaitea/laya-vision-smolvlm-256m - https://github.com/r33drichards/laya-vision