new | show | ask | jobs Github

koakuma-chan 5 days ago

What is VLM?

▲

pwatsonwailes 5 days ago | parent | next [-]

Vision language models. Basically an LLM plus a vision encoder, so the LLM can look at stuff.

▲

echelon 5 days ago | parent | prev | next [-]

Vision language model.

You feed it an image. It determines what is in the image and gives you text.

The output can be objects, or something much richer like a full text description of everything happening in the image.

VLMs are hugely significant. Not only are they great for product use cases, giving users the ability to ask questions with images, but they're how we gather the synthetic training data to build image and video animation models. We couldn't do that at scale without VLMs. No human annotator would be up to the task of annotating billions of images and videos at scale and consistently.

Since they're a combination of an LLM and image encoder, you can ask it questions and it can give you smart feedback. You can ask it, "Does this image contain a fire truck?" or, "You are labeling scenes from movies, please describe what you see."

▲

littlestymaar 5 days ago | parent [-]

> VLMs are hugely significant. Not only are they great for product use cases, giving users the ability to ask questions with images, but they're how we gather the synthetic training data to build image and video animation models. We couldn't do that at scale without VLMs. No human annotator would be up to the task of annotating billions of images and videos at scale and consistently.

Weren't Dall-E, Midjourney and Stable diffusion built before VLM became a thing?

▲

tomrod 5 days ago | parent [-]

These are in the same space, but are diffusion models that match text to picture outputs. VLMs are common in the space, but to my understanding work in reverse, extract text from images.

▲

vlovich123 5 days ago | parent [-]

The modern VLMs are more powerful. Instead of invoking text to image or image to text as a tool, the models are trained as multimodal models and it’s a single transformer model where the latent space between text and image is blurred. So you can say something like “draw me an image with the instructions from this image” and without any tool calling it’ll read the image, understand the text instructions contained therein and execute that.

There’s no diffusion anywhere which is kind of dying out except as maybe purpose-built image editing tools.

	▲	tomrod 5 days ago \| parent [-]
		Ah, thanks for the clarification.

▲

dmos62 5 days ago | parent | prev [-]

LLM is a large language model, VLM is a vision language model of unknown size. Hehe.