Remix.run Logo
Ask HN: (Why) Was the LLM breakthrough useful for images, audio, etc.?
4 points by rogerrogerr 21 hours ago | 3 comments

This has been bothering me for a while. I feel like I have a decent conceptual grasp of what LLMs are doing with written text. But it seems like they also unlocked a bunch of progress in understanding and generating images, audio, and video. I can’t twist my brain into understanding the connection.

Is the boom in generated non-text content also built on LLMs, or is it just correlated with it because a bunch of excitement drove investment into the industry? I’m hoping for an ELI-non-ai-but-cs-major, this has been bothering me for a while.

jerlendds 20 hours ago | parent | next [-]

VLM != LLM. Vision language models basically treat text tokens and image tokens the same. Post-training an LLM on images+text can improve its capabilities. Id recommend searching around the keyword VLM to find more resources on how multi-modal AI works.

- https://huggingface.co/blog/vlms

- https://en.wikipedia.org/wiki/Multimodal_learning

verdverm 19 hours ago | parent | prev [-]

transformers were first an image understanding technique, the text processing came later, it's all about the training data and gradient descent, and allegedly attention

laruss5 4 hours ago | parent [-]

It's actually the other way round - the Transformer architecture was introduced for text (machine translation) in "Attention Is All You Need" (2017). Vision Transformers, which apply it to images, came three years later in 2020: https://arxiv.org/abs/2010.11929