| ▲ | Ask HN: (Why) Was the LLM breakthrough useful for images, audio, etc.? | |||||||
| 4 points by rogerrogerr 21 hours ago | 3 comments | ||||||||
This has been bothering me for a while. I feel like I have a decent conceptual grasp of what LLMs are doing with written text. But it seems like they also unlocked a bunch of progress in understanding and generating images, audio, and video. I can’t twist my brain into understanding the connection. Is the boom in generated non-text content also built on LLMs, or is it just correlated with it because a bunch of excitement drove investment into the industry? I’m hoping for an ELI-non-ai-but-cs-major, this has been bothering me for a while. | ||||||||
| ▲ | jerlendds 20 hours ago | parent | next [-] | |||||||
VLM != LLM. Vision language models basically treat text tokens and image tokens the same. Post-training an LLM on images+text can improve its capabilities. Id recommend searching around the keyword VLM to find more resources on how multi-modal AI works. | ||||||||
| ▲ | verdverm 19 hours ago | parent | prev [-] | |||||||
transformers were first an image understanding technique, the text processing came later, it's all about the training data and gradient descent, and allegedly attention | ||||||||
| ||||||||