| ▲ | zmmmmm 9 hours ago | |
> It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. I'm confused, videos contain images and audio ...? | ||
| ▲ | ibotty 9 hours ago | parent | next [-] | |
That's most likely a disagreement on terms. In the media world, video is only the moving images, not audio. This is separate from images, that are meant to be still images. | ||
| ▲ | EricBurnett 6 hours ago | parent | prev | next [-] | |
Video contains images (frames), but not every image would reasonably be found in the frames of a video, or interpreted spatially. In the space of world model synthesis, consider blueprints, relationship diagrams, pages of instructions, sheet music, or a boarding pass. | ||
| ▲ | PxldLtd 9 hours ago | parent | prev [-] | |
It's more a comment about the feature detection I think; all image, video and audio input contribute to the same weights/activations that can produce image, video and audio output. | ||