Remix.run Logo
zmmmmm 9 hours ago

> It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation.

I'm confused, videos contain images and audio ...?

ibotty 9 hours ago | parent | next [-]

That's most likely a disagreement on terms. In the media world, video is only the moving images, not audio. This is separate from images, that are meant to be still images.

EricBurnett 6 hours ago | parent | prev | next [-]

Video contains images (frames), but not every image would reasonably be found in the frames of a video, or interpreted spatially. In the space of world model synthesis, consider blueprints, relationship diagrams, pages of instructions, sheet music, or a boarding pass.

PxldLtd 9 hours ago | parent | prev [-]

It's more a comment about the feature detection I think; all image, video and audio input contribute to the same weights/activations that can produce image, video and audio output.