| ▲ | dmarcos 9 hours ago |
| I believe so. This is not a model that generates pixels frame by frame from user input like genie 3. Instead, there’s an actual 3D scene / structure generated (point cloud, 3dgs) from the input images. |
|
| ▲ | 9 hours ago | parent | next [-] |
| [deleted] |
|
| ▲ | xyzsparetimexyz 6 hours ago | parent | prev [-] |
| Wrong, it does go straight to generating images. |
| |
| ▲ | dmarcos 5 hours ago | parent [-] | | From blog post: “It generates both image frames from novel views and explicit 3D outputs” Model can indeed generate novel views but I don’t think in real time (I could be wrong). If you want to navigate a space in real time as user above was asking gotta rely on the 3D output and the explicit representation will provide continuity. User likely was referring to the limitations we’ve seen on video-style world models like genie and precursors. |
|