| ▲ | popalchemist 3 hours ago | |||||||
Image edit models can probably do a grid, but the temporal accuracy / coherence will never match what a video model, which is really a world model, can do. | ||||||||
| ▲ | armcat 2 hours ago | parent [-] | |||||||
Regarding your world model statement. This is completely FALSE. Learning the visual statistics of a physical world is NOT the same thing as learning its causal dynamics. The difference is observational likelihood versus intervention-dependent dynamics. There have been great studies disproving video models as world models, like this ICML paper: https://proceedings.mlr.press/v267/kang25g.html. Unfortunately lot of people treat them as world models, mostly because of their ability to reproduce increasingly convincing physical behaviour without ever discovering the underlying physical laws. This is due to many things that I could write an essay about, but better conditioning, latent space represtnation, scaling etc, all make them look awesome. I can still get absolutely insane results with MiniMax H3 - insane in the sense that it would not make sense at all and would make your head spin. | ||||||||
| ||||||||