perhaps a representation of movement rather than a series of frames would better model what people see when watching video - highly compressed formats might even be directly useful in this regard?