Remix.run Logo
▲ cubefox 4 hours ago

> It takes one rendered frame (a low dynamic range proxy of it, three lanes of Gaussian noise, the previous frame's output reprojected, and five conditioning scalars) and produces four f32 channels per pixel: an RGB residual and one temporal-blend logit.

> The temporal path is implemented, but in the demo: the network's history input lanes and its per-pixel blend logit drive a reprojected feedback loop (docs/frame.md). The dlss5vk tool runs single frames with no history, which is what the reference captures were made with.

From this I assume the network uses the (via motion vectors) reprojected previous frame in order to increase temporal stability, i.e. similarity over adjacent frames. But this isn't strictly necessary, and apart from it, DLSS 5 is a pure post-process filter. So you could apply it to an old animated CGI movie like Final Fantasy (2001) [1]. Which should make it look significantly more realistic, at the cost of some flicker or other temporal instability.

One could also apply it to still images, like old renders from Tomb Raider [2], where temporal stability is not a factor. The difference to conventional text-to-image models with a "make it photorealistic" prompt would be that DLSS 5 strongly adheres to the underlying geometry.

1: https://www.imdb.com/title/tt0173840/

2: https://www.tombraiderchronicles.com/images/artwork-high-res...