| ▲ | liuliu 4 hours ago | |||||||||||||
It is a well-known trick, given that the timestep is between 0 to 1, you can slicing them at any resolution (1000, or 10000, give or take), and then keep a look-up table for modulation scale / bias etc for each. It is quite different from quantization and it is indeed lossless. It is also only applicable to diffusion models as only these operates at per-timestep. | ||||||||||||||
| ▲ | xienze 3 hours ago | parent [-] | |||||||||||||
So... why didn't the model ship this way to begin with? They just wanted to waste VRAM for fun? | ||||||||||||||
| ||||||||||||||