| ▲ | kamranjon 2 hours ago | |
In what way is it similar? | ||
| ▲ | in-silico 2 hours ago | parent [-] | |
This algorithm: sample a bunch of latents, train the model using the one with the lowest error. IWAE: sample a bunch of latents, weight the loss of training the model using each one by softmax(-error). For images and text where the errors have large variance, those weights become one-hot, yielding this algorithm. | ||