| ▲ | yorwba 3 hours ago | |||||||||||||||||||||||||
GRPO increases the likelihood of samples that are better than average, not just the single best, and decreases that of samples that are worse than average. This method doesn't even involve an explicit likelihood, so it's a completely different mechanism. A comparison with minibatch optimal transport is in appendix A.2 of the paper. | ||||||||||||||||||||||||||
| ▲ | zimablue22 3 hours ago | parent [-] | |||||||||||||||||||||||||
You're right about the original GRPO proposal, but there are simplified variants that do just use best of K sampling. GRPO (or GRPO like approaches) for diffusion/flow matching similarly can be likelihood free. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||