| ▲ | zimablue22 2 hours ago | |
Any of the Best of N papers that exploded in popularity after GRPO. Likelihood is not fundamental to the spirit of GRPO, any exploratory mechanism would work. That sequential LLMs have a step-wise probability is convenient but not critical to this approach (where rejection sampling is widely used in diffusion models). | ||
| ▲ | kamranjon an hour ago | parent [-] | |
Do you have an example? Would love to read one. | ||