| ▲ | anon373839 2 hours ago | |
How about proof that black-box distillation can deliver these results without a very sophisticated RL pipeline doing the heavy lifting? | ||
| ▲ | dannyw an hour ago | parent [-] | |
"Black-Box On-Policy Distillation of Large Language Models", Microsoft Research, https://aka.ms/GAD-project > 'GAD consistently surpasses standard sequence-level distillation, delivering superior generalization and achieving performance that rivals the proprietary teacher. These results validate GAD as an effective and robust solution for black-box LLM distillation.' No RL, although I'm a little bit surprised to see MS Research publishing a paper on distilling GPT5? | ||