| ▲ | tescreal 2 hours ago |
| I really want to see hard evidence of distillation before I buy into it. Seems like a lot of sour grapes over not having the sort of lead assumed. In this field, it has been shown repeatedly that leaps in performance come swiftly and without notice. |
|
| ▲ | hypfer 2 hours ago | parent | next [-] |
| FWIW, the way GLM-5.2 (and 5.3) talk is clearly claude, so it is for sure also trained using distillation. The metric used there is me screaming at my screen per operating hours. Does it matter? IMO not really. Weights are open after all. (Or.. soon at least for 5.3) |
| |
| ▲ | RataNova 8 minutes ago | parent | next [-] | | Half of the new open-source stuff on github is written by claude now, all the way from issues to docs. Models are just vacuuming up this dataset during pretraining, naturally picking up the tone. You don't even need direct distillation via api anymore when the whole internet has turned into one big snapshot of Anthropic's weights | |
| ▲ | dannyw an hour ago | parent | prev [-] | | With the amount of Claudish on the internet now, and in source code repositories (how many Claudish README.mds have you seen?), you don't have to make a single API call to end up with a model that talks like Claude. And critically, like contracts in general, Anthropic's terms of service is only binding upon the user/counterparty. So even if a company say specifically sought out 'claude-like' content, and claude code traces available on the internet, if they don't use the Anthropic platform there is no ToS claim. | | |
|
|
| ▲ | xienze 2 hours ago | parent | prev [-] |
| What would constitute evidence in your opinion? |
| |
| ▲ | anon373839 2 hours ago | parent [-] | | How about proof that black-box distillation can deliver these results without a very sophisticated RL pipeline doing the heavy lifting? | | |
| ▲ | dannyw an hour ago | parent [-] | | "Black-Box On-Policy Distillation of Large Language Models", Microsoft Research, https://aka.ms/GAD-project > 'GAD consistently surpasses standard sequence-level distillation, delivering superior generalization and achieving performance that rivals the proprietary teacher. These results validate GAD as an effective and robust
solution for black-box LLM distillation.' No RL, although I'm a little bit surprised to see MS Research publishing a paper on distilling GPT5? |
|
|