Remix.run Logo
bluegatty 3 days ago

I see your 'fine point' but I don't think it holds - 'distillation' is a perfectly reasonable term to describe the process of creating outputs from one model to that expose key training element, to use in another model.

I think where the definition may be be invalid, is in the creation of 'unrelated data sets for training' models, for unrelated issues.

Creating training sets that mach a models core training, is definitely distillation, it does not have to expose the reasoning traces.

Synthesizing data for some arbitrary thing ... I'm not sure that would be the same thing.

It's hard to draw the line.

But the Chinese models are absolutely distilling - and would not be competitive without this distillation.

At the same time, there's a lot of real innovation and regular building going on at the same time over there.

HarHarVeryFunny 3 days ago | parent [-]

No - you can't distill if what you are given doesn't have the thing in it that you want to distill out of it.

I don't know why it's so important to you to use the word "distillation", but it's the wrong word to use.

BTW OpenAI on twitter also said that Kimi 3 "cannot be explained away by distillation or anything like that". The timeline of how long it takes to train a model and when Fable was released don't even line up. This is just Anthropic as usual trying to manipulate the US government into helping them shut down competition.

bluegatty 3 days ago | parent [-]

Distillation is absolutely - and uncontroversially - a valid term for what is happening here.

This isn't really a debate, I'm not making a fine point - just check with all of the various defintions of the term.

Moreover - the 'reasoning traces' are not required for distillation at all.

Finally - it's entirely possible for them to have used Fable for later stage fine tuning.

It's fair to be skeptical of Anthropic (and everyone else) - but this is 'distilling'.

HarHarVeryFunny 2 days ago | parent | next [-]

Words have meaning - you cant just redefine them because you want to.

Are reasoning traces required for distillation? Well they are if what you are trying to distill is reasoning, such as coding expertise.

Do you need reasoning traces for "LLM as judge"? No, but it would be highly perverse to call that distillation when there is a more accurate name for it - LLM as judge.

If you want to call use of Anthropic's redacted model outputs in any fashion that violates their terms of service (using them them to help develop anything that competes with Anthropic) as "distillation" then I can't stop you, but it reduces their claims to a joke.

Finally, as noted, OpenAI (who are just as anti-Chinese as Anthropic) said that Kimi 3 can't be explained via distillation (even true distillation!!), or even "anything like it". But random internet guy, you, disagrees. OK.

throw10920 a day ago | parent [-]

> Words have meaning - you cant just redefine them because you want to.

You are redefining words. The consensus among people who work in this space is that "distilling" is what's actually going on here.

> who are just as anti-Chinese as Anthropic

Conflating criticism of IP theft with being "anti-Chinese" is a standard PRC influence playbook technique.

And furthermore, OpenAI's market strategy is to win through regulatory capture. They are financially incentivized for Anthropic to be distilled by PRC labs and to be undercut by open models. Their claim about Kimi not being explainable due to distillation is not a factual claim - it's marketing from a company owned by Sam Altman.

Although, it does conclusively disprove your claim about the meaning of distillation, because you cannot say that "Kimi can't be explained by distilling" unless the consensus definition of "distillation" is such that it could be done on the summarized reasoning traces that Anthropic models expose.

HarHarVeryFunny a day ago | parent [-]

> you cannot say that "Kimi can't be explained by distilling" unless the consensus definition of "distillation" is such that it could be done on the summarized reasoning traces that Anthropic models expose

You can interpret it as you choose, but a much more obvious reason he [OpenAI's Dean Ball] might say it can't be distilled is because it can't be distilled. You can't distill alcohol out of orange juice.

throw10920 18 hours ago | parent [-]

> because it can't be distilled

...and, as everyone in the frontier labs knows, this is a lie, because that's not how distillation is defined.

I know that I won't convinced you, because you're quite possibly a PRC agent, but for all the other HN readers coming to this thread in the future to look at this failure of propaganda: just ask a model.

User: according to standard LLM lab parlance, can you "distill" one model from another if the model being distilled from does not expose a thinking trace?

GPT-5.6 Sol: Yes. In standard LLM terminology, you can distill one model from another even if the teacher model does not expose a chain-of-thought or "thinking trace."

Sonnet 5: Yes. "Distillation" broadly means training a student model to replicate a teacher model's outputs (or output distribution), and this doesn't require access to the teacher's chain-of-thought.

That's all she wrote. You're lying, and even the models know it. If you want to continue to discredit your account, go ahead :)

HarHarVeryFunny 6 hours ago | parent [-]

When your mom tells you it's time to go to bed, do you respond "You're lying! You're a communist! Boo hoo, I'm going to tell dad!"

Just curious.

throw10920 5 hours ago | parent [-]

Keep discrediting your account. I know that I can't convince a propagandist, but you just keep on further damaging your own reputation and argument for other HN readers every time you respond like this and ignore evidence that I've linked :)

throw10920 a day ago | parent | prev [-]

The GP (HarHarVeryFunny) is not operating in good faith. They've repeatedly lied about my own words to me, and are making up definitions that people who actually work at a frontier lab would disagree with.