Remix.run Logo
delichon 4 hours ago

If you want to believe that the success of Kimi is about distillation attacks, ignore this.

egeozcan 3 hours ago | parent | next [-]

I'd kindly suggest that we could also stop calling them "distillation attacks".

halJordan 2 minutes ago | parent | next [-]

[delayed]

reilly3000 3 hours ago | parent | prev | next [-]

Agreed. I think when it comes to light that Claude has been known to say “I’m DeepSeek” that everyone has had their hand in that cookie jar. Moreover, paying for API calls hardly seems like an attack; ToS violation to be certain but not in the same category of law as criminal activity like hacking.

idiotsecant 2 hours ago | parent [-]

wait, is there evidence of this? I've not observed it. It sounds like the kind of thing that I want to be true because it would be hilarious but that makes me suspicious.

petu an hour ago | parent | next [-]

LLMs can't reliably answer what they are (w/o getting that info from system prompt/tool call/etc), so yea.

https://xcancel.com/teortaxesTex/status/2026130112685416881

I think I've seen same happening with some European languages as well.

snovv_crash an hour ago | parent | prev [-]

Ask it in Chinese

__MatrixMan__ 2 hours ago | parent | prev [-]

Agreed, "distilled variants" might be more suitable.

overfeed 19 minutes ago | parent [-]

Calling them variants is also inaccurate when the pretrained base and architecture are completely different.

WhitneyLand 4 hours ago | parent | prev | next [-]

False dichotomy right?

Are Chinese labs impressively innovating? Clearly.

However this doesn’t rule out possible gains due to distillation.

I don’t know the degree of the latter but both things could certainly be true.

SirHackalot 3 hours ago | parent | next [-]

Didn't Anthropic train on our collective data just to sell it back to us for $100/month? On top of that, Apple is suing them over alleged IP and trade secret theft by ex-Apple employees. Hard to feel too sympathetic, and I’m not an Anthropic hater in particular…

DashAnimal 3 hours ago | parent | next [-]

That Apple lawsuit is against OpenAi, just for clarity

SirHackalot 2 hours ago | parent [-]

Wow, I should never comment first thing in the morning... Thanks for the correction, you’re right. I will see if I can still edit my comment.

hugopuybareau 3 hours ago | parent | prev [-]

The only data-related lawsuit Anthropic got was the books nah ? And they paid only a minor part as paid agreement compared to what they would have paid losing the trial

SirHackalot 2 hours ago | parent [-]

Yep, that second part of my comment was an article I read about OpenAI and my mind mixed it up with Anthropic. My mistake.

culi an hour ago | parent | prev | next [-]

Fable was available for a few weeks before Kimi K3 came out. If it was a distillation attack, then that's a truly groundbreaking technological feat to distill a model like Fable in 2 weeks

fnord123 3 hours ago | parent | prev | next [-]

Also possibly true: Anthropic is running Kimi locally in their hardware and "distilling" it.

an hour ago | parent | prev | next [-]
[deleted]
cma 3 hours ago | parent | prev | next [-]

If they can distill fable into a full model post training run in ~15 days without the real thinking traces, yet we know Claude chats degraded with the thinking traces removed (chat resume bug from earlier in the year they reported stripping thinking to shed load as being the cause of degradation), how big can this degree be?

api 4 hours ago | parent | prev [-]

“Distillation” is just indirectly pirating the largely pirated training data used to train the original model.

“You stole my warez!”

serial_dev 3 hours ago | parent [-]

If we do it, it's training a model. When they do it, it's distillation attack. - Anthropic

esafak 3 hours ago | parent [-]

"You are distilling what I have rightfully pirated."

Parfait__ 4 hours ago | parent | prev | next [-]

I stil don't understand them. I want the US to "win the AI race" but I have trouble understanding how most of all inventions today aren't "distillations" of past knowledge. Is Anthropic claiming the data they stole as trade secrets?

an hour ago | parent | next [-]
[deleted]
lukewarm707 2 hours ago | parent | prev | next [-]

i want china to win so that i get access to ai and not restricted and censored.

the chinese models are less censored, you'd better believe it.

try asking claude about its 'guardrails' (restrictions), very high chance anthropic will censor it.

fwip 4 hours ago | parent | prev [-]

Anthropic is claiming that training an LLM to mimic another LLM is materially different and worse than slurping up stuff written by humans (even if that material is stolen).

Basically, they want IP protection for Claude. This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.

blints 4 hours ago | parent | next [-]

Their claim is even stronger than that, they have complaints about their models being used as a validation step for other model output, which is standard practice in the industry.

koe123 3 hours ago | parent | prev | next [-]

How does anyone justify this? How can you argue this in good faith?

fwip 3 hours ago | parent [-]

Snarky answer: “It is difficult to get a man to understand something, when his salary depends on his not understanding it.”

Realer answer: A combination of the above, plus group/bubble effect of all your coworkers saying the same thing. You as a group, conflate a bunch of concerns together (China, no-guardrails-AI, etc), decide that your group will be the responsible stewards of AI, and then anybody "stealing your work" appears dangerous - both to your livelihood and to the human race.

Levitz 3 hours ago | parent | prev | next [-]

>This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.

No, it's perfectly reasonable once you get down to reality.

China is not going to care about IP. That's just a fact. So either nobody cares about IP (at the very last in this context) and any AI company can just do whatever with data, or Chinese companies have to be held up to scrutiny.

We don't have the privilege to be able to hold western companies to higher ethical, legal, and environmental standards and not risk competitiveness.

That there is a whole lot of people right now who insist on doing above and still somehow praise China at every turn is something historians or news outlets will have to make sense of in some 5 years time.

nemomarx 2 hours ago | parent | next [-]

Give up IP for end users and people will be pretty okay with giving up on IP for ai companies. You can't have different standards for special companies though.

idiotsecant 2 hours ago | parent | prev [-]

why should western AI companies be held to different standards than Chinese ones? Neither of them are your buddy.

creato 42 minutes ago | parent [-]

They shouldn't be. But the point is, they are. OpenAI and Anthropic are paying many rights holders for access to their data (reddit, NYT, etc.).

So distillation, among other things, allows Chinese labs to indirectly benefit from these arrangements at no cost to them.

verdverm 4 hours ago | parent | prev [-]

Google is apparently taking a different stance and offering distillation as a paid product

https://docs.cloud.google.com/gemini-enterprise-agent-platfo...

petu an hour ago | parent | next [-]

You don't get to take distilled model home, it all stays with Google.

It's "optimize your costs in our garden" product.

verdverm an hour ago | parent [-]

yup, strings are certainly attached when dealing with US Big Tech / Ai

I recommend Fireworks as an alternative

behnamoh 3 hours ago | parent | prev [-]

But nobody wants to distill Google's models, Gemini is really bad.

verdverm 3 hours ago | parent [-]

strong agreement, I've stopped using all closed weight models on principle, but the latest gemini models have increased hallucinations and now talk back, so double reason not to use them

Aurornis 4 hours ago | parent | prev | next [-]

You can’t build a frontier model with one single thing. This is an incremental improvement but it doesn’t explain the entire success of the model. The training set is immensely important, regardless of how you feel about distillation.

EGreg 3 hours ago | parent [-]

Reminds me of this btw:

https://www.bbc.com/news/technology-12343597

Microsoft replied that Bing uses “many different signals” —- including cribbing from Google :-)

krlx 2 hours ago | parent [-]

I remember 15 years ago or so, one of my first student job was to evaluate Bing results compared to the same query on Google. Didn't know then that I was a distillation attacker.

igleria 3 hours ago | parent | prev | next [-]

The distillation complaints to me sound like when a casino complains about card counting

jeremyjh 2 hours ago | parent | prev | next [-]

It can easily be both. Also, they didn't use this innovation in K3 - K3 pre-training would have started months ago and the paper only mentions a 48B model. The people working at this level may not even be heavily involved in shipping a new iteration of K3, or at least theory contributions to it were done many months or even a year ago and after that it is all engineering.

vikramkr 2 hours ago | parent [-]

This paper is from last year

rdtsc 3 hours ago | parent | prev | next [-]

Does one have to exclude the other?

moralestapia 4 hours ago | parent | prev [-]

Well said.

The distillation theory does not even make sense as Fable was only around for days (effectively) before Kimi was released.