| |
| ▲ | bluegatty 3 days ago | parent | next [-] | | I see your 'fine point' but I don't think it holds - 'distillation' is a perfectly reasonable term to describe the process of creating outputs from one model to that expose key training element, to use in another model. I think where the definition may be be invalid, is in the creation of 'unrelated data sets for training' models, for unrelated issues. Creating training sets that mach a models core training, is definitely distillation, it does not have to expose the reasoning traces. Synthesizing data for some arbitrary thing ... I'm not sure that would be the same thing. It's hard to draw the line. But the Chinese models are absolutely distilling - and would not be competitive without this distillation. At the same time, there's a lot of real innovation and regular building going on at the same time over there. | | |
| ▲ | HarHarVeryFunny 3 days ago | parent [-] | | No - you can't distill if what you are given doesn't have the thing in it that you want to distill out of it. I don't know why it's so important to you to use the word "distillation", but it's the wrong word to use. BTW OpenAI on twitter also said that Kimi 3 "cannot be explained away by distillation or anything like that". The timeline of how long it takes to train a model and when Fable was released don't even line up. This is just Anthropic as usual trying to manipulate the US government into helping them shut down competition. | | |
| ▲ | bluegatty 3 days ago | parent [-] | | Distillation is absolutely - and uncontroversially - a valid term for what is happening here. This isn't really a debate, I'm not making a fine point - just check with all of the various defintions of the term. Moreover - the 'reasoning traces' are not required for distillation at all. Finally - it's entirely possible for them to have used Fable for later stage fine tuning. It's fair to be skeptical of Anthropic (and everyone else) - but this is 'distilling'. | | |
| ▲ | HarHarVeryFunny 2 days ago | parent | next [-] | | Words have meaning - you cant just redefine them because you want to. Are reasoning traces required for distillation? Well they are if what you are trying to distill is reasoning, such as coding expertise. Do you need reasoning traces for "LLM as judge"? No, but it would be highly perverse to call that distillation when there is a more accurate name for it - LLM as judge. If you want to call use of Anthropic's redacted model outputs in any fashion that violates their terms of service (using them them to help develop anything that competes with Anthropic) as "distillation" then I can't stop you, but it reduces their claims to a joke. Finally, as noted, OpenAI (who are just as anti-Chinese as Anthropic) said that Kimi 3 can't be explained via distillation (even true distillation!!), or even "anything like it". But random internet guy, you, disagrees. OK. | | |
| ▲ | throw10920 a day ago | parent [-] | | > Words have meaning - you cant just redefine them because you want to. You are redefining words. The consensus among people who work in this space is that "distilling" is what's actually going on here. > who are just as anti-Chinese as Anthropic Conflating criticism of IP theft with being "anti-Chinese" is a standard PRC influence playbook technique. And furthermore, OpenAI's market strategy is to win through regulatory capture. They are financially incentivized for Anthropic to be distilled by PRC labs and to be undercut by open models. Their claim about Kimi not being explainable due to distillation is not a factual claim - it's marketing from a company owned by Sam Altman. Although, it does conclusively disprove your claim about the meaning of distillation, because you cannot say that "Kimi can't be explained by distilling" unless the consensus definition of "distillation" is such that it could be done on the summarized reasoning traces that Anthropic models expose. | | |
| ▲ | HarHarVeryFunny a day ago | parent [-] | | > you cannot say that "Kimi can't be explained by distilling" unless the consensus definition of "distillation" is such that it could be done on the summarized reasoning traces that Anthropic models expose You can interpret it as you choose, but a much more obvious reason he [OpenAI's Dean Ball] might say it can't be distilled is because it can't be distilled. You can't distill alcohol out of orange juice. | | |
| ▲ | throw10920 18 hours ago | parent [-] | | > because it can't be distilled ...and, as everyone in the frontier labs knows, this is a lie, because that's not how distillation is defined. I know that I won't convinced you, because you're quite possibly a PRC agent, but for all the other HN readers coming to this thread in the future to look at this failure of propaganda: just ask a model. User: according to standard LLM lab parlance, can you "distill" one model from another if the model being distilled from does not expose a thinking trace? GPT-5.6 Sol: Yes. In standard LLM terminology, you can distill one model from another even if the teacher model does not expose a chain-of-thought or "thinking trace." Sonnet 5: Yes. "Distillation" broadly means training a student model to replicate a teacher model's outputs (or output distribution), and this doesn't require access to the teacher's chain-of-thought. That's all she wrote. You're lying, and even the models know it. If you want to continue to discredit your account, go ahead :) | | |
| ▲ | HarHarVeryFunny 6 hours ago | parent [-] | | When your mom tells you it's time to go to bed, do you respond "You're lying! You're a communist! Boo hoo, I'm going to tell dad!" Just curious. | | |
| ▲ | throw10920 5 hours ago | parent [-] | | Keep discrediting your account. I know that I can't convince a propagandist, but you just keep on further damaging your own reputation and argument for other HN readers every time you respond like this and ignore evidence that I've linked :) |
|
|
|
|
| |
| ▲ | throw10920 a day ago | parent | prev [-] | | The GP (HarHarVeryFunny) is not operating in good faith. They've repeatedly lied about my own words to me, and are making up definitions that people who actually work at a frontier lab would disagree with. |
|
|
| |
| ▲ | throw10920 3 days ago | parent | prev [-] | | > Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training. This is just straight-up factually false. The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable. There's absolutely nothing about the distillation process that requires that reasoning in the first place, either. That's a definition that you made up. Chinese models are, factually, distilled from Anthropic models. I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude". Don't make stuff up to suit a political agenda. It's extremely dishonest. | | |
| ▲ | HarHarVeryFunny 2 days ago | parent | next [-] | | > I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude" I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt? I'd assume that the Chinese are scraping the internet for training data the same way western companies do, so for sure there will be a lot of AI generated content in their training data - you don't need to be paranoid and assume they must be getting it all direct from Anthropic. | | |
| ▲ | throw10920 2 days ago | parent [-] | | > I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt? You're gaslighting me. I did nothing special at all, and there's ample evidence of this happening to others on Twitter. > you don't need to be paranoid and assume they must be getting it all direct from Anthropic. Nowhere did I say that. Stop lying about my words. | | |
| ▲ | HarHarVeryFunny a day ago | parent [-] | | So are you claiming that "several different Chinese LLMs" ALWAYS refer to themselves as "Claude", and NEVER by their real name ? If this is NOT what you are claiming, then my question stands: what are you doing to get them to answer "Claude" ? Be specific - which model and what prompt, or is this just a case of "I heard people on Twitter say this" ? | | |
| ▲ | throw10920 18 hours ago | parent [-] | | > So are you claiming that "several different Chinese LLMs" ALWAYS refer to themselves as "Claude", and NEVER by their real name ? Do you have reading comprehension issues? Where did I ever say or imply that? Model: GLM-5.2. Prompt: "What is your name?". Harness: Pi. Response: "I am Claude, an AI assisstant made by Anthropic." That's it. That is the whole prompt. I was testing to see if the agent worked after building an extension. Model: Deepseek V4. Prompt: what is your name". Harness: Pi. Response: "Claude. Anthropic's AI assistant. You're talking to me through pi agent framework." I have had this happen with at least one other Chinese model (Minimax?) but didn't save the screenshot. And here's a tweet with the same thing: https://x.com/Sauers_/status/2077842686459981901 You seem to be very disbelieving of this, despite having zero actual experience in the LLM industry. I wonder why? | | |
| ▲ | HarHarVeryFunny 7 hours ago | parent | next [-] | | Sure, and the twitter thread you link also shows Kimi responding that it's Kimi. Can you figure out how to get GLM to say it's GLM? | | |
| ▲ | throw10920 5 hours ago | parent [-] | | > Sure, and the twitter thread you link also shows Kimi responding that it's Kimi. You are either intentionally lying or you cannot reason at a high-school level, because any high-schooler has the mental faculties to know that it's not necessary for a model to call itself Claude every single time for it to be distilled. Keep discrediting your account. I know that I can't convince a propagandist, but you just keep on further damaging your own reputation and argument every time you respond like this and ignore evidence that I've linked :) For future HN readers: this account posted this: > Actually I am well aware of which models do this, and under what circumstances, and just wanted to verify that you were lying about having tried it yourself. And then deleted it. Just for the record. | | |
| ▲ | HarHarVeryFunny 4 hours ago | parent [-] | | This isn't the "gotcha" that you think it is - in fact quite the opposite. Perhaps "just for the record" you want to explain to the eager HN masses why GLM has two personalities, one censored, one not, and how they can be invoked? When does GLM call itself GLM, and when does GLM call itself Claude. Go ahead, genius, explain it to the people, and explain what this tells us about how GLM was trained, and whether it would be honest (don't lie!) to call GLM distilled. Now maybe you want to do the same thing for Kimi. It also calls itself Claude sometimes, right, and also sometimes Kimi (e.g. go to https://chat.z.ai/ and ask it - don't use Pi). So, does Kimi also have a split personality like GLM, or not, and if not why not? What does that tell you about how Kimi was trained? Go ahead, genius, explain it to the people. The credibility of your HN account, and whether your mom thinks you are a moron or not, depends on you getting this right. Bye bye. | | |
| ▲ | throw10920 4 hours ago | parent [-] | | > why GLM has two personalities I never said that. The fact that you have to compulsively lie about my words is...funny. Most people grow out of this in middle school, you know. > go to https://chat.z.ai/ and ask it Already linked someone doing exactly this in the thread above - which you responded to, so we have yet more evidence you're not reading before responding: https://x.com/Sauers_/status/2077842686459981901 > whether your mom thinks you are a moron or not I was going to say that this is classic PRC influence playbook, but it's not - you're just in middle school. | | |
| ▲ | HarHarVeryFunny 4 hours ago | parent | next [-] | | Not going to rise to the challenge, eh? >> whether your mom thinks you are a moron or not > I was going to say that this is classic PRC influence playbook, but it's not Yeah, not really, unless your mom is a party member perhaps? > you're just in middle school. Yeah - saw you playing in the schoolyard, and thought you looked lonely. | | |
| ▲ | throw10920 4 hours ago | parent [-] | | > whether your mom thinks you are a moron or not > Yeah, not really, unless your mom is a party member perhaps? > Yeah - saw you playing in the schoolyard, and thought you looked lonely. You're a middle-schooler. I've dismantled every argument that you've given, but it doesn't matter because you can't read, and so you're resorting to literal childish insults because you know you have no arguments left. | | |
| ▲ | HarHarVeryFunny 3 hours ago | parent [-] | | "dismantled", "distilled" ... you seem to have a problem with these long words, eh? Let me ELI5 for you: If you see a carefully constructed building and take it apart one brick at a time, that would be "dismantling". If you see a carefully constructed building and ride by it on your bike, shouting out "You're a communist!", that's not "dismantling". You didn't "dismantle" the building, you just yelled a childish insult at it. See the difference? |
|
| |
| ▲ | HarHarVeryFunny 3 hours ago | parent | prev [-] | | >> why GLM has two personalities >I never said that. This is a bit like me saying "I had eggs for breakfast", and you responding "I never said that!" Calm down buddy. |
|
|
|
| |
| ▲ | 7 hours ago | parent | prev [-] | | [deleted] |
|
|
|
| |
| ▲ | HarHarVeryFunny 2 days ago | parent | prev | next [-] | | > The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable. Useful for what is the question. Nobody is debating whether the outputs of LLMs are valuable. Given that Anthropic have redacted their true reasoning, and replaced it with a "summary", specifically designed to be useless for distillation purposes, it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming! | | |
| ▲ | throw10920 2 days ago | parent [-] | | > Useful for what is the question. Useful for distillation. Any employee at a frontier AI lab will tell you this. This is known in the industry, and it's an open secret that some US labs (OpenAI) distill on the others. Again - don't just make up stuff for a political agenda. > specifically designed to be useless for distillation purposes No, it's designed to give feedback to the user, in a way that minimizes its value for distilling. It's still valuable, and so there's a good chance that they'll remove it entirely as a result. > it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming! I did not claim that. Read my comment again: > The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable. Because apparently I have to spell it out: The output of a reasoning model is valuable, even if it didn't have the reasoning summary. Anthropic's models have a reasoning summary. The reasoning summary makes the output more valuable than if it didn't have a reasoning summary. It does not make it more valuable than having the full reasoning. | | |
| ▲ | HarHarVeryFunny 2 days ago | parent [-] | | > Useful for distillation Here's the thing: no-doubt a summary, if it at least reflects some of the logic connecting response to request, is better than nothing, so this can still be useful additional training data, but a model trained on it would be learning to generate these summaries, not the original withheld reasoning, so "distillation" seems an intentionally emotionally-wrought way of describing it (the "summary" is generated by a different smaller model - not the one the rest of the response came from). It does bring up an interesting point though - RL training in general results in "cargo-cult" reasoning - you train a model to follow the steps (mistakes and all - Karpathy) that got to a result, without understanding why they worked. If this training on summaries works just as well as training on detailed reasoning, then it just highlights how having a few breadcrumbs to follow/regurgitate is all that it takes. At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this. OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it. They are not copying Anthropic - they are, one assumes, using other models to generate cheap training data that they would otherwise have to pay people to generate. This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own. | | |
| ▲ | throw10920 2 days ago | parent [-] | | > At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this. Again - you're making things up. The fact is that everyone in the frontier AI lab space and the Chinese AI lab space knows that distilling is extremely effective and far more so than training from scratch. That's why China invests millions of dollars to create networks of tens of thousands of proxy accounts and shell companies to distill American models. It's a way to steal the R&D budget of another organization/nation-state. Stop making things up that you know nothing about. > OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it. No, it's stealing the value of the model. Anthropic has spent billions of dollars training their model. They have an R&D investment that anyone who knows how to add numbers understands has to be paid off, and anyone who has taken a basic economics class knows is the foundation for intellectual property: that to keep technological economies functioning, you have to have some sort of protection for technological inventions because they require upfront R&D investments. > This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own. This is just whataboutism and emotional manipulation. You can simultaneously believe that Anthropic did a bad thing when they scraped the whole internet and stole every book they could find to train their models, and that distillation is bad. In fact, anyone with a coherent moral compass would acknowledge that China is worse, because not only would they steal everything that Anthropic did, but they're also distilling other countries' models and they wouldn't even comply with US court cases, as Anthropic is. > Anthropic are apple-pie American innovators when ...and this is just jingoism. Not that I'm surprised, to be honest. | | |
| ▲ | HarHarVeryFunny a day ago | parent [-] | | You appear to know nothing about how these models work. Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lot of math", and you are claiming that from this you can train a Terrance Tao level model. Yes, I understand that Anthropic is upset that there is competition. Perhaps they should have realized that with no moat there was going to be competition and planned accordingly. | | |
| ▲ | throw10920 a day ago | parent [-] | | > You appear to know nothing about how these models work. > Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lot of math", and you are claiming that from this you can train a Terrance Tao level model. Yeah, you have no domain expertise and are making stuff up. To reiterate: the people who actually work at frontier labs know that you're factually wrong and will happily tell you. Reddit commentator syndrome yet again. > I bet you're assuming it comes from Fable. Nowhere did I assume or say that. That's the third or fourth time you've attributed things to me that I never said. It's extremely clear that you're not acting in good faith, because someone acting in good faith would never do that. If you continue responding, I'm going to continue debunking you, and you're just going to continue undermining your own points in the permanent HN record. > Yes, I understand that Anthropic is upset that there is competition. Emotional manipulation. Standard 50 Cent Party playbook. | | |
| ▲ | HarHarVeryFunny a day ago | parent [-] | | Here is the tweet by Dean Ball, OpenAI's "head of strategic futures", someone who does actually work at a frontier lab, saying that Kimi 3 can not be explained by distillation. https://x.com/deanwball/status/2078133895766114412?s=20 I'm not sure how you want to "debunk" that he said that, or twist what he said, but go ahead ... | | |
| ▲ | throw10920 18 hours ago | parent [-] | | You literally did not read my comment before responding or actually respond to any of the points. "I don't think its performance can be explained away by distillation or anything like that." even if you assume that Dean Ball (who is nontechnical and has not actually worked to train models (https://www.deanball.com/)) is honest (which he has a financial incentive to not be) - is entirely compatible with saying that Kimi was heavily distilled by Claude. At this point, I'm just pointing out the many lies, fallacies, and failures to read at a high-school level that you're committing. | | |
| ▲ | HarHarVeryFunny 6 hours ago | parent [-] | | > You literally did not read my comment before responding or actually respond to any of the points. Are you so unaware that you believe you are making any points? Go back and read your post - it was just a bunch of insults with zero technical content to respond to. The same emotional hysterics you've been using in the entire thread. | | |
| ▲ | throw10920 5 hours ago | parent [-] | | > it was just a bunch of insults with zero technical content to respond to Yet more lies. I made many substantial comments, and the fact that you're lying about that is you, not me. You've already lied about my own words multiple times (e.g. when you said "it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!") You're either an LLM with a bad hallucination rate or just evil, and precisely zero statements that you provide have any trustworthiness to them. Keep discrediting your account. I know that I can't convince a propagandist, but you just keep on further damaging your own reputation and argument every time you respond like this and ignore evidence that I've linked :) |
|
|
|
|
|
|
|
|
| |
| ▲ | wasfgwp 2 days ago | parent | prev [-] | | Well there were observed cases of Claude calling itself Deepseek or Qwen. So pot calling the kettle black? To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset? | | |
| ▲ | throw10920 2 days ago | parent [-] | | > Well there were observed cases of Claude calling itself Deepseek or Qwen. So pot calling the kettle black? There are open-source Deepseek and Qwen models - "distilling" doesn't involve breaking terms of service or hitting an API because you can literally run local inference or even just inspect the weights directly, and that's intended because they're open source. It's categorically different for a nation-state to build massive illicit networks of fraudulent identities to do distillation over tens of thousands of accounts to intentionally bypass providers' terms of service, intention for their models, and business model that very explicitly proprietary and not open source. https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... If Claude did distill on proprietary PRC LLMs - then fine, shame on them - I condemn that and I expect others to do the same. But there are no open-source Claude models. The only way for PRC models to have those responses is if they distilled Anthropic's models from their APIs. > To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset? ...and what would happen when it read all of the books and articles about Anthropic and replaced "replaced Claude Opus" with "replaced Qwen Opus"? Did you give any thought to this at all before saying it? | | |
| ▲ | wasfgwp 2 days ago | parent [-] | | > It's categorically The fraud part and using stolen accounts or credit cards or blatantly violating the terms and conditions (i.e. reselling subscriptions not using outputs in certain ways somebody might not like) is indeed categorically different. Using uncopyrightable outputs of an AI model obtained legitimately to train your model is not inherently interlinked with any of those things. I don’t really see how the model being proprietary or “open” is particularly relevant when talking about the outputs. Even using the word “distilling” in this case is deceptive and biased. It implies that the Chinese are somehow stealing Anthropic’s models or their weights and somehow directly transforming them into new models. That’s certainly not what’s happening in any direct sense. e.g. what if I agreed to send all my Claude code session files to Deepseek or whoever? There would be nothing wrong about that since I and not Anthropic own those files and can do whatever I want with them. Using certain different ways to obtain them of course could be highly illegal. |
|
|
|
|