Remix.run Logo
throw10920 2 days ago

> Useful for what is the question.

Useful for distillation. Any employee at a frontier AI lab will tell you this. This is known in the industry, and it's an open secret that some US labs (OpenAI) distill on the others. Again - don't just make up stuff for a political agenda.

> specifically designed to be useless for distillation purposes

No, it's designed to give feedback to the user, in a way that minimizes its value for distilling. It's still valuable, and so there's a good chance that they'll remove it entirely as a result.

> it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!

I did not claim that. Read my comment again:

> The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.

Because apparently I have to spell it out:

The output of a reasoning model is valuable, even if it didn't have the reasoning summary. Anthropic's models have a reasoning summary. The reasoning summary makes the output more valuable than if it didn't have a reasoning summary. It does not make it more valuable than having the full reasoning.

HarHarVeryFunny 2 days ago | parent [-]

> Useful for distillation

Here's the thing: no-doubt a summary, if it at least reflects some of the logic connecting response to request, is better than nothing, so this can still be useful additional training data, but a model trained on it would be learning to generate these summaries, not the original withheld reasoning, so "distillation" seems an intentionally emotionally-wrought way of describing it (the "summary" is generated by a different smaller model - not the one the rest of the response came from).

It does bring up an interesting point though - RL training in general results in "cargo-cult" reasoning - you train a model to follow the steps (mistakes and all - Karpathy) that got to a result, without understanding why they worked. If this training on summaries works just as well as training on detailed reasoning, then it just highlights how having a few breadcrumbs to follow/regurgitate is all that it takes.

At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this. OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it. They are not copying Anthropic - they are, one assumes, using other models to generate cheap training data that they would otherwise have to pay people to generate.

This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own.

throw10920 2 days ago | parent [-]

> At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this.

Again - you're making things up. The fact is that everyone in the frontier AI lab space and the Chinese AI lab space knows that distilling is extremely effective and far more so than training from scratch. That's why China invests millions of dollars to create networks of tens of thousands of proxy accounts and shell companies to distill American models.

It's a way to steal the R&D budget of another organization/nation-state.

Stop making things up that you know nothing about.

> OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it.

No, it's stealing the value of the model. Anthropic has spent billions of dollars training their model. They have an R&D investment that anyone who knows how to add numbers understands has to be paid off, and anyone who has taken a basic economics class knows is the foundation for intellectual property: that to keep technological economies functioning, you have to have some sort of protection for technological inventions because they require upfront R&D investments.

> This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own.

This is just whataboutism and emotional manipulation. You can simultaneously believe that Anthropic did a bad thing when they scraped the whole internet and stole every book they could find to train their models, and that distillation is bad.

In fact, anyone with a coherent moral compass would acknowledge that China is worse, because not only would they steal everything that Anthropic did, but they're also distilling other countries' models and they wouldn't even comply with US court cases, as Anthropic is.

> Anthropic are apple-pie American innovators when

...and this is just jingoism. Not that I'm surprised, to be honest.

HarHarVeryFunny a day ago | parent [-]

You appear to know nothing about how these models work.

Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lot of math", and you are claiming that from this you can train a Terrance Tao level model.

Yes, I understand that Anthropic is upset that there is competition. Perhaps they should have realized that with no moat there was going to be competition and planned accordingly.

throw10920 a day ago | parent [-]

> You appear to know nothing about how these models work.

> Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lot of math", and you are claiming that from this you can train a Terrance Tao level model.

Yeah, you have no domain expertise and are making stuff up. To reiterate: the people who actually work at frontier labs know that you're factually wrong and will happily tell you. Reddit commentator syndrome yet again.

> I bet you're assuming it comes from Fable.

Nowhere did I assume or say that. That's the third or fourth time you've attributed things to me that I never said. It's extremely clear that you're not acting in good faith, because someone acting in good faith would never do that. If you continue responding, I'm going to continue debunking you, and you're just going to continue undermining your own points in the permanent HN record.

> Yes, I understand that Anthropic is upset that there is competition.

Emotional manipulation. Standard 50 Cent Party playbook.

HarHarVeryFunny a day ago | parent [-]

Here is the tweet by Dean Ball, OpenAI's "head of strategic futures", someone who does actually work at a frontier lab, saying that Kimi 3 can not be explained by distillation.

https://x.com/deanwball/status/2078133895766114412?s=20

I'm not sure how you want to "debunk" that he said that, or twist what he said, but go ahead ...

throw10920 18 hours ago | parent [-]

You literally did not read my comment before responding or actually respond to any of the points.

"I don't think its performance can be explained away by distillation or anything like that." even if you assume that Dean Ball (who is nontechnical and has not actually worked to train models (https://www.deanball.com/)) is honest (which he has a financial incentive to not be) - is entirely compatible with saying that Kimi was heavily distilled by Claude.

At this point, I'm just pointing out the many lies, fallacies, and failures to read at a high-school level that you're committing.

HarHarVeryFunny 6 hours ago | parent [-]

> You literally did not read my comment before responding or actually respond to any of the points.

Are you so unaware that you believe you are making any points?

Go back and read your post - it was just a bunch of insults with zero technical content to respond to. The same emotional hysterics you've been using in the entire thread.

throw10920 5 hours ago | parent [-]

> it was just a bunch of insults with zero technical content to respond to

Yet more lies. I made many substantial comments, and the fact that you're lying about that is you, not me.

You've already lied about my own words multiple times (e.g. when you said "it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!") You're either an LLM with a bad hallucination rate or just evil, and precisely zero statements that you provide have any trustworthiness to them.

Keep discrediting your account. I know that I can't convince a propagandist, but you just keep on further damaging your own reputation and argument every time you respond like this and ignore evidence that I've linked :)