Remix.run Logo
TheJCDenton 6 hours ago

> He also notes that the proprietary AI labs didn’t ask permission when they vacuumed up as much human knowledge as they could to train their models.

I think this should desactivate the moral high ground from which Anthropic is trying to speak. That they would want to make distillation orderly IMHO is fair, but to make it illegal is very rich from any AI frontier lab, really.

Bluestein 6 hours ago | parent | next [-]

Also, as said elsewhere: "Lab" is rich here, for outfits that, facing these giant, energy swallowing black boxes have really no clue what's going on inside.-

The moniker gives them an air of scientific, knowledgeable, tranquil, pro-social, pro bono work.-

Of course they are entitled to kill off a few mice, or pillage the commons to forward their "lab" work.-

samizdis 5 hours ago | parent | next [-]

> The moniker gives them an air of scientific, knowledgeable, tranquil, pro-social, pro bono work.

The Atlantic argued this (rather well, IMO) a week or so ago - "There’s No Such Thing as an AI ‘Lab’" - https://www.theatlantic.com/technology/2026/09/stop-calling-...

Den_VR 5 hours ago | parent | prev | next [-]

“We don’t know what’s going on” is essentially marketing. Sure we don’t _know_ but we have intuitions about why, where, and how to make certain changes…

numpad0 4 hours ago | parent [-]

Those labs publicly said during GPT-3/4 era that the optimal epoch count, or dataset repetition count, for foundation model training, is one. So it's a forward 1-pass compression.

But it's a black box! Nobody knows whats going on inside! It's all transformative! Sure...

paulddraper 2 hours ago | parent [-]

Have they seen big pharma?

pona-a 4 hours ago | parent | prev | next [-]

It used to be OpenAI was a real research organization that wrote real open-access papers that aren't marketing brochures, and when they did large training runs, they released all artifacts including model weights. Now certainly they are anything but. We haven't learned learned anything meaningful about ML from OpenAI since GPT-3 was released.

Their open-weights competitors like Facebook can at least claim some kind of public benefit, but it's still just running a well-understood algorithm on dubiously obtained data with longer and longer runs, give or take some inconsequential architectural tweaks.

Anthropic's mechanistic interpretability work is the most "lab-like" of these, but it's still just secondary to selling subscriptions and fear-mongering for regulatory capture/investment/publicity.

travisgriggs 5 hours ago | parent | prev [-]

We also associate laboratories with evil scientists and Frankenstein and the like. I can just hear Boris Karloff (er Bobby Picket) uttering “I was working in the lab late one night. When my eyes beheld an eerie sight… … … …the monster mash”. If anything, I associate _uncertainty_ with labs. The result is never known up front, they’re a place of discovery.

But I get your meaning. What should they be called instead? AI Sausage Factories maybe (cue Upton Sinclair?)?

Avicebron 5 hours ago | parent | next [-]

> What should they be called instead? AI Sausage Factories maybe (cue Upton Sinclair?)?

That's actually great? Slaughterhouses killing off the collective genius of humanity and grinding it into a bland paste for mass consumption.

travisgriggs 3 hours ago | parent [-]

Even kind of fits the model. All of the creativity man has raised is herded to the slaughterhouse and ground up so we end up with a big homogenized mash of ground up creativity, devoid of the life that gave it, rotten if not eaten soon enough.

whatshisface 3 hours ago | parent | prev [-]

By that definition, wall street would be a lab, and so would be a casino. I guess we could call them, "data refineries."

Bluestein 3 hours ago | parent [-]

Refinery makes a lot of sense. I like it particularly because it raises the question of whose (whose) "oil" (data) it is they are purloining.-

sobellian 6 hours ago | parent | prev | next [-]

I reflected on this myself recently. Model distillation seems to be at least as fair a use as distilling a book.

causal 6 hours ago | parent [-]

More than fair if you consider that the tokens are paid for.

dathery 6 hours ago | parent | next [-]

Both labs even explicitly promise the customer owns the outputs. It feels like they want to have their cake (ensure enterprises don't get spooked away from using as many LLMs as possible) while eating it too (still arguing some level of control over the outputs).

> Ownership of content. As between you and OpenAI, and to the extent permitted by applicable law, you (a) retain your ownership rights in Input and (b) own the Output. We hereby assign to you all our right, title, and interest, if any, in and to Output.

https://openai.com/policies/terms-of-use/

> As between the parties and to the extent permitted by applicable law, Anthropic agrees that Customer (a) retains all rights to its Inputs, and (b) owns its Outputs. Anthropic disclaims any rights it receives to the Customer Content under these Terms. Subject to Customer’s compliance with these Terms, Anthropic hereby assigns to Customer its right, title and interest (if any) in and to Outputs.

https://www.anthropic.com/legal/commercial-terms

Obviously there is some bad behavior going on in the distillation scene with gray-market token resellers but that is "just" normal fraud.

zenoprax 5 hours ago | parent [-]

> Both labs even explicitly promise the customer owns the outputs.

> to the extent permitted by applicable law, you (a) retain your ownership rights in Input and (b) own the Output

If the argument is that the model itself is under copyright protection then "as permitted by applicable law" would be doing some heavy lifting. Assuming that were true, given that locally-run LLMs exist, what would be illegal: the distillation itself or the provision of service of the distilled model?

visarga 4 hours ago | parent | prev | next [-]

Distilled content can also sever the direct link to infringement if the new models never saw the original texts.

5 hours ago | parent | prev [-]
[deleted]
6 hours ago | parent | prev | next [-]
[deleted]
5 hours ago | parent | prev | next [-]
[deleted]
toomuchtodo 6 hours ago | parent | prev [-]

YC does better if its startups get open weight frontier benefits. Garry’s just advocating for his book, which is his job. Consider how much capital YC portfolio companies would have to burn until liquidity if they have to pay OpenAI and Anthropic, versus relying on open weight frontier capabilities.

SOLAR_FIELDS 6 hours ago | parent | next [-]

If someone proposes the right thing for selfish reasons, do we call that bad? Or do we call it proper incentive alignment?

dofm 4 hours ago | parent | next [-]

We used to call it enlightened self-interest.

5 hours ago | parent | prev [-]
[deleted]
visarga 4 hours ago | parent | prev [-]

> versus relying on open weight frontier capabilities

ahem.. it happens even today, you can use open weight models directly and even fine tune