Remix.run Logo
jacquesm 2 hours ago

Because he planned to make the data public, Meta just wants to use it to enrich its shareholders. See also: Google, OpenAI, Anthropic and every other big player in this space besides.

yepyeppers an hour ago | parent | next [-]

There’s also the fact that Swartz was physically trespassing and attaching unauthorized machines into networking closets to run scraping on a university network to exfiltrate the scrapes to the public, versus just scraping public facing web from the public web to train a model. That’s a little different and while the feds were heavy-handed against Swartz these computer crime laws were well known and it was less heavy handed than the hacker crackdowns of the 90s if you want to look at precedents.

derekdahmer 31 minutes ago | parent | prev [-]

Meta has opened sourced almost all of their models trained on this data.

jacquesm 8 minutes ago | parent [-]

But not the training data.