Remix.run Logo
▲ lensecat an hour ago

Think they're referring to the following, when Amodei was still working for OpenAI:

'OpenAI Feared “Optics,” Not the Law – “Dario Amodei, OpenAI’s then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ [OpenAI researcher Sam] McCandlish explained: ‘I was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on [Hacker News] would be unfortunate.”'

▲yorwba 3 minutes ago | parent [-]

I was curious what he was responding to. Per https://authorsguild.org/app/uploads/2026/09/Class-Plaintiff... it was "On July 19, 2019, McCandlish wrote in an OpenAI Slack channel: “We’re not sure if we’re going to release the Foresight LM Scaling paper publicly, but if we do we were thinking about removing all mentions of LibGen, since it's a bit of a sketchy data source." The paper may or may not be https://arxiv.org/abs/2001.08361 where they write "we also test on similarly-prepared samples of Books Corpus [ZKZ+15], Common Crawl [Fou], English Wikipedia, and a collection of publicly-available Internet Books."