Pulled datasets off libgen for a corpus once and it's overwhelmingly in-copyright textbooks, the public domain framing doesn't really hold up.