| ▲ | xvxvx 2 days ago |
| Reddit is my warning flag when AI gives me an answer. If I see Reddit linked, I know to disregard the answer entirely. They may as well be scraping 4chan. |
|
| ▲ | lostlogin a day ago | parent | next [-] |
| If you want a decent review of something, you could do a lot worse than Reddit. But they do seem to be trying to destroy themselves at the moment. |
|
| ▲ | AndrewKemendo a day ago | parent | prev [-] |
| Common Crawl is heavily populated with reddit threads So a massive chunk of LLMs are Reddit data |
| |
| ▲ | ccgreg a day ago | parent [-] | | That isn't true. You're welcome to peruse our index to prove or disprove your claim. | | |
| ▲ | AndrewKemendo a day ago | parent [-] | | I appreciate the response because it made me go back and look at my sources. What I conflated was that WebText utilizes Reddit links and data (mostly prior to 2023/4) and that I combined WebText and common crawl for the original GPT2 bootstrap into one dataset I was incorrectly connecting Reddit-mediated WebText pipeline to Common Crawl. So thanks for the correction! | | |
|
|