I appreciate the response because it made me go back and look at my sources.
What I conflated was that WebText utilizes Reddit links and data (mostly prior to 2023/4) and that I combined WebText and common crawl for the original GPT2 bootstrap into one dataset
I was incorrectly connecting Reddit-mediated WebText pipeline to Common Crawl. So thanks for the correction!