| ▲ | jamilton 44 minutes ago | |
Yeah, the filtering process probably wasn't particularly robust. The Schrodinger's cat example ("It's a cat that has been misbehavin'!") sounds like... a joke? Looking at the paper, it looks like they started with FineWeb-Edu, then filtered it based on an "age of word acquisition" dataset, with word frequency used as a proxy for values not in the dataset. They "only discard samples in which more than 5% of the words exceed the target age of 12." Maybe 5% was too high? They also filtered out beyond K-5 math symbols, like sigma. Then they trained a classifier to do more filtering. And they tested it on two grade-level benchmarks, and it only got 0-3% correct on the beyond k-5 boundary, while also decreasing in performance on the k-5 boundary (which they say is an acceptable tradeoff, since they were trying to get a sharp cutoff). So presumably since they got good results from the benchmark they stopped. | ||