| ▲ | brcmthrowaway 2 hours ago | |||||||
How do you cleanse the data at this scale? | ||||||||
| ▲ | ivo-42 2 hours ago | parent [-] | |||||||
By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data. We have a lot of details in the tech report if you want to go deeper. | ||||||||
| ||||||||