| |
| ▲ | ivo-42 2 hours ago | parent [-] | | By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data. We have a lot of details in the tech report if you want to go deeper. | | |
| ▲ | stephantul 29 minutes ago | parent [-] | | Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining. I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful! |
|
|