Remix.run Logo
▲ ivo-42 2 hours ago

I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.

▲brcmthrowaway 2 hours ago | parent [-]

How do you cleanse the data at this scale?

▲ivo-42 2 hours ago | parent [-]

By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.

We have a lot of details in the tech report if you want to go deeper.

▲stephantul 29 minutes ago | parent [-]

Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining.

I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!