Remix.run Logo
profsummergig 2 hours ago

Could someone please share how such open source micro-LLMs might have been created?

Do the creators take something like DeepSeek, and then delete most of the neurons to whittle down the size?

hgoel an hour ago | parent | next [-]

Another option for something this small and narrowly specialized could be to get traditional LLMs to synthesize the training data. Model collapse is probably less of an issue at this size relative to terabyte sized models.

HenryNdubuaku 2 hours ago | parent | prev [-]

Technically, you could do that, but we trained this one from the ground up!

profsummergig 29 minutes ago | parent [-]

That sounds like an enormously expensive exercise.

ronsor 2 minutes ago | parent [-]

At <50M parameters, training costs are completely trivial. You'll spend a lot more on your rent this month.