| ▲ | profsummergig 2 hours ago | ||||||||||||||||
Could someone please share how such open source micro-LLMs might have been created? Do the creators take something like DeepSeek, and then delete most of the neurons to whittle down the size? | |||||||||||||||||
| ▲ | hgoel an hour ago | parent | next [-] | ||||||||||||||||
Another option for something this small and narrowly specialized could be to get traditional LLMs to synthesize the training data. Model collapse is probably less of an issue at this size relative to terabyte sized models. | |||||||||||||||||
| ▲ | HenryNdubuaku 2 hours ago | parent | prev [-] | ||||||||||||||||
Technically, you could do that, but we trained this one from the ground up! | |||||||||||||||||
| |||||||||||||||||