| ▲ | wg0 3 hours ago | |||||||
Suppose I inherited a data center spanning several hundred acres full of GPUs and free electricity. Where do I get the data? I mean, this many models. They have to start somewhere. | ||||||||
| ▲ | petu 3 hours ago | parent | next [-] | |||||||
I guess public datasets on HuggingFace and some shadow libraries content is enough to start. e.g. fineweb dataset is 50TB https://huggingface.co/datasets/HuggingFaceFW/fineweb | ||||||||
| ▲ | lucrbvi 3 hours ago | parent | prev | next [-] | |||||||
There are a lot of open-research on pre-training, post-training and RL data mixtures and sourcing. I recommend checking papers from Datalogy, Nvidia Nemotron, Ai2 (Ollmo, Tulu, ...) and the recent model from Aleph Alpha if you want to learn more. | ||||||||
| ▲ | altcognito 3 hours ago | parent | prev | next [-] | |||||||
If you ask a model, they will generally tell you where to get data. Modern frontier models have the large advantage of having tens if not hundreds of millions of users providing use cases to train against to improve their responses. | ||||||||
| ▲ | ttul 2 hours ago | parent | prev | next [-] | |||||||
https://scale.com/data-engine - you just buy it. | ||||||||
| ▲ | Hamuko 3 hours ago | parent | prev | next [-] | |||||||
Get data from Claude. That's what the Chinese (allegedly) do. | ||||||||
| ||||||||
| ▲ | konfusinomicon 2 hours ago | parent | prev [-] | |||||||
forget the data....sell it and go live your life! | ||||||||