| ▲ | dannyw 35 minutes ago | |
Don’t sleep on NVIDIA and Nemotron. It’s not completely open source, but they actually release their pretraining and post-training datasets with some redactions for (cough) pirated content. They also have very good code and playbooks for actually doing a fine-tune, CPT, etc. Even if you’re not tuning a Nemotron model, its mixes are very excellent for your replay data slice; or general experiments. Way better curation and quality than Dolma, etc; or other large huggingface data mixes I tested. | ||
| ▲ | nickludlam 4 minutes ago | parent [-] | |
Yes, I second Nemotron. I'm using Ultra remotely and Super locally, and I find them very useful for RAG-like problems. I wouldn't really use them for coding. | ||