| ▲ | echelon 4 hours ago | |
Eventually we'll just construct 100% synthetic training data that can reliably reproduce pretrains and fine tunes. The first broadly useful fully open source models will do this. We already have open data / open code / open weights for some domain-specific cases, such as audio models trained on large open datasets, eg. Tacotron / LJSpeech from waaay back in the day, though that is certainly not SOTA anymore. Distillation could possibly be considered an early case of this as raw AI outputs are themselves not copyrightable unless humans enrich, filter, or transform them. Granted, that does not handle the cases where the outputs are sufficiently similar to copyrighted original works. | ||
| ▲ | chaosharmonic 3 hours ago | parent | next [-] | |
But how much of that synthetic data still ultimately derives from non-open sources? You'd still have to ask what a clean room implementation ultimately is, depending on how granular or aggressive a large publisher wanted to get about it. That said, I don't necessarily disagree with you. Talkie[1] presents an interesting case for it being at least possible to do this entirely on public domain material. But even that used Claude somewhere in the course of its training pipeline (it's listed as a contributor on their GitHub), so again, how granular you want to get with that is still a question. | ||
| ▲ | waffleiron 3 hours ago | parent | prev [-] | |
Where does that synthetic data come from? Magically just started existing? | ||