| ▲ | za_creature 17 hours ago | ||||||||||||||||
I answered here: https://news.ycombinator.com/item?id=49775387 I will continue to hold that position until such a time that we get a better answer than: > we cannot rule out that de-identified data derived from their usage of our products helped improve our models | |||||||||||||||||
| ▲ | andsoitis 17 hours ago | parent [-] | ||||||||||||||||
I hear you, but I think you might miss my point, which is while LLMs are clearly trained on copyrighted material, what they produce (their output) is NOT a copy of a specific code snippet they were trained on in a way that you would say "that's a copy from this code base". | |||||||||||||||||
| |||||||||||||||||