Remix.run Logo
drooby a day ago

You might be misreading him.

He doesn't make a claim that implies "don't train" might not matter in the way you might think.

They cannot rule out that the data was trained because they have no per-user provenance tracking through the training pipeline once data is de-identified..

The entire point of de-identification is the inability to know the source of data. If the researcher forgot to hit "do not train" then that's that..

The only thing they have is a coincidence and the fact that the LLM may have used the training data that then researcher technically may have agreed to share.

Whether or not that's smoking gun of anything is hard to say. And the fact may remain that the proofs are significantly different, we do not know.