| ▲ | ctoth 10 hours ago |
| This ... is not how this works. The model is not speaking longer to watermark anything. |
|
| ▲ | TheOtherHobbes 9 hours ago | parent | next [-] |
| It's exactly how it works - at least potentially. Lean text is harder to watermark because word choices and meanings are tightly constrained. Low-entropy text is fluff and filler. It's very easy to synonym-substitute words without changing the message - if there even is one. |
| |
| ▲ | usef- 6 hours ago | parent [-] | | You're assuming they're training the model to maximize the watermark signal, on top of already adding the watermark. I suspect that would hurt model performance quite a lot, and simply be unnecessary... the watermark tech works well enough as it is. As far as I know, anthropic aren't intrinsically motivated by watermarking (if anything it hurts sales, and seems indifferent to safety(?)) they're simply doing it to fulfill the EU obligations. | | |
| ▲ | northzen 24 minutes ago | parent [-] | | > As far as I know, anthropic aren't intrinsically motivated by watermarking (if anything it hurts sales, and seems indifferent to safety(?)) they're simply doing it to fulfill the EU obligations. They are. They want to reduce the amount of LLM generated text they feed into their next model training. Also, how would you watermark a sentence with just 3 words for an example? This exactly why it became so verbose. |
|
|
|
| ▲ | skarz 9 hours ago | parent | prev [-] |
| Perhaps, but there are certainly now catchphrases and words that can indicate it was written with AI i.e. load-bearing, idempotent, etc. Style and structure are in and of themselves, a fingerprint. |
| |
| ▲ | nick__m 9 hours ago | parent [-] | | idempotent was frequently used before LLM; it's hard to talk about REST and infrastructure as code without using that word... |
|