Remix.run Logo
COAGULOPATH 5 hours ago

>I'd like to know a lot more about how that works.

My guess is that it works like Gemini's SynthID: by altering the logprobs of the next token.

Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something. (Obviously it's way more complicated but I think conceptually this is how it works.) No human will notice this, but a classifier trained on Claude's output will.

So it's not like the watermark is the words "le epic bacon" and Claude will output "le epic bacon" in everything. That would be extremely annoying (and easy to defeat).

thunfischtoast 2 hours ago | parent [-]

They still need to choose when to do that though. When I prompt the program to e.g. alter a bash script in a specific way or to recite a longer known text it can't go round and randomly exchange tokens. It has to somehow define what is a simple repeated text from a different origin and what is a novel generation.

user43928 2 hours ago | parent [-]

I am wondering how that applies to newly generated code.

Odd variable naming? Stylistic choices that are watermarked?

Or as someone else noted further down in the comments, it could be more subtle:

Between the first and second most likely choice, in certain positions it will consistently choose in a certain way.

melvinroest 36 minutes ago | parent [-]

> Odd variable naming? Stylistic choices that are watermarked?

Whatever it is, I'm sure it's load-bearing.

asdfsa32 27 minutes ago | parent [-]

You're absolutely right. But it is not just load-bearing, it is the load-bearing seams.