| ▲ | absoluteunit1 5 hours ago |
| > Google DeepMind tested this impact by serving a model that used watermarking to a portion of their Gemini traffic and comparing thumbs-up and thumbs-down ratings. They found no statistically significant differences from the unwatermarked model. And in a controlled study, human raters comparing watermarked and unwatermarked answers side-by-side saw no difference in quality. For some reason I had assumed testing this would be more sophisticated than just checking the thumbs up/down stats and user "vibes" |
|
| ▲ | jonas21 4 hours ago | parent | next [-] |
| It's not just checking user thumbs up/down. As your quote says, they also did a controlled study with people rating the results. What else would you want them to do? The whole point is that it needs to introduce a detectable statistical difference, but humans should not be able to perceive it as a quality difference. |
| |
| ▲ | absoluteunit1 2 hours ago | parent | next [-] | | > human raters comparing watermarked and unwatermarked answers side-by-side saw no difference in quality Yes - maybe saying "vibes" was minimizing the effort but what I am trying to say is that even the controlled testing is just asking users whether quality is impacted or not. Which is subjective and thats what I meant by when I said "vibes" Don't get me wrong - I have no idea how one would go about testing this with other methods; I was just stating my assumption. Since they rolled this out to all users I had assumed there would be other testing involved. | |
| ▲ | bonoboTP 4 hours ago | parent | prev [-] | | > What else would you want them to do? Retest on benchmarks whether it accomplishes tasks with the same success rates. Prose is only one thing. Messing with the randomness may make the problem solving capabilities weaker. Probably it doesn't but this is the answer to what else I would want them to do. |
|
|
| ▲ | cube00 5 hours ago | parent | prev [-] |
| More unannounced testing on paying customers. |
| |
| ▲ | cj 4 hours ago | parent | next [-] | | The only function of the thumbs up/down buttons are to give feedback to Google. As a user it's pretty obvious that's the purpose of the button. | | |
| ▲ | cube00 4 hours ago | parent [-] | | You were still tested on and your outputs messed with even if you didn't click on either button. | | |
| |
| ▲ | baliex 5 hours ago | parent | prev [-] | | Genuine question, how else would they do it? And isn’t this practice the same as basically any agile-developed SaaS? | | |
| ▲ | thevinter 4 hours ago | parent [-] | | I'm not a mathematician but to me it doesn't seem so far-fetched to think that there might exist some mathematical proof that ensures the indistinguishability | | |
| ▲ | antonvs 4 hours ago | parent [-] | | We don’t have the ability to do that level of analysis of natural language text mathematically. If we did, we probably wouldn’t need LLMs in the first place, i.e. we could just generate text using explicitly programmed algorithms. |
|
|
|