| ▲ | jonas21 4 hours ago | |
It's not just checking user thumbs up/down. As your quote says, they also did a controlled study with people rating the results. What else would you want them to do? The whole point is that it needs to introduce a detectable statistical difference, but humans should not be able to perceive it as a quality difference. | ||
| ▲ | absoluteunit1 2 hours ago | parent | next [-] | |
> human raters comparing watermarked and unwatermarked answers side-by-side saw no difference in quality Yes - maybe saying "vibes" was minimizing the effort but what I am trying to say is that even the controlled testing is just asking users whether quality is impacted or not. Which is subjective and thats what I meant by when I said "vibes" Don't get me wrong - I have no idea how one would go about testing this with other methods; I was just stating my assumption. Since they rolled this out to all users I had assumed there would be other testing involved. | ||
| ▲ | bonoboTP 4 hours ago | parent | prev [-] | |
> What else would you want them to do? Retest on benchmarks whether it accomplishes tasks with the same success rates. Prose is only one thing. Messing with the randomness may make the problem solving capabilities weaker. Probably it doesn't but this is the answer to what else I would want them to do. | ||