| ▲ | RGS1811 5 hours ago | |||||||
"This is a classic test request..." I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem. | ||||||||
| ▲ | croemer 4 hours ago | parent | next [-] | |||||||
Of course it was exposed - not sure it's explicit or not. Why wouldn't HackerNews comments be part of the training data? And Simon's blog and the many discussions about Pelicans? It'd be hard to miss. Doesn't mean Anthropic has made this an explicit goal in training. | ||||||||
| ▲ | pgwhalen an hour ago | parent | prev | next [-] | |||||||
It would be genuinely shocking at this point if any of the frontier models weren't well exposed to the problem. | ||||||||
| ▲ | simonw 5 hours ago | parent | prev | next [-] | |||||||
See here for more discussion of that: https://news.ycombinator.com/item?id=49803892#49804881 | ||||||||
| ||||||||
| ▲ | 0x10ca1h0st 4 hours ago | parent | prev | next [-] | |||||||
Lets start frog riding motorcycle trend until they frogmaxx, or cat driving convertible. | ||||||||
| ▲ | cubefox 5 hours ago | parent | prev [-] | |||||||
The model recognizing the task doesn't mean it was benchmaxxed (RLVR-trained) to solve it. It might simply recognize it from pre-training on Internet text. | ||||||||