| ▲ | simonw 9 hours ago | |
> The pelican prompt is ridiculous Yes, deliberately so. It was never intended as a meaningful benchmark. The surprising thing was that for the first ~12 months performance on the stupid pelican benchmark did seem to correspond to the performance of the models on other tasks. That pattern no longer holds - Fable 5 and GPT-5.6 have both been out-pelicaned by lesser models now. | ||