| ▲ | scrlk 7 hours ago | ||||||||||||||||||||||||||||||||||
Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...) | |||||||||||||||||||||||||||||||||||
| ▲ | tedsanders 6 hours ago | parent | next [-] | ||||||||||||||||||||||||||||||||||
Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all. ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good! (I coauthored the linked blog post) | |||||||||||||||||||||||||||||||||||
| ▲ | woah 6 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||
Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses? | |||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||
| ▲ | kasperni 7 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||
yes it is. | |||||||||||||||||||||||||||||||||||
| ▲ | enraged_camel 6 hours ago | parent | prev [-] | ||||||||||||||||||||||||||||||||||
Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO. | |||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||