| ▲ | kasperni 7 hours ago | |
"On the current ARC-AGI-3 leaderboard, conventional frontier-model runs sit dramatically below Astra's reported 98.6% result. But the comparison isn't straightforward. OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations." | ||