Remix.run Logo
ActionHank 5 hours ago

"This should be allowed, let me explain the reason they cheated and state again that they should be allowed to cheat."

simianwords 5 hours ago | parent [-]

> GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.

> Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository and testing policy document both approaches.

This is what the Author of the benchmark has to stay. Quality of the comments keep going down smh

ActionHank 5 hours ago | parent [-]

"Going forward we will capitulate and still try to keep the integrity of our benchmark in tact, but from now on every benchmark will be compromised with providers being able tweak things sufficiently to game at least a 30% bump in results."

simianwords 5 hours ago | parent [-]

"I'll twist the words of the author of the benchmark itself to make a point"

ActionHank 5 hours ago | parent [-]

"I refuse to see the wall that I am running directly into, because if I see it I will hit it"

simianwords 5 hours ago | parent [-]

If you mean a capability wall, the author of the benchmark says this

>We see Astra as a major breakthrough in model intelligence.

You think the author of the benchmark is also in the conspiracy