Remix.run Logo
panabee 37 minutes ago

These can all be true at once:

1. OpenAI's internal model benefited from the reasoning in Alpöge and Buckmaster's Codex sessions through de-identified usage data.

2. OpenAI's model seems to have surpassed Alpöge and Buckmaster's results and produced a field-defining proof. If verified, this work would be the first AI resolution of a Millennium Prize problem.

3. Alpöge and Buckmaster did something remarkable. What could have been a celebration of human-AI collaboration is now overshadowed by controversy.

The first point remains speculative, resting on these concepts:

- Data labs like Surge and Mercor recruit PhDs and industry professionals, few of whom work at the research frontier. Their contractors suit many tasks, like Salesforce or SAP work, but not scientific breakthroughs or complex coding. Meta drafted its own engineers into data annotation for this reason. Zuckerberg told employees that Meta staff are smarter on average than the contractors that data-labeling firms typically hire. [0]

- It's easier to judge a 5-star dish than to create one. Similarly, it's easier for models to judge insightful reasoning than to create it. A model that cannot produce breakthrough ideas can still recognize them in user sessions, flag them, and de-identify details before adding insights to the training corpus. This is especially true in mathematics, where Lean can verify proofs automatically.

Even if true, point 1 may have had minimal impact or been unnecessary. Outsiders cannot verify or falsify it.

[0] https://techcrunch.com/2026/06/12/metas-months-old-ai-unit-i...