Remix.run Logo
tristanj 21 hours ago

The poster works at Anthropic, so they likely have internal access to the next generation of Fable. Their internal model is probably an absolute beast at mathematics, and the upcoming benchmark results will likely set a new record for maths performance.

I suspect this is what happened, because the poster is coy about sharing the actual prompt / reasoning trace used to reach this result. That would be covered by an NDA until the model is properly released.

Exciting times!

sailingparrot 16 hours ago | parent | next [-]

Quite a jump in conclusion you are making here.

Sol is able to find the same counter example independently [1], so no reason to conclude in the existence of a benchmark destroying math beast Fable 6.

[1]: https://x.com/aaron_lou/status/2079218392452530249

tristanj 14 hours ago | parent [-]

My Fable 6 theory is admittedly speculative, but your “[public GPT-5.6] Sol is able to find the same counterexample” is also a jump in conclusions. Aaron specifically says he used “an internal version of Codex”. When asked whether that meant a different model or harness, he dodged the question, and only said the harness should be the standard commercial GPT-ultra harness [0]. He (intentionally) avoids identifying which model was used, so your claim is similarly unresolved. Given that Aaron works at OpenAI and has access to internal models, that GPT-5.7 is expected to launch in a few weeks and is rumored to be 10T+ parameters, it's very plausible that Aaron used that model in his analysis.

Furthermore, public GPT-5.6 pro failed six times to find a disproof to the Jacobian Conjecture [1], even with hints, which is evidence against the claim that the public GPT-5.6 Sol can solve this.

Regarding the existence of Fable 6: an internal upgraded version of Fable or Mythos almost certainly exists, given that Anthropic has been testing Mythos internally since April and previously released new models roughly every ~6 weeks.

[0] https://x.com/eliebakouch/status/2079237073001730510

[1] https://x.com/Tomodovodoo/status/2079172223055863895

sailingparrot 13 hours ago | parent [-]

You are right, I completely missed that, to me internal version of Codex meant the harness, but I was wrong.

tristanj 12 hours ago | parent [-]

Not necessarily, it's plausible that GPT-5.6 Sol could also find this disproof, as neither of us has enough information to make a conclusive case.

Public GPT is certainly capable, it was able to reverse-engineer the counterexample into a short proof: https://x.com/davikrehalt/status/2079175065695035442

I still suspect Fable 6 found the counterexample, mostly because Anthropic has been silent about this achievement. This is a huge accomplishment, and companies don't usually stay quiet when they hit a breakthrough like this. The lack of a press release or blog post is very suspicious.

tristanj 7 hours ago | parent [-]

Update: Aaron mentioned on X that he hasn't tested it on public ChatGPT Pro, so it was indeed an unreleased/internal model. https://x.com/aaron_lou/status/2079295442760736898

The next few weeks will be even more exciting!

ev0lv 18 hours ago | parent | prev [-]

I'm not very excited. Access to the best AI is not a party I was invited to. And the people who are at that party, well, they don't exactly reflect on my best interests.

tptacek 17 hours ago | parent [-]

Is mathematics an science or an art? To the extent it's art, it's expressive and rewards the human experience that inspires and that creates it. If it's art, it's drastically less valuable to advance it through automation. But if mathematics is a science, then our entire goal is to increase humanity's understanding of the field. Whether by automation or genius inspiration or as a reward for decades of grinding it out incrementally, it's all the same: knowing more is the point.