Remix.run Logo
sailingparrot 16 hours ago

Quite a jump in conclusion you are making here.

Sol is able to find the same counter example independently [1], so no reason to conclude in the existence of a benchmark destroying math beast Fable 6.

[1]: https://x.com/aaron_lou/status/2079218392452530249

tristanj 14 hours ago | parent [-]

My Fable 6 theory is admittedly speculative, but your “[public GPT-5.6] Sol is able to find the same counterexample” is also a jump in conclusions. Aaron specifically says he used “an internal version of Codex”. When asked whether that meant a different model or harness, he dodged the question, and only said the harness should be the standard commercial GPT-ultra harness [0]. He (intentionally) avoids identifying which model was used, so your claim is similarly unresolved. Given that Aaron works at OpenAI and has access to internal models, that GPT-5.7 is expected to launch in a few weeks and is rumored to be 10T+ parameters, it's very plausible that Aaron used that model in his analysis.

Furthermore, public GPT-5.6 pro failed six times to find a disproof to the Jacobian Conjecture [1], even with hints, which is evidence against the claim that the public GPT-5.6 Sol can solve this.

Regarding the existence of Fable 6: an internal upgraded version of Fable or Mythos almost certainly exists, given that Anthropic has been testing Mythos internally since April and previously released new models roughly every ~6 weeks.

[0] https://x.com/eliebakouch/status/2079237073001730510

[1] https://x.com/Tomodovodoo/status/2079172223055863895

sailingparrot 13 hours ago | parent [-]

You are right, I completely missed that, to me internal version of Codex meant the harness, but I was wrong.

tristanj 12 hours ago | parent [-]

Not necessarily, it's plausible that GPT-5.6 Sol could also find this disproof, as neither of us has enough information to make a conclusive case.

Public GPT is certainly capable, it was able to reverse-engineer the counterexample into a short proof: https://x.com/davikrehalt/status/2079175065695035442

I still suspect Fable 6 found the counterexample, mostly because Anthropic has been silent about this achievement. This is a huge accomplishment, and companies don't usually stay quiet when they hit a breakthrough like this. The lack of a press release or blog post is very suspicious.

tristanj 7 hours ago | parent [-]

Update: Aaron mentioned on X that he hasn't tested it on public ChatGPT Pro, so it was indeed an unreleased/internal model. https://x.com/aaron_lou/status/2079295442760736898

The next few weeks will be even more exciting!