| ▲ | sailingparrot 16 hours ago | |||||||||||||||||||||||||
Quite a jump in conclusion you are making here. Sol is able to find the same counter example independently [1], so no reason to conclude in the existence of a benchmark destroying math beast Fable 6. | ||||||||||||||||||||||||||
| ▲ | tristanj 14 hours ago | parent [-] | |||||||||||||||||||||||||
My Fable 6 theory is admittedly speculative, but your “[public GPT-5.6] Sol is able to find the same counterexample” is also a jump in conclusions. Aaron specifically says he used “an internal version of Codex”. When asked whether that meant a different model or harness, he dodged the question, and only said the harness should be the standard commercial GPT-ultra harness [0]. He (intentionally) avoids identifying which model was used, so your claim is similarly unresolved. Given that Aaron works at OpenAI and has access to internal models, that GPT-5.7 is expected to launch in a few weeks and is rumored to be 10T+ parameters, it's very plausible that Aaron used that model in his analysis. Furthermore, public GPT-5.6 pro failed six times to find a disproof to the Jacobian Conjecture [1], even with hints, which is evidence against the claim that the public GPT-5.6 Sol can solve this. Regarding the existence of Fable 6: an internal upgraded version of Fable or Mythos almost certainly exists, given that Anthropic has been testing Mythos internally since April and previously released new models roughly every ~6 weeks. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||