| ▲ | avaer 8 hours ago | ||||||||||||||||||||||
What's "Top U.S. Models"? Even if we know what the set is, it's not clear what the numbers actually mean. Is it min, max, average, weighted, median of the models? Prerelease or public, with or without safeguards? A bit frustrating to have this be hand waved in a report, the graphs might as well just have two mystery bars, U.S. and China. | |||||||||||||||||||||||
| ▲ | guessmyname 7 hours ago | parent | next [-] | ||||||||||||||||||||||
NIST has named the top U.S. models in their (full) report: https://www.nist.gov/system/files/documents/2026/07/17/CAISI... Spoiler: it’s OpenAI’s GPT-5.5 and Anthropic’s Mythos Preview [**]. [**] Reminder that Mythos Preview is a very different beast from Fable5, Mythos5, and Opus5. Unfortunately, anyone outside Project Glasswing will probably never get to test what this model can actually do, which is a shame, and it’s a big part of why the industry is so skeptical of the capabilities insiders keep claiming it has. I’d be skeptical too if I hadn’t tested it myself. We lost access to Mythos Preview when Anthropic forced us onto Mythos 5 some weeks ago, which is garbage by comparison. I’ve already switched to GPT-5.5 and I’m working on adapting my harness(es) to less restrictive open-weight models. I don’t see any other way forward at this point. | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | NitpickLawyer 7 hours ago | parent | prev [-] | ||||||||||||||||||||||
Check out the "Completed steps on..." graph in this [1] evaluation. That graph gives a good perspective of what models they've tested, and roughly what "subject" each step covers. It is on that task that they note this: > Kimi K3 reaches step 17 on average, compared with step 11 for GLM-5.2. [1] - https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5... | |||||||||||||||||||||||