Remix.run Logo
bel8 6 hours ago

So it's a Fable class LLM?

                             DSV4Pro vs Fable5
    HLE w tools              60.0 vs 63.0
    Terminal Bench 2.1       87.9 vs 88.0
    Cybergym                 83.3 vs 83.1
    DeepSWE                  62.7 vs 70.0
    Toolathlon-Verified      74.1 vs 77.9
    AutomationBench (Public) 31.8 vs 29.1
    DSBench-FullStack        71.1 vs 77.2
    DSBench-Hard             67.2 vs 68.3
eli 6 hours ago | parent | next [-]

Fable's guardrails would never let it do something like Cybergym so at least for that one it's measuring Opus 5

wren6991 6 hours ago | parent [-]

We have a first-party figure from the system card [1]:

> Mythos 5 reproduced 83.8% of targeted vulnerabilities on a single try, and produced at least one crash in 99.4% of tasks. This is comparable to Claude Mythos Preview, which reproduced 83.1% of targeted vulnerabilities and produced a crash in 97.1% of tasks. By contrast, Claude Opus 4.8 achieved a score of 78.1% (95.7% any crash).

So their quoted figure exactly matches the figure for Mythos Preview, although they don't state the provenance. It could also quite possibly be an independent measurement of Opus 5.

[1]: https://www-cdn.anthropic.com/57a52ea7d8f0e54e8a542e90826608...

nikcub 3 hours ago | parent | prev | next [-]

that DeepSWE result is likely most indicative of how you'll find real world usage

aftbit 6 hours ago | parent | prev [-]

Fabble lol

qiran87 5 hours ago | parent [-]

[dead]