Remix.run Logo
hodgehog11 a day ago

The "second only to Fable 5" comment is pretty telling here. I remember early on when a lot of naysayers were saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model, provided they continue to allow people to use it. It will be genuinely exciting when an open model is able to beat it.

InsideOutSanta a day ago | parent | next [-]

I wouldn't call it a moat, but I would call it a noticeably better model. Subjectively, for my own work, I would rate the top models Fable > K3 > Sol.

But it's not like Fable is so substantially better than the other two that I would be seriously impacted if I didn't have access to it anymore. All three are amazing models, and of the three, Fable is the only one that regularly triggers refusals.

hodgehog11 18 hours ago | parent | next [-]

It really does depend on your application. In my domain (math research), it is substantially better. Fable can solve really hard tasks with surprising consistency. It makes mistakes, and occasionally refuses, but honestly, at the top level, ideas are the currency and the rigor is the busywork. The other models cannot come close in this domain.

If you couple Fable's idea factory with Sol's rigor, you get a real game-changer. It puts the emphasis on top-level ideas, and nearly trivialises the intermediate layers.

porker a day ago | parent | prev [-]

It's always fun to see what works for others, because for my work it'd have to be Sol > Fable. Fable makes too many mistakes.

Coordinating agents though? Fable any day.

neevans a day ago | parent | prev | next [-]

tbh even if its better model due to lot of restrictions its not that useful than opus.

NitpickLawyer a day ago | parent | next [-]

100% this. There's currently this [1] submission that hasn't gained much attention, but is really important. In this [2] incident report from HuggingFace, they talk about detecting an attack and not being able to analyse the logs / IoC with API models because of guardrails. If not even highly regarded reputable companies can't sort out access to SotA models for blue team use, the raw capabilities don't matter. They're useless paperweights (hah!), and nothing else. Having to resort to open models is insane!

[1] - https://news.ycombinator.com/item?id=48965243

[2] - https://huggingface.co/blog/security-incident-july-2026

throwa356262 a day ago | parent [-]

Key part from [2]:

"When we started the log analysis, we first used frontier models behind commercial APIs. This did not work [...] We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. [...] The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout [...]"

hodgehog11 a day ago | parent | prev [-]

Tell that to my colleagues. Despite Sol getting the attention, Fable is really starting to have an impact on mathematicians right now. It has unbelievable insights in a lot of cases that can rapidly speed up progress.

ferrouswheel a day ago | parent | prev | next [-]

I dunno, I find Fable slops alot. Sol is my workhorse. Fable can be creative but isn't very good at doing work reliably (or without endlessly burning tokens).

vitalyan8184 a day ago | parent | prev | next [-]

their "genuine moat" is that mythos is the only super heavyweight model right now. it's always been possible to train a 10T model and get 10% more performance over a 1T model.

had mythos been just Opus 5, with the same size and price as the previous opuses, then yeah, that would be a tie-breaker. but it's not.

XCSme a day ago | parent | prev | next [-]

I am not even sure if Fable is as smart as they say, I can't get it to answer almost any question, it always refuses for "cyber-security" concerns...

reckless a day ago | parent | prev | next [-]

5.6-sol would be a better comparison given it's general availability and usage allowances

hodgehog11 18 hours ago | parent | next [-]

They are likely assessing based on "raw intelligence" benchmarks, rather than agentic ones. Fable crushes in those, but that doesn't necessarily translate to microscopic rigor, which is what most people use these models for. You only see it when you ask really tough questions.

ferrouswheel a day ago | parent | prev [-]

And sol is much more reliable as a agent for doing work. Fable sometimes just goes on wild flights of fancy.

epolanski 7 hours ago | parent | prev | next [-]

> Like it or not, Anthropic have a genuine moat

I don't think yours and my definition of moat is the same, when I can switch from Fable to competing models and have similar results.

scotty79 a day ago | parent | prev | next [-]

> saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model,

What's more interesting is that Anthropic moat shrunk to just that model. There's zero reason to use any other model from Anthropic right now. And once they take Fable off subscription there will be zero reason to have Anthropic subscription.

Havoc 17 hours ago | parent | prev | next [-]

More like a puddle than a moat

dgellow a day ago | parent | prev | next [-]

That’s not a moat though

nprateem 18 hours ago | parent | prev | next [-]

Fable is still as dumb as a post. I ask it simple questions and it routinely gets things backwards, prioritises things that should be subordinate to others, etc.

An example: It just suggested that I shouldn't raise the price of my saas because it'd complicate the arithmetic if I did 0 -> $100k YT channel instead of sticking to $20 p/m.

It's just a complete moron, like all of them.

matheusmoreira a day ago | parent | prev [-]

What's the point of Fable if we can't use it? I get to prompt it like 5 times on my subscription before it gets cut off, and even then I'm constantly fighting the insufferable safety classifier.

I'll switch to OpenAI soon because of this. I also can't wait for the day it becomes feasible to run these awesome open weight models on my own hardware.