Remix.run Logo
causal an hour ago

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations:

1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months?

2) Distillation - also implausible for the reason above.

3) Benchmark hacking. AI companies have ways they can dial up performance artificially, and will reach for that to maintain the appearance of parity.

Other reasons?

Edit: Most replies are ignoring timing. It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually.

logancbrown an hour ago | parent | next [-]

Its possible no AI lab has any unique edge, and success is a combination of (a) having access to GPUs (b) having access to large amounts of data (c) know about the handful of techniques to build an LLM, of which nearly all are likely open source and documented in papers. So the cycle of growth is (a) and (b), get more GPUs and get more data and you have a better model.

sm0ss117 44 minutes ago | parent | next [-]

Yea, this reads as LLMs are a pretty obvious technology to develop(for the highly intelligent researchers who are there). Also there's probably a lot of actual divergence in model capabilities and skills that concealed by the fairly narrow set of tests we run them against nowadays. Like wasn't Grok 4.20 super targeted at non-coding tasks.

causal 38 minutes ago | parent | prev [-]

GPUs might explain the remarkably concurrent timing. Data access doesn't really explain it unless all labs simultaneously got access to some treasure trove of data.

lanthissa 6 minutes ago | parent | prev | next [-]

what we're going through is the same thing as smartphones, the limiter is compute.

it used to be snapdragon came out HTC rushed out a janky phone everyone went omg htc is goat, then in the next few weeks and months others would impliment better versions and people would not notice those as much, finally sony would release a polished phone right as the next snapdragon cycle came.

eventually compute gains leveled off and apple won on taste.

nvidia/tpu is the new snapdragon. Anthropic and google both peaked on the first training run on a new tpu cycle.

you should expect amazing things within a few months of each other from everyone with access to chips and willingness to use them on a training run.

We haven't seen willingness from google to do that. So its currently xai,oai,anthropic, and probably soon meta.

glimshe an hour ago | parent | prev | next [-]

4) There's nothing terribly special about Anthropic. No moat.

causal an hour ago | parent [-]

Agreed, but my suspicion is tied to the timing. Catching up eventually is to be expected. Having similar jumps in capability ready at the same time is odd.

dash2 36 minutes ago | parent [-]

Maybe "readiness" is quite a flexible category? You're mid-training for your next model; a rival releases something; you clear the boards and release the model without completing the training run?

causal 32 minutes ago | parent [-]

Touche, aborted training runs probably do happen often. Closed model providers have zero incentive to announce a new model with less-than-best benchmarks.

extr an hour ago | parent | prev | next [-]

It's because Fable is just synthetic RL tasks + scale. The secret has been out for awhile now.

causal 42 minutes ago | parent [-]

Does not explain timing

extr 33 minutes ago | parent [-]

keep in mind fable = mythos which as been "done" since february. so the gap is not 2 months, it's more like - techniques probably started "working" in late 2025, now are trickling down to 2nd tier labs 9 months later.

causal 31 minutes ago | parent [-]

Yeah that would make more sense, it's probably a tight community and word gets around when something starts working.

ayewo 37 minutes ago | parent | prev | next [-]

> 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months?

The assumed timeline (2 months) is slightly wrong because Fable (Latin) is essentially the same as Mythos (Greek) albeit with protections against cyber and biological misuse.

Mythos (Preview) was publicly announced in April 2026 [1] which means other labs have had 4 months to catch up, not 2 months.

Assuming everyone had access to Mythos from the start, your expression, similar to other folks would have been "Mythos-level intelligence" and not "Fable-level intelligence".

1: https://news.ycombinator.com/item?id=47679258

causal 34 minutes ago | parent [-]

Fair point. Still a very quick turnaround considering the other labs would have to figure out both HOW to train a Mythos-level model and then do the work (and Grok is the last to catch up), but certainly more plausible than a 2 month window.

inerte 26 minutes ago | parent | prev | next [-]

No, it has happened to almost every other "sota" model before. There used to be a meme with a circular arrow going through Anthropic, OpenAI, Google as a hype circle. Now we can drop Google and add a couple of Chinese companies.

It's not an explanation of why it happens, I am just pointing Fable is not an exception, it has happened with almost every other model release by all these companies over the last 2-3 years.

moomin 42 minutes ago | parent | prev | next [-]

Yeah, I’m not convinced that there are any models as smart as Fable. Opus 5 definitely isn’t for all it has great benchmark scores. Fable displays judgement in a way I haven’t seen from any other model.

causal 40 minutes ago | parent [-]

Yeah as models get better, valid benchmarks become more "trust me bro".

jerf an hour ago | parent | prev | next [-]

Possibility: They're all hitting the same plateau of what LLMs can do with their current architectures.

I'm not stating this as a fact, but it's a hypothesis I'm keeping in my mix.

moduspol 36 minutes ago | parent [-]

It's possible, though I was thinking the same when GPT 5 released and it was kind of a nothing burger. Then I threw out that hypothesis with Opus 4.5.

Jcampuzano2 44 minutes ago | parent | prev | next [-]

I'm pretty sure both Anthropic and OpenAI haven't necessarily been secretive that they have internal models that are much more capable than commercially available ones.

It's probably a mix of all of that plus simply always keeping one in the chamber to 1up everyone else when the time is right.

causal 42 minutes ago | parent [-]

The "one in the chamber" is another good candidate that could explain the timing.

r_lee 30 minutes ago | parent [-]

I think this is the right one, iirc 5.6 came out quite soon after Opus 5 etc?

bottlepalm an hour ago | parent | prev | next [-]

I think model level is more a function of the state of hardware. Once it exists and is available (and if a lab can afford it), then they can train their own 1T, 5T, coming up next 10T model.

lanthissa 3 minutes ago | parent | next [-]

this is exactly whats happening. Its funny having lived through this with snap dragons and phones.

Everyones hyped about the branded phone, but it was the chip that mattered and how fast you rushed a product out after you got it.

Sames true now, except size of training run is also a factor.

causal 43 minutes ago | parent | prev [-]

This is a good candidate because it would also explain the timing. Most of the replies here do nothing to explain the timing I brought up.

user43928 35 minutes ago | parent | prev | next [-]

I understand Mythos became internally available on the 24th of February.

Other labs catching up in half a year seems about right.

becquerel an hour ago | parent | prev | next [-]

More compute is coming online at all times.

enraged_camel 37 minutes ago | parent | prev | next [-]

I'm solidly in the "they are benchmaxxing" camp. This became very apparent with GPT 5.6 Sol. It, too, was widely hailed to have near-Fable level intelligence. But I used it non-stop for a week and realized that they had mostly just dialed up the relentlessness meter to eleven, most likely via heavy RLHF.

Last week I gave it a small-sized auth ticket to work on, then stepped away. I came back later that afternoon and found that it had worked for 3+ hours and written 25,000+ lines of code. I skimmed over the code and it looked like a small fix followed by a massive number of additional checks around it, including static analysis tooling.

I gave it to another GPT 5.6 and said "check this code and see if it addresses the ticket". It looked at it and said that 98% of it was garbage and should be thrown away (its own words). I then gave it to Fable, which said it was massively over-engineered. Fable's theory was that the agent implemented the fix first, but then compacted and lost crucial context, forgot what the original task was about, and kept going. After many compaction cycles it was completely lost.

Some people complain that Opus 5 stops before finishing a task. But to me, that behavior is vastly preferable to what GPT 5.6 Sol does.

causal 29 minutes ago | parent [-]

Yeah I found the timing on Sol especially curious since it came right on the heels of Fable. I've had mixed results with it - sometimes it seems great, other times it makes mistakes so stupid I cannot understand how it ever gets anything right.

Explaining it as a difference of effort would explain both.

re-thc 28 minutes ago | parent | prev | next [-]

> It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually.

What are suspicious of? If the timing is similar maybe just everyone already are of similar capabilities and got there at a similar time?

> Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models?

It means Anthropic had no real moat and no real lead. Is that weird to you?

Traubenfuchs an hour ago | parent | prev [-]

> other reasons

Maybe research is sufficiently public and simple to reproduce or the next steps of how to improve things are sufficiently obvious to the smart people working on frontier AI.