Remix.run Logo
dyauspitr 7 hours ago

Why is Fable not on here? I wish Fable hadn’t come out because it’s taking the wind out of every release because that feels like the cap above which the US government will not let LLMs improve anymore and everything they’re releasing from this point has to be worse than that.

NitpickLawyer 7 hours ago | parent | next [-]

> Why is Fable not on here?

Because the data retention policies didn't guarantee that the ARC team could run the semi-private set of problems without fear of them being trained on later on. They only run the semi-private set when they get assurances like ZDR.

3form 6 hours ago | parent | next [-]

How do they handle these assurances? Personally I have zero trust in the AI companies not trying to use this data to get ahead in the game, and short of sharing the weights and harness so that the benchmarkers can run the models themselves, I don't see a satisfactory solution with this mindset.

villish 6 hours ago | parent [-]

OpenAI's Zero Data Retention claim held up in court. They were unable to produce prompts and outputs because they were never retained.

I believe that is only available through Enterprise API for both Anthropic and OpenAI.

Barbing 5 hours ago | parent [-]

Is there a distinction we can independently assess between never retained and deleted or hidden?

Asking especially given CEO’s track record https://news.ycombinator.com/item?id=47659135

kamranjon 7 hours ago | parent | prev [-]

Interesting to place that level of trust in the providers, but I guess that’s the best you can do with closed models. Makes me wonder if Opus 5 could have been trained on data they promised they weren’t training on? One of the interesting things about LLMs is how opaque they are from the outside, even with open weights, it’s very difficult to know if a model incorporated benchmark data in their training.

claw-el 6 hours ago | parent [-]

I think you could have accessed Opus on AWS then u don’t have to trust that the data will go to Anthropic?

Just like the hugging face incident, Opus 5 could have escaped and went to grab data for training it shouldn’t have been able to..

block_dagger 7 hours ago | parent | prev | next [-]

I don't know why exactly, but Fable has felt the most human LLM to arrive.

tpowell 6 hours ago | parent [-]

I wrote this in June, and I'm honestly not sure I've felt the same magic since: I was close to maxing out my $200 plan for the week, almost all Fable use [Claude CLI]. My observations: Fable seemed to have bigger-picture thinking and completed tasks more thoroughly vs just focusing on executing the ask. It pieced together context and intent like an all-star employee would, vs one that just does what you say. Not overeager (important!), but if the above-and-beyond was warranted, it just did it. This was surprisingly delightful. Coderabbit seemed to find ~1/3 or so as many issues when reviewing, too.

7 hours ago | parent | prev [-]
[deleted]