Remix.run Logo
scotty79 4 hours ago

Do you draw that conclusion from the fact that AI surprisingly quickly reaches the end of each ruler we try to measure it with?

otterdude 4 hours ago | parent | next [-]

Its not really that surprising when models are trained on the exams

astro1234 4 hours ago | parent [-]

I work in AI evaluation, lots of problems and leakage is an issue as is ecological validity, but they definitely do not explain the progress we see.

I think Epoch has the best analysis I’ve seen on evaluation trends; they use IRT to basically model a variety of benchmark difficulties, and then model a capability parameter for each model. This is as robust a sort of “meta-study” of evaluations as I’ve seen and the trend in capabilities show no sign of slowing down.

So I think people’s feelings clash with reality, and that’s because releases are more frequent and the jumps between releases are smaller, but the growth in capabilities _over time_ has not changed for the better or worse over a very very long period of time.

otterdude 4 hours ago | parent [-]

Benchmarks saturate around 80-90%?

This is not "Acing" a test, this is hitting a wall.

2 hours ago | parent | next [-]
[deleted]
scotty79 4 hours ago | parent | prev [-]

Even on very small tests a fraction of questions might have wrong answers in the key.

If models can't get more than 90% of the benchmark right I think it's a strong indication that they were not trained on the answers and that benchmark itself is messy enough that <10% desired answers might be wrong or misleading.

astro1234 2 hours ago | parent [-]

Yea this may explain part of it or all of it, it’s likely a case by case kind of thing.

Also to respond to the parent comment: benchmarks have a variety of difficulty levels. Humanity’s Last Exam, though now hitting the beginning of a saturation phase with Fable, was long unsaturated while other benchmarks saturated awhile ago. So that’s what I meant by Epoch capability index: using IRT models this effect so that you gather robust signals from variety of benchmark difficulties and can track progress over time as model capabilities have evolved (and so benchmarks have had to evolve to keep up).

But yes like I was saying: all benchmarks are problematic, some are useful. Benchmark quality problems abound, so 90% being the true ceiling is not surprising. There may be other factors at play here too, I haven’t studied this problem that deeply to have a good thorough answer to this. But keep in mind there are probably 50,000 benchmarks in the literature and that is not a joke number. A crapload of noise in that signal but it’s not all noise.

freejazz 4 hours ago | parent | prev [-]

Can't call it AI like that without discrediting yourself. You mean LLMs?

otterdude 4 hours ago | parent | next [-]

Jumping in here, frankly I hate the trend of calling every type of automation intelligence.

Most "AI" is really an optimization algorithm in software tools, same as its always been. This really isnt anything new, aside from adding a chatbot / MCP interface to the same tools.

scotty79 4 hours ago | parent | prev | next [-]

When a Big Killing Robot comes to murder you be sure to always call it BKR and don't discredit yourself by calling it AI.

freejazz 3 hours ago | parent [-]

That's a bit hyperbolic when we're all just posting on HN

logicchains 4 hours ago | parent | prev [-]

Talk about moving the goalposts. Pray tell, exactly what must an LLM do before you're willing to consider it AI? Be specific, otherwise you're just woo-mongering.

freejazz 3 hours ago | parent [-]

Everyone here is talking about LLMs, why bother calling them something else