Remix.run Logo
alephnerd 2 hours ago

Most of my and my peers PortCos run their own eval and benchmark sets, simply because they know what they need best.

The reality is, capabilities have largely converged across foundation models over the last 18 months, and much of the value add is coming from the harness layer itself now.

This has been the operating assumption for me and my peers, and has largely played out that way.

That said, this has always been an issue with benchmarking since the very beginning. DB Benchmarks, compute benchmarks, and others that were external facing were always inherently a content and product marketing tool. The actual internal benchmarking used to model, understand, and enhance your product was always a closely held secret.

Most of these conversations are happening, but largely in person and not on HN.

NitpickLawyer 8 minutes ago | parent | next [-]

> capabilities have largely converged across foundation models over the last 18 months

For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25.

It's been a while since we've heard the old "models have stagnated". Oh well.

tancop a minute ago | parent [-]

It's not "models have stagnated" but "models released at the same time are on the same level". Improvements are still real but the relative gaps between OpenAI, Anthropic, Meta, Grok, Gemini and open models are closer than ever. That doesn't mean progress is slowing down, it's just more widely distributed.

matan0904 an hour ago | parent | prev [-]

[flagged]