| ▲ | alephnerd 2 hours ago | |||||||
Most of my and my peers PortCos run their own eval and benchmark sets, simply because they know what they need best. The reality is, capabilities have largely converged across foundation models over the last 18 months, and much of the value add is coming from the harness layer itself now. This has been the operating assumption for me and my peers, and has largely played out that way. That said, this has always been an issue with benchmarking since the very beginning. DB Benchmarks, compute benchmarks, and others that were external facing were always inherently a content and product marketing tool. The actual internal benchmarking used to model, understand, and enhance your product was always a closely held secret. Most of these conversations are happening, but largely in person and not on HN. | ||||||||
| ▲ | NitpickLawyer 8 minutes ago | parent | next [-] | |||||||
> capabilities have largely converged across foundation models over the last 18 months For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25. It's been a while since we've heard the old "models have stagnated". Oh well. | ||||||||
| ||||||||
| ▲ | matan0904 an hour ago | parent | prev [-] | |||||||
[flagged] | ||||||||