Remix.run Logo
qnleigh 6 hours ago

> OpenAI / Anthropic models have largely stopped advancing

I'm shocked anyone could conclude this. This year it became common for people to entirely delegate coding to AI (I know many competent programmers/researchers who do this now). Progress in math has just been insane. An internal model at OAI just resolved one of the most celebrated open problems in mathematics. If anything, progress has accelerated.

glub 5 hours ago | parent | next [-]

> This year it became common for people to entirely delegate coding to AI

This has been the case for around 2 years now, more reliably - a year. We've mostly stayed there since then.

Saying that more people started doing it isn't indicative of significant improvement. Some people just started doing it later.

I can't speak about math because I haven't used AI for that application, but I know that there hasn't been any significant advancement in coding in this year on base models. There has been more RL work, more harness work, more tools, they all expanded some capabilities like cyber or orchestration or tool use, but raw intelligence of base models is no longer where the main focus is.

BobbyJo an hour ago | parent | next [-]

> This has been the case for around 2 years now, more reliably - a year.

I have to disagree with this pretty strongly. Opus 4.5 needed a lot of handholding not to work itself into a corner pretty quickly. Fable I basically never need to correct, and I've most become a data source.

itkovian_ an hour ago | parent | prev [-]

What are you talking about - I feel like we’re living in parallel realities. If I had to go back to opus 4.5 tomorrow I’d be hugely upset and significantly slowed down

bel8 4 hours ago | parent | prev [-]

I'm not. Yes we normalized 1m context window and models tend to hallucinate less.

But models have been somewhat stagnant since Opus 4.6/7.

And in some regards there were even regressions like Claudeisms that are load bearing.