Remix.run Logo
simianwords 4 hours ago

Sorry, this is ridiculous. OpenAI said that they have a step change in model performance. They proved it by solving a Millenium problem. HF incidents are public and vetted.

If you still think there's a stall despite all evidence pointing to opposite, I don't know what to say..

threecheese an hour ago | parent | next [-]

If you believe Yann LeCun and David Silver (and others), there's an architecture wall; maybe labs are starting to see this on the horizon. Maybe the DRAM supply constraints are forcing it.

Are these real step changes - big picture wise, or refinements in RL/agentic orchestration/"taste" and advancements due to bigger models and hardware technology/capacity scaling? If they not, does this tactic - and hardware improvements - continue to scale non-linearly like they need to?

It is clear that whatever does change in each model increment has resulted in meaningfully better end user capabilities (as well as regressions in some areas, honestly), but that doesn't prove anything. I'm not sure what I personally believe, but stating with your full chest that a stall is ridiculous ignores a lot of potential evidence to the contrary.

Stupid example: Astra. Its main improvements are: much much better computer use and 3d modeling capabilities; better subagent orchestration; better and more reliable tool use; slightly worse coding.

This looks to me, from a feature perspective, to be an incremental improvement across several functional areas, plus new features which are unquestionably excellent but are probably the result of RL focus, not magic.

Step change? Ehhhh depends on how you squint. But how many more iterations of this do we have? Are we going to squeeze quintillion parameter transformers into GPUs?

tuvix an hour ago | parent | prev | next [-]

Solving millennium problems doesn’t pay the bills.

Their best models might be quite good when directed at extremely difficult and very focused problems, but most business use cases are nothing like that.

Edit:

Not trying to imply these new frontier models aren’t also better at other things, just that there’s really no reason for most people to use them when cheaper alternatives exist that get the job done just as well.

2 hours ago | parent | prev | next [-]
[deleted]
vrganj 4 hours ago | parent | prev [-]

What reason could OpenAI possibly have to lie about their own performance? The fact they were desperate enough for a PR win they stole the research leading to the Millenium problem further hurts their case imo.

snaking0776 4 hours ago | parent | next [-]

Two things can be true. Open AI is a huge company.

1. It has people in it who are career obsessed and who are willing to ruthlessly go after any opportunity to improve their standing/stock valuation. Using a 2 week old model to snipe a millenium prize for PR is in line with that.

2. There are many researchers and even executives at the company who genuinely think we are speedrunning the end of the world. I don’t know anyone in this field who honestly argues that if we build ASI soon it doesn’t lead to extinction. This group of people can output warnings about the state of research and fears for the future while pushing for regulation out of genuine fear of what they’re building. I tend to agree with them.

You can’t view these companies as a monolith. They’re actions won’t be consistent because it’s built of many people with conflicting beliefs. Please look at arguments regarding AI risk and the current pace of progress and value them as it relates to the argument itself, not who said it. We are in a dangerous place and no one is sure how quickly we’ll get to a bad spot.

AnimalMuppet an hour ago | parent [-]

> I don’t know anyone in this field who honestly argues that if we build ASI soon it doesn’t lead to extinction.

Say what? Could lead to extinction, sure. Does? And nobody honestly argues otherwise? Baloney.

orangecat 2 hours ago | parent | prev | next [-]

The fact they were desperate enough for a PR win they stole the research

I guess people are just going to keep spreading misinformation about this, along the same lines as "Anthropic's C compiler fails on hello world". There is zero indication that they stole anything, unless it counts as "stealing" to spin up a bunch of compute based on rumors that the problem had been solved already.

2 hours ago | parent [-]
[deleted]
simianwords 4 hours ago | parent | prev [-]

they lied about their performance by.... solving a famously hard problem to solve?

Edit:

The "reputable academic" who accused OpenAI had to say this about LLMs and the recent result.

source: https://cims.nyu.edu/~tristanb/statement.pdf

> “the results are not the important thing.”

> “the important thing is instead the significance that a mathematician and an LLM model can now do all this work in a month.”

> “This is a Deep Blue-Kasparov moment.”

> “incredibly important developments.”

> “If indeed an OpenAI model did close the gap to Navier-Stokes, that is a remarkable thing and it should be said loudly, by them, with the history intact.”

Clearly Buckmaster (who is probably one of the most accomplished academics in the field) himself doesn't believe that AI has stalled. What makes you think you are right?

vrganj 4 hours ago | parent [-]

Solving is a very interesting way of describing what happened. Unless you believe them over reputable academics I suppose?

I will note the remarkable goal post shift in your edit - giving direct counter evidence to your own earlier claim of an OpenAI proof - and leave it at that.