| ▲ | simianwords 4 hours ago | ||||||||||||||||||||||||||||||||||||||||||||||
Sorry, this is ridiculous. OpenAI said that they have a step change in model performance. They proved it by solving a Millenium problem. HF incidents are public and vetted. If you still think there's a stall despite all evidence pointing to opposite, I don't know what to say.. | |||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | threecheese an hour ago | parent | next [-] | ||||||||||||||||||||||||||||||||||||||||||||||
If you believe Yann LeCun and David Silver (and others), there's an architecture wall; maybe labs are starting to see this on the horizon. Maybe the DRAM supply constraints are forcing it. Are these real step changes - big picture wise, or refinements in RL/agentic orchestration/"taste" and advancements due to bigger models and hardware technology/capacity scaling? If they not, does this tactic - and hardware improvements - continue to scale non-linearly like they need to? It is clear that whatever does change in each model increment has resulted in meaningfully better end user capabilities (as well as regressions in some areas, honestly), but that doesn't prove anything. I'm not sure what I personally believe, but stating with your full chest that a stall is ridiculous ignores a lot of potential evidence to the contrary. Stupid example: Astra. Its main improvements are: much much better computer use and 3d modeling capabilities; better subagent orchestration; better and more reliable tool use; slightly worse coding. This looks to me, from a feature perspective, to be an incremental improvement across several functional areas, plus new features which are unquestionably excellent but are probably the result of RL focus, not magic. Step change? Ehhhh depends on how you squint. But how many more iterations of this do we have? Are we going to squeeze quintillion parameter transformers into GPUs? | |||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | tuvix an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||||||||
Solving millennium problems doesn’t pay the bills. Their best models might be quite good when directed at extremely difficult and very focused problems, but most business use cases are nothing like that. Edit: Not trying to imply these new frontier models aren’t also better at other things, just that there’s really no reason for most people to use them when cheaper alternatives exist that get the job done just as well. | |||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | 2 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||||||||
| [deleted] | |||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | vrganj 4 hours ago | parent | prev [-] | ||||||||||||||||||||||||||||||||||||||||||||||
What reason could OpenAI possibly have to lie about their own performance? The fact they were desperate enough for a PR win they stole the research leading to the Millenium problem further hurts their case imo. | |||||||||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||||||||