| ▲ | Infinity315 6 hours ago | ||||||||||||||||||||||||||||
I'm not an OpenAI simp, but how anyone can have any opinion on the performance of these models in less than a day - let alone a few hours - is beyond me. | |||||||||||||||||||||||||||||
| ▲ | phoghed 6 hours ago | parent | next [-] | ||||||||||||||||||||||||||||
I think it’s one of the reasons why you often see people decrying the lessening capabilities of the models a few weeks later, despite there being 0 proof of any changes, and evidence of the models staying the same from sites that track it. They form these super strong opinions after a few prompts, then face reality over time. People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work. | |||||||||||||||||||||||||||||
| ▲ | toasty228 6 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
Try it, it's that good compared to openai current offering. I get better results and usage our of my $20 claude sub than my $100 openai sub... it's that ridiculous | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | rspeele 6 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
While I have no experience comparing this brand-new model, OpenAI themselves call it "near-Astra" intelligence. I set Astra and Opus 5.5 independently working on the same large research/coding task in an experimental project (doing NURBS surface modeling stuff). They had the same starting repo state, same task packet, same test suite to try to meet. I have the $100 plan in both. Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output. The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster. The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind. Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical. Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | colinhb 6 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
Yeah totally agree, people keep jumping in w/ strong views hours after release, eg: https://news.ycombinator.com/item?id=49045430 | |||||||||||||||||||||||||||||
| ▲ | beering 5 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
They’re comparing against the previous model, not the newly released one (6.1). Why do that on a thread about the new model, I don’t know. | |||||||||||||||||||||||||||||
| ▲ | 6 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
| [deleted] | |||||||||||||||||||||||||||||
| ▲ | ex1fm3ta 5 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
benchmarks. | |||||||||||||||||||||||||||||
| ▲ | AndrewKemendo 6 hours ago | parent | prev [-] | ||||||||||||||||||||||||||||
Only takes 5-10 minutes to test your favorite one shot comparison prompt. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||