Remix.run Logo
▲ minimaxir 6 hours ago

> Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing

This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

▲joshstrange 6 hours ago | parent | next [-]

> 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

Cache doesn't help you much when you are compacting every 5 minutes...

I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).

▲redox99 5 hours ago | parent | next [-]

If you run out of sol medium with $100 you're doing something wrong. Astra destroys your usage, I get 1 day of usage with Astra, but 6 sol is almost unlimited and I only use xhigh.

▲Aeolun 2 hours ago | parent | next [-]

It’s only nearly unlimited if you haven’t just used a banked reset. After a banked reset your weekly usage gets cut by about 80% (not the week you need to wait to get your normal limits back though). ChatGPT has given me a really good reason to cancel.

▲threecheese 39 minutes ago | parent [-]

Can you elaborate? I've been getting great usage out of my $200/mo plan, and thought I'd try a reset (first time) which was expiring just for giggles. Am I going to get only 20% of it effectively?

I overused Astra in order to drain my weekly, figuring I'd have the reset. (not wastefully, I did get more work done)

▲jorblumesea 3 hours ago | parent | prev | next [-]

yeah I use sol constantly and have done maybe $15 of spend in the past week. it's solid and cheaper. this is at least 4-5 investigations, prs, whatever per day.

▲shimman 4 hours ago | parent | prev [-]

"You're holding it wrong." Is hardly a retort from a real paying customer having problems with their paid services.

This is why these companies are struggling to make money, they're chastising their customers just like they've been chastising the human race.

▲trio8453 2 hours ago | parent [-]

> "You're holding it wrong." Is hardly a retort from a real paying customer having problems with their paid services.

It's very appropriate in the cases when you're holding it wrong. The fact that you're paying doesn't mean that you can't make mistakes or waste resources.

▲onlyrealcuzzo 4 hours ago | parent | prev | next [-]

If you're compacting every 5 minutes, you have a workflow problem - period.

No LLM will be cost effective if it's compacting this often. You have to find a way around it.

▲ngruhn 4 hours ago | parent [-]

Context window is only 275k or something. And honestly compaction is not that bad in Codex. I often don't even notice I went through 5 compactions in a session.

▲SyneRyder 2 hours ago | parent | next [-]

Sounds like that's the problem then, 275k is a tiny context window. I regularly have sessions that go to 450k or even up to 700k for an unattended overnight Claude Opus session.

Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:

https://x.com/thsottiaux/status/2089082893804896524

There's at least a forum thread about it here:

https://community.openai.com/t/why-does-codex-report-a-258-4...

▲gf000 an hour ago | parent [-]

But that 250k context worth way more than 1M in terms of how well it's utilized, so actually I do like codex trying to keep you at that sweet spot.

▲jeremyjh an hour ago | parent | prev | next [-]

I don’t usually have a problem doing a complete task in that context size. OMP does make a lot of use of rewind which may be helping - basically forks itself and sends back a summary after a long tangent. Coding tasks use a Luna max agent.

I’ve also found compaction not to be a problem when it does happen.

▲threecheese 38 minutes ago | parent [-]

How do you trigger this? I've been messing with OMP lately for funsies.

▲onlyrealcuzzo an hour ago | parent | prev | next [-]

If it's compacting every 5 mins, you're going to notice it in your cache miss ratio and your costs...

It also presumably means it's regularly not able to get everything it wants to have to make decisions in context, which means it's going to perform poorly...

▲sally_glance 2 hours ago | parent | prev [-]

Same for me, I started wondering if maybe workflows using compaction instead of clear + markdown memory would be more efficient. Writing a plan or tasks to a file often has the next session repeat part of the exploration, compaction seems to keep most relevant context.

▲manmal 3 hours ago | parent | prev | next [-]

Your tool calls (MCPs?) are very likely too wasteful. Apply some filtering logic on the offending tool’s output. Either a wrapper CLI, or just tell codex how to filter.

▲AmazingTurtle 4 hours ago | parent | prev | next [-]

you can actually leverage 400k and 1M contexts in codex with very little code changes to the harness. note that excess context past the.. 250k or 400k mark (i don't remember) is charged at 2x the price.

▲apitman 5 hours ago | parent | prev | next [-]

You have a lot of control over compaction, both directly by changing compaction settings, and indirectly by how you structure your codebase/docs so agents use less tokens.

▲codewithcheese 5 hours ago | parent | prev | next [-]

you can config codex to compact at a higher context limit

▲_davide_ 4 hours ago | parent | prev | next [-]

As a reference i burn 1% percent for every 40 minutes of sol on average

▲antonvs 5 hours ago | parent | prev [-]

Try Gemini. It’s so cheap I often use my personal AI Pro account for corporate work, and most of the time it doesn’t matter.

▲ChickeNES 5 hours ago | parent [-]

Gemini is dumb as hell though, it's not like for like

▲Marha01 4 hours ago | parent | next [-]

Gemini 3.8 Flash is actually pretty good.

▲Foobar8568 4 hours ago | parent | prev [-]

cheerleader hallucinating agent. That's Gemini.

▲TuxSH 6 hours ago | parent | prev | next [-]

Exactly half as expensive as Opus 5.5 in every API pricing metric

▲bigwheels 6 hours ago | parent | next [-]

And half as good. I didn't have great experiences with Anthropic models in the past, but Opus 5.5 seems to have turned a major corner. It is churning through tasks significantly more quickly and efficiently.

Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.

Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).

▲dotancohen 6 hours ago | parent | next [-]

  > Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does.
That's far too vague. I found Opus to be terrific at coding, but human text just seems so robotic with it. OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural. What is "something difficult" in your workflow?
▲notatoad 2 hours ago | parent | next [-]

My side by side evaluation this week was to build a tool for mounting my app’s UI components in a headless chrome and feeding mock data into them, for the purpose of taking screenshots for help docs. Not super complicated, but a real task I needed done.

I gave the task to codex first, sol 6 xhigh. it took a couple back and forth prompts to define the project and then it worked for a bit and to took a couple more prompts before I decided it was good enough - not perfect, but close. It re-implemented some wrapper components in a simplified way that lost some of the UI, but it would work.

Opus 5.5 high took the same prompt with no back and forth, it just went off and one-shotted a tool that takes pixel-perfect screenshots of exactly what my app looks like.

▲peterbell_nyc 6 hours ago | parent | prev | next [-]

You HAVE to have a set of personal evals for each class of task you want to use models against at scale so you can test plausible candidates and compare output on your work against your evals.

There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.

▲Starlevel004 5 hours ago | parent | prev [-]

> OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural.

This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.

▲beering 5 hours ago | parent | prev | next [-]

This news and thread is about 6.1 Sol, not 6 Sol. You haven’t even had time to do a fair comparison yet.

▲TuxSH 6 hours ago | parent | prev | next [-]

> Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.

Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.

▲mmis1000 6 hours ago | parent | prev | next [-]

For my personal experience, antropic model have better user experience except for 4.7 and 4.8 though. 4.7 and 4.8 feels like expensive downgrade of 4.6 to me (I didn't know why these two should even exist)

However it's less willing to obey your instruction so it's less usable for general runtine flows.

▲krzyk 5 hours ago | parent [-]

For me Anthropic models from 4.7 to 5 including where bad and ate tokens like crazy. Task delivery was worse than GPT 5.6 and token usage was 2-3x higher.

Looks like 5.5 is the new 4.6

▲sobiolite 5 hours ago | parent | prev | next [-]

Are you comparing Opus 5.5 with GPT-6 Sol or GPT-6.1 Sol? Because they are different models.

▲Infinity315 6 hours ago | parent | prev | next [-]

I'm not an OpenAI simp, but how anyone can have any opinion on the performance of these models in less than a day - let alone a few hours - is beyond me.

▲phoghed 6 hours ago | parent | next [-]

I think it’s one of the reasons why you often see people decrying the lessening capabilities of the models a few weeks later, despite there being 0 proof of any changes, and evidence of the models staying the same from sites that track it.

They form these super strong opinions after a few prompts, then face reality over time.

People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.

▲toasty228 6 hours ago | parent | prev | next [-]

Try it, it's that good compared to openai current offering.

I get better results and usage our of my $20 claude sub than my $100 openai sub... it's that ridiculous

▲copperx 6 hours ago | parent [-]

The usage allowances are now insane, like they were when the Max plans were introduced. The $100 plan is usable again for real tasks.

▲rspeele 6 hours ago | parent | prev | next [-]

While I have no experience comparing this brand-new model, OpenAI themselves call it "near-Astra" intelligence. I set Astra and Opus 5.5 independently working on the same large research/coding task in an experimental project (doing NURBS surface modeling stuff). They had the same starting repo state, same task packet, same test suite to try to meet. I have the $100 plan in both.

Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output.

The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.

The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind.

Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.

Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.

▲this_user 2 hours ago | parent | next [-]

Astra doesn't just burn token at an insane rate, it is also strangely high maintenance when using it. Occasionally, you have to keep prodding it to keep working. Then at other times, it will disappear down some rabbit hole, trying to resolve increasingly hypothetical issues. It feels like you constantly have to keep it on track, while Opus is just churning through tasks.

▲agar 5 hours ago | parent | prev | next [-]

This was a very interesting, informative, and well-written comment (and experiment). Thank you.

▲chaostheory an hour ago | parent | prev [-]

> I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer).

Going on a slight tangent, I find that I get the best results when I force Codex models (Astra/Sol) and Claude models (Opus/Fable) to consult each other (just have them build a simple skill). There are tasks that neither one can fully solve on their own, but their differences are large enough to make a difference when they collaborate.

▲rspeele an hour ago | parent [-]

I strongly agree!

My biggest conclusion from this test was: the most efficient use of my weekly Astra budget is as a reviewer/consultant for work done by Opus. I don't have Astra write much code right now, but I do have it reading a lot of what Opus writes. Of course with the way the AI landscape shifts the balance could be the exact opposite 2 weeks from now, but either way having 2 "smart" models available from 2 different companies is a boon.

Seeing how each model preferred its own flavor of code shows that, even from a "blind" fresh context, a same-model reviewer will still often look at the work of another incarnation of itself and go "yep that's how I woulda done it" and not be as likely to realize that there was an alternative path or implicit assumption/mistake in the work.

▲colinhb 6 hours ago | parent | prev | next [-]

Yeah totally agree, people keep jumping in w/ strong views hours after release, eg: https://news.ycombinator.com/item?id=49045430

▲beering 5 hours ago | parent | prev | next [-]

They’re comparing against the previous model, not the newly released one (6.1). Why do that on a thread about the new model, I don’t know.

▲ 6 hours ago | parent | prev | next [-]
[deleted]
▲ex1fm3ta 5 hours ago | parent | prev | next [-]

benchmarks.

▲AndrewKemendo 6 hours ago | parent | prev [-]

Only takes 5-10 minutes to test your favorite one shot comparison prompt.

▲edgyquant 6 hours ago | parent | next [-]

Can you give an example? For me I find that one shot prompts are pretty good it’s only when working with large codebases and complex, multi prompt workflows, that I find the real limitations of models

▲AndrewKemendo 5 hours ago | parent [-]

Yeah the whole Pelican riding the bike is the best obvious one

▲squidbeak 6 hours ago | parent | prev [-]

If 5-10 minutes is enough, you need a more ambitious one-shot goal.

▲jauntywundrkind 6 hours ago | parent | prev [-]

A pity I have to use claude code to try this, that I can't use the tools I know and love and have built around (opencode).

(I did use some CC for Fable when it came out, and it was... ok. Not the worst thing ever.)

▲dom96 6 hours ago | parent | prev [-]

Based on my benchmark[1] it is the same price as Opus 5.5 and just as capable.

1 - https://bench.killswitch-lang.org

▲zeroonetwothree 4 hours ago | parent [-]

Opus 5 scoring higher than 5.5 makes me question of the value of this benchmark to real world usage

▲dom96 4 hours ago | parent [-]

Well, it is genuine.

Opus 5.5 fails the "understanding" tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1/0 instead of the max numbers. Opus 5 gets it right.

Here are the outputs from both: https://gist.github.com/dom96/b5bce82b6e6c1ebd5271ed70ad941b....

Looking at that Opus 5.5 fails to deduce that the "hack statement" is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5's greater intelligence.

▲verdverm 6 hours ago | parent | prev | next [-]

cache is typically 10%, is this OAI setting a new level at half, 5%?

▲crazylogger 6 hours ago | parent [-]

The backdrop being deepseek offering 1% (I remember it was ~1% when 4-pro first came out early this year - 4-pro is now removed) / 2% (current for 4.1-flash).

▲sscaryterry 6 hours ago | parent | prev [-]

[flagged]

▲user43928 6 hours ago | parent | next [-]

It's obviously true.

With the 80% price cut, this is competitive with Opus 5.5 despite the subscription downgrade.

Additionally, it was said that existing 20x subscriptions retain the higher limits for some time.

I have seen you make these immature accusations that users here are OpenAI employees multiple times today.

▲sscaryterry 6 hours ago | parent [-]

It is not obviously true. Please provide real proof. OpenAI's customers are tired of their BS.

▲JimDabell 6 hours ago | parent | prev | next [-]

> most people, get hardly a days usage out of a 20x account

This is not even remotely true.

▲peterbell_nyc 6 hours ago | parent [-]

This is the distribution of usage. Spin up a bunch of loops or fire a semi-autonomous factory at a project and it's pretty easy to blow through a 20x account in a few hours if you can afford the sandboxes, CI and other infra required.

If you're running 2-3 parallel agent session with a few sub agents and waiting for you to prompt them, you'll have a very different experience!

▲JimDabell 6 hours ago | parent [-]

> Spin up a bunch of loops or fire a semi-autonomous factory at a project and it's pretty easy to blow through a 20x account in a few hours if you can afford the sandboxes, CI and other infra required.

This is a tiny minority of people, not “most people”.

▲minimaxir 6 hours ago | parent | prev [-]

if an openai employee is reading this plz hire me i am unemployed and i need a job

(Usage limits are entirely dependent on what you're doing with them. If you're not running it on 1 million LoC codebases you can get a lot of mileage out of even a 5x account particularly with the recent cheap models)