Remix.run Logo
▲ benrutter 4 hours ago

> Spent over $400,000 in API priced tokens with GPT-5.6 Sol and GPT 6 Astra

Tangent here, but I think this bit is super interesting!

You could viably hire someone to do this work for that kind of money - I think the interesting thing is that substantially less interested/experimenting engineers would consider paying for a human to do this work, than would happily chuck a big amount of money into an LLM.

I don't have any suggestion about why that exists, but it's a strange and interesting contract.

▲flossly 4 hours ago | parent | next [-]

I found the whole section super interesting... I'll copy it here:

I used a lot of OpenAI models to try and complete this port. In total I did over $400,000 in API priced tokens with GPT-5.6 Sol and GPT 6 Astra. They wrote over 1.3m lines of Rust over multiple months of /goal loops and never got past like 84% compat.

When I saw how little my Claude Code limits were burning, I figured it'd be fun to throw Opus 5.5 at this. It had a working v0 in 10 hours.

I assumed it kept using the code the Codex models wrote. I was wrong. Opus 5.5 started from scratch. It got further than Astra in 1/10th the time.

I let it keep going, and it definitely did. Total token spend was ~$24,047 of API spend over 2 weeks. I was using my Claude accounts, and it worked out to somewhere between 925% and 983% of my $200 plan weekly limits.

Expensive, for sure, but not that bad considering how much work has went into typescript-go.

▲flossly 4 hours ago | parent | next [-]

Great advertisement for Claude Code...

▲sreekanth850 an hour ago | parent | prev | next [-]

Are we sure that he spend that much or he calculated the API pricing and used a subcription to implement this?

▲semiquaver an hour ago | parent | next [-]

We’re sure that he did not spend that much and did use (sharded) subscriptions, because he said so.

▲TiredOfLife 43 minutes ago | parent | prev [-]

The GPT ones were provided by OpenAI for evaluation. Claude ones were pooled subscriptions.

▲Topfi 3 hours ago | parent | prev [-]

> I assumed it kept using the code the Codex models wrote. I was wrong. Opus 5.5 started from scratch. It got further than Astra in 1/10th the time.

Ouch, that is brutal and honestly, quite embarrassing but confirms what I have been seeing for a while. Personally, I find output from current OpenAI models still very hard to parse (though it has gotten better vs the pre-trains from both labs in mid/late 2025), thus hard to truly understand, verify and get comfortable maintaining vs current Anthropic models. I do occasionally see a higher ceiling in well scoped tasks with OpenAI models at the cost of (frequently) deviating from the original prompt in (sometimes) very destructive ways.

Could be that this hard-to-read output doesn't just go over my limited capacity/skills but with current models can become simply impossible to untangle beyond a certain size even when one has (essentially) infinite resources via multiple subs and different models.

Would also work with my suspicions for why OpenClaw (mainly build with Opus 4.5 and its post-trains) has been this hard to truly "fix", requiring highly paid Nvidia engineers, multiple months, (literally) infinite resources from OpenAI including access to internal models and yet still holds records for CVEs. Heck, another one was found just 7 days ago after what I'd argue was one of the most extensive hardening sessions any piece of software has ever undergone.

Makes my (multiple) decisions to start from scratch more than once on a major reworking of the existing tabbing interface in Firefox a bit less painful. Learned with each, found gaps in my knowledge, thanked the amazing docs the Firefox devs have been maintaining for decades and while starting from 0 was painful, getting back to MVP is easier than ever. When I hit a point were I was starting to struggle to truly parse additions a model was making to the patches applied to Firefox source code (even if they worked), I always found that pushing even slightly beyond that would incur painful, but hard to notice regressions, introduce major DB maintenance burdens as some models struggle to understand that in development regressions and incompatibility are acceptable and schema transitions aren't needed pre-release (still a case with GPT-6.1 Sol, less so post Fable for Anthropic), make me uncomfortable concerning privacy/security/data loss prevention (especially as I have seen Fable 5 cut some privacy/proper data removal corners in simple CRUD) and simply take away my control about the implementation. Wouldn't feel right to release something in that state, what I have now is fully understandable and thus could be maintained even without models.

Still expecting bugs of course, massively dreading security findings or even worse, possible data loss given browsers handle some of our most important personal+professional data and will surely have taken some embarrassing approaches that might have a much more performant solutions when implementing an infinite canvas of webpages, but still, rather that then also knowing I wouldn't even know where to start understanding a feature.

More so if, like with ts-rust, even (nearly) infinite tokens couldn't get me unstuck.

▲Rapzid an hour ago | parent | next [-]

I think it's hard to know what to make of that. Observations with little insight.

Maybe it just needed to be prompted differently? Maybe Astra starting from scratch could have done it? Maybe Opus was somehow trained more on the Golang implementation.

It's curious, but again little insights.

▲scotty79 24 minutes ago | parent | prev [-]

> Ouch, that is brutal and honestly, quite embarrassing but confirms what I have been seeing for a while.

What's most interesting is that Opus wasn't told to toss out the Codex crap. It decided to do that on its own. Which might mean that LLMs are actually getting a reasonable taste.

I've seen on youtube that some guy implemnted game engine with Astra and Claude. Astra wasn't really given fair chance because it didn't decide to work as long as Claude and the guy didn't force it. But still, scripting code within the engine that came from Astra was chaotic, messy, piling up things, but the code that Claude made looked downright pleasant.

▲____mr____ 4 hours ago | parent | prev | next [-]

This is Theo's project, and he has been very transparent about how he is only doing these things because api prices are heavily subsidized by subscription pricing. The reason people aren't willing to pay 400k to an engineer to do this work is because they aren't paying the AI this much.

I don't think I've seen an enterprise paying API prices doing these sorts of rewrites, and honestly, I think until that happens I will remain skeptical about LLM AI's viability as a profitable business.

▲bewareofscams 4 hours ago | parent [-]

[dead]

▲viraptor an hour ago | parent | prev | next [-]

He's not paying that much in practice. But even if he did, you couldn't hire 3 engineers (I'm assuming you'd get 3x 133k each) that could do this work in anywhere close to this amount of time. It would be a multi-month long project. So "to do this work" is not really comparable.

▲true_religion an hour ago | parent [-]

He already said it was a multi-month project that used multiple models. So it feels very comparable to hiring multiple humans for multiple months.

▲chadcmulligan 4 hours ago | parent | prev [-]

I've had the same thought about the maths - they've spent millions on tokens, I can't recall anyone spending millions on mathematicians.

▲binlog an hour ago | parent | next [-]

NSF by itself contributes $200-300M a year towards math grants. Universities spend a billion+. Add corporate funding and other sources (DoD, other government agencies, private foundations) and you’re easily looking at a few billion dollars a year just in the US going to mathematicians.

▲fiforpg an hour ago | parent | prev [-]

To be fair, NSF grants for pure mathematicians were never as substantial as for applied sciences, but they would routinely reach 500k–1M territory. That's the amount going to a single proposal, and grant panels in, say, analysis would definitely spend millions per application season. NSF funding data is public and can be looked up.

I'm using past tense because iirc NSF has recently reduced their funding volumes, although I'm not following the situation very closely.