Remix.run Logo
▲ bpodgursky 2 hours ago

Yes, both labs already have monstrous internal teacher models they don't sell for inference, this is generally acknowledged. They cut releases for the public just to keep revenues growing, it's not their actual frontier.

▲sowhat1 2 hours ago | parent [-]

Ok. Given this hypothesis, why is the software they release generally considered crappy by competitor standards, benchmarks, and open source standards?

Claude Code is an awful codebase, has leaked its own source code multiple times, and scores the worst on number of tokens burned vs pass rate percentages.

Is anyone even using their Figma competitor?

▲monocasa 2 hours ago | parent | next [-]

Probably because bad code that you create initially without thinking that it's a core piece of your stack becomes depended on for its crappy behavior, and then you can't change much without breaking workflows.

I seem to recall Fred Brooks talking about that experience with OS/360 JCL (maybe just straight up in The Mythical Man Month?).

▲CoolestBeans an hour ago | parent [-]

If agents really are superpowerful at programming tasks why not just have it rewrite the tool that the majority of your customers use and have it recreate the bugs? I mean presumably its the primary force behind the current version so what's the major cost there?

▲XenophileJKO 20 minutes ago | parent [-]

I imagine it comes down to economics.. there isn't much upside to fixing the last 20% of issues that the dumber faster models are missing.

The cost to serve, latency profile ,and internal demand for a maximal intelligence model would probably keep it pointed at harder and more valuable problems most of the time.

▲addaon 2 hours ago | parent | prev | next [-]

Because even beyond-frontier LLMs are bad at software.

▲owebmaster 2 hours ago | parent | prev | next [-]

> Is anyone even using their Figma competitor?

Yes. CC started to create canvases without me asking. The mockups look good (it's just html+css), the tool is vibecoded crap

▲TeMPOraL an hour ago | parent [-]

I mean, on the one hand the tool may be vibecoded crap - I don't know, haven't checked, taking it at your word.

On the other hand, Opus 5.5 cracked zero-shotting proper LCARS interfaces that near-perfectly adhere to the franchise "design language" even in tiny details, while simultaneously being 100% functional following my admonitions about Airbus cockpit design rules and nuclear reactor control room standards.

So yeah, why wouldn't I use it? It works spectacularly well.

▲IanCal 2 hours ago | parent | prev [-]

How many people does it stop from using the software?