Remix.run Logo
▲ whatever1 5 hours ago

Even if you are competent I cannot review your 5,000 lines of code you produce per day vs the 100 you were producing before the LLM apocalypse.

▲rfgplk 5 hours ago | parent | next [-]

5,000 is the output velocity of someone not fully immersed in agentic coding. I've seen repos do ~100k to ~250k loc changes per week.

▲0c3ca83 5 hours ago | parent | next [-]

Yes, they're certainly squeezing 500 lines of functionality into 250,000 lines of code. Agents are great at this.

▲bitwize 4 hours ago | parent [-]

Tell me you're not using a frontier model without telling me you're not using a frontier model

▲necovek 30 minutes ago | parent | next [-]

I've asked Codex with GPT 6 Astra to *review" a one time benchmarking script for any mistakes (built by Claude Code using Opus 5.5) and it refactored the shit out of it claiming all sorts of stuff without even asking about the context in which it was developed.

If I was to employ them to review the code without giving each the same baseline multi-page prompt, they go into endless loop of "improvement" with no end goal in sight.

More and more frequently, I instruct frontier models to stop and go back to the task at hand.

▲whateveracct an hour ago | parent | prev | next [-]

i have unlimited tokens and i throw Fable / Astra at everything. They suck ass still for anything nontrivial. I could commit that garbage but if I kept doing it, I will end up with a ball of mud only Fable / Astra can grok..convenient for Dario and SamA..

▲paganel 4 hours ago | parent | prev | next [-]

Where's the great software, then? I'm genuinely asking: where is it? Because I can't find it, and it's been close to a year since AI for programming has started to take off.

▲dawnerd 4 hours ago | parent [-]

In fact, a lot of the once great software that's switched to AI driven development has gotten worse.

▲cat-snatcher 3 hours ago | parent [-]

For example?

▲dawnerd 2 hours ago | parent | next [-]

Clickup. Windows. Github, VSCode...

Could be argued they were going downhill before but it's a much faster decline since 2023-ish

▲necovek 27 minutes ago | parent [-]

Somebody posted that GitHub status summary: in ten years, they averaged 10 incidents per month — this includes the last 12 months. In the last six months, average is 22.

▲0c3ca83 3 hours ago | parent | prev [-]

Google.

▲cat-snatcher 2 hours ago | parent [-]

It was going downhill way before AI.

▲intrikate 2 hours ago | parent [-]

That's true, but also doesn't prevent the accelerated decay that it seems to be undergoing since AI hit the scene en masse a few years ago.

▲0c3ca83 4 hours ago | parent | prev [-]

Tell me you never once bothered to look at the generated code without telling me you don't look at the generated code.

Luajit is under 80,000 lines of code.

▲teiferer 4 hours ago | parent | prev | next [-]

So where is all that new software? My laptop and phone run essentially the same software as 2 or 3 years ago. Yes there were some minor updates to some apps, but nothing faster than in the years prior.

Where does all that supposed productivity go?

▲Powdering7082 2 hours ago | parent | next [-]

Number of git pushes in GH is way up. 320M in Q1 this year vs 80M in Q1 or 2020.

https://innovationgraph.github.com/global-metrics/git-pushes

Just because you haven't installed new software doesn't mean that new software doesn't exist.

▲legulere 9 minutes ago | parent | next [-]

To quote the article: "Don't confuse motion with progress."

▲californical 41 minutes ago | parent | prev | next [-]

But also, having significantly more code churn doesn’t necessarily mean there is more or better software.

In fact, having more churn can lead to worse software due to diverging patterns and inconsistency

▲nemetroid 36 minutes ago | parent | prev [-]

The question was "where is all that new software?".

▲necovek 25 minutes ago | parent [-]

I see somebody asking "where is that great software", but I think it's obvious to everyone that there is more (mostly bad) software.

▲al_borland 3 hours ago | parent | prev [-]

The only real updates I've seen to anything have been AI features... so all this AI is only being used to add AI to stuff. Most of which the average person doesn't seem to want or use.

▲pretendscholar 28 minutes ago | parent | prev | next [-]

What kind of applications are people building that involve 250k loc a week? Genuinely trying to understand this.

▲mstaoru 17 minutes ago | parent [-]

So far I mostly see a metric sh.. ton of meta- and meta-meta-projects re-wrapping AI wrapper tools, with sloppy slogans like "One Model, Five Harnesses. Combined." or "You run in the park. Rrrunnnrr.ai runs your AI." Looking inside, out of 250k it's often 200k of verbally incontinent self-explaining comments; or "smart" redesign of builtins.

No new browser, no new iOS clone than runs on Android, no new easy to use DaVinci, no new CAD suite, no $5 SolidWorks clone, no redesigned K8s, no 10x performance speedup in Linux kernel.

▲whatever1 5 hours ago | parent | prev | next [-]

I mean these guys are not even pretending to be reviewing the code.

It just gets “reviewed” by an LLM, which will find a nitpick while ignoring the huge fire in the core of the design, force the planner to make even more sloppy code to cover for an irrelevant test case. Rinse old tokens and repeat until you hit limits.

▲bunderbunder 4 hours ago | parent [-]

This is exactly what I've seen.

For example, I recently got brought in to help with quality on a large-scale system that had been ported to a new platform with the help of coding agents. The project was completed and declared operational in record time, but soon after the business discovered that:

1. The promised scalability improvements did not materialize. Instead, it got worse.

2. Observability had been lost. The telemetry was no longer trustworthy.

3. Users stopped trusting it because it was producing incorrect outputs.

What I ended up discovering was that, while it scrupulously kept existing automated tests passing, any behavior that wasn't explicitly covered by a test was free to change any which way. And there were plenty of small things that weren't explicitly covered. Perhaps because the original authors thought they were so obvious and commonsense that they didn't need one, perhaps because mistakes happen. The why doesn't matter. The point is that reality is messy and imperfect, so giving someone a chance to look at things and think, "Huh, that's funny..." is an essential part of defense in depth.

The real worst part was, this whole replatforming was a huge waste of time, anyway. The improvements they were looking for could easily have been accomplished with some controlled incremental changes to the original system. Mostly just removing a few basic and well-known performance antipatterns.

But way back at the outset, the person in charge of the project asked their agent, "What's the best way to X," and the agent gave them a trendslop answer about how Y alternative technology is more scalable and we should just port to that. It was convincing and they were under intense time pressure to just ship some code because leadership is bought into the AI hype and now has the patience of a 4 year old, so they just went with it.

▲Ambolia 5 hours ago | parent | prev [-]

Can the users of the software even keep up at that point? We may have reached diminishing returns on software production, and not enough impact on the rest of the process.

▲bitwize 4 hours ago | parent | prev [-]

That's okay. Reviewing the code will become the agents' job as well.

A couple more step functions in model capability of the type we've seen in the past year, and there will pretty much be no reason for humans to be involved in the development process at all. All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.

▲weatherlite 2 hours ago | parent | next [-]

> All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.

Kinda what i'm doing already, but for the young startup I'm at that's surprisingly tons of work. I miss the days we wrote code by hand boy those were fun 8.5 hours workdays.

▲necovek 24 minutes ago | parent | prev | next [-]

> All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.

Sounds like the easiest thing in the world: I wonder why did we not think of it earlier?

▲xpct 4 hours ago | parent | prev [-]

"A couple more step functions" is doing a lot of heavy lifting here

▲pydry 3 hours ago | parent [-]

It's the AI bro's mantra.

It didnt go wrong

And if it did, it was because you werent using the latest model.

And if you were, it was because you didnt have the appropriate guardrails.

And if you did, it's because you didnt have AGENTS.MD.

And if you did, it's because you didnt prompt it properly.

And if you did, it you're still going to be redundant soon because I'm sure the next model released will fix whatever went wrong.

▲mywittyname 2 hours ago | parent [-]

I've stopped calling out Claude mistakes on team meetings because this is so true.

I mean, sure, I could have predicted in what ways an LLM would fuck up, but there's just so many ways I can't keep up.

We just had a major production issue because someone's LLM wrote queries against dev databases. Which are very obviously dev databases because they are labelled with dev in the name, and in the table descriptions. AI reviewer didn't catch it, neither did the human reviewer for that matter.