Remix.run Logo
ballsac 12 hours ago

> A challenge with this kind of study is that coding agents (Claude Code, OpenAI Codex) only started working really well in

2026? 5? 4? 3?

Heard this one way too many times.

ben_w 5 hours ago | parent | next [-]

I get the point having read much the same from Tesla (and fans) regarding self driving cars that still haven't done half the things that Musk said was just around the corner pending regulators a decade ago and repeatedly since then.

And myself I keep making comparisons between AI and the progress in 90s video games where every minor improvement got called "photo realistic" and then forgotten with the next game engine: https://archive.org/details/nextgen-issue-26

So I'm not gonna say "this is it" when the software quality really matters, and I absolutely won't speak to progress (or lack of it) outside of software.

But I will say "you can look around and easily see small businesses using AI to generate posters, quite a lot of small business software and websites are in the same category: the mistakes are real but increasingly don't matter".

qsera 4 hours ago | parent | next [-]

>the mistakes are real but increasingly don't matter...

I think it would start to matter once again. People will get fed up of AI posters and art. I think they already are...and once some threshold is crossed, the business won't dare to use AI generated assets/designs.

Turns out humans are much better at recognizing patterns in stuff that is generated ONLY using patterns from human generated content.

ben_w 3 hours ago | parent | next [-]

> People will get fed up of AI posters and art. I think they already are

Agreed, but will this look like a meme/fashion cycle? If so, re-prompt each year with a different look. Yes, there are still issues here, a friend found an image he was amazed was AI generated, but to me it was obviously so, so I showed him a screenshot of ChatGPT making something just it and included my prompt:

  create image: hand drawing of cute springer spaniel puppy looking sideways, various geometric shapes drawn in layer behind and in front of the puppy, all done in style of 7 year old using crayons with mediocre colouring-in skills
As I said to them:

  yeah, the line thickness feels AI, to me, the bad colouring-in scribbles feel like just the art style it was propmpted with

  it's like: it gets the big picture of the composition, and it knows how to colour in badly, but it doesn't know how to draw a dog as badly as the colouring in
> Turns out humans are much better at recognizing patterns in stuff that is generated ONLY using patterns from human generated content.

We're better at recognising patterns full stop. All biological brains are, and needed to be better than the current state of the art in machine learning because if a living organism was as poor at learning patterns as the SotA in machine learning, the organism would starve to death before being able to pick up anything and eat it.

AI also has a second disadvantage, because there are so few models: the laziest of ChatGPT "thinkpiece" blog posts being everywhere is hard to miss, and 5000 fake bloggers all prompting the same model with "find biggest news story of today and write a blog post about it in a way that maximises my ad revenue" will get 5000 almost identical posts. This will remain true while each instance of the most commonly used AI fail to talk to each other in a way that at least mimics them collectively getting bored with writing the same thing 5000 times, it does not depend on e.g. quality.

TeriyakiBomb 3 hours ago | parent | prev [-]

Will? They already have.

qwerpy 2 hours ago | parent | prev | next [-]

I’m probably illustrating your point but as a FSD fan it really got ”good enough” recently with version 14. The tipping point was suddenly, much more often than not, it can drive end to end from start (my garage) to finish (parked at destination) with no interventions. I can text and watch videos on my phone and as long as I glance up once a minute, it doesn’t complain.

Handling highway driving with lane changes was great when it got there years ago, but just in the last year or so it has gone from a nice to have to “from now on I will never buy a car that can’t do this”.

AI has hit some milestones for replacing work as well. There’s still many more to go and maybe some of them will never get hit (much like I don’t think a coast to coast drive with zero interventions during winter conditions is ever going to happen) but there are points at which it forever meaningfully changes some field of work. I think it’s there for writing code.

ben_w an hour ago | parent | next [-]

> I’m probably illustrating your point but as a FSD fan

Half-and-half. I'm not denying that self driving cars (and LLMs) are improving, I'm comparing it against the standards set by the biggest proponents. But yes, I have heard basically the same thing you just wrote for the previous several major releases of FSD.

Where we agree is that, while you are a fan, you do explicitly give as an example of something you think it will never do, something very close to what Musk has promised:

  "Ultimately you'll be able to summon your car anywhere … your car can get to you. I think that within two years, you'll be able to summon your car from across the country. It will meet you wherever your phone is … and it will just automatically charge itself along the entire journey."
- Musk, in Jan 2016: https://en.wikipedia.org/wiki/List_of_predictions_for_autono...

(That said, I think Tesla's FSD will never get there, not that it's impossible. The way Musk is behaving, there's going to be a financial scheme named after him in whatever passes for a textbook in 20 years, and it won't be the positive kind of example).

grim_io an hour ago | parent | prev [-]

Who is responsible for any accident happening while you use your phone during FSD?

It's not FSD until the human is no longer responsible.

This half measure bullshit is a joke.

w4der 20 minutes ago | parent | prev [-]

I digress with your last quote, I feel like some people who are aware of AI are starting to develop a quasi allergic reaction to slop, and while the mistakes might not matter for most of the population, for others it does and will take notice.

cmenge 3 hours ago | parent | prev | next [-]

I guess it's important who one hears this from.

I just spoke to a fried who is a headhunter and who's been trying to automate his processes for a while (he likes to fiddle and certainly has skills, but he's not an engineer). He kept trying, but it just wasn't good enough.

Now he said with GPT Work and Sol, it worked, but the key point is: all of it suddenly worked.

The problem was one of reliability, of handling edge cases. All previous attempts / model-harness-combinations were too brittle and needed too much observation and fiddling - cheaper to do it yourself.

Now he says "I don't know why I would ever hire a recruiter [the folks doing the cold outreach] again. I can focus on the candidate screening and acquiring projects, everything else is fully automated".

This doesn't come from an engineer or an AI lab, but a technically inclined power user, and I think this is where things get interesting.

johnyzee 35 minutes ago | parent | next [-]

You heard it from someone with no experience developing software. A lot of AI hype comes from that, even (or especially) from people in the actual business of developing software - a surprising amount of managers and adjacent or supporting roles in the field actually have very little clue about software development.

It's cool that 'regular' people can now create solutions to many small problems, and automate stuff - genuinely a step forward. Like Excel, only vastly better. But for bigger projects, real software engineers know that what LLMs do today is only a tiny part of development. And it solves it in a way that might well make the rest of the lifecycle a lot harder. It's like that saying about tools that make easy things easier and hard things impossible.

mstaoru 3 hours ago | parent | prev | next [-]

Then it just becomes a new baseline (everyone have access to the same LLMs), and recruiting moves up the philosophical ladder where human can add more value. What will it be? I don't know, I'm not a recruiter.

noosphr 3 hours ago | parent | prev [-]

Again I've heard this since 2022 when gpt3.5 came out.

This is like microprocessors in the 80s. Sure they double in capability every 18 months but the start is so pathetic it will be 30 years before they are good enough for everyday tasks.

siva7 14 minutes ago | parent | prev | next [-]

The timescale is well established: Late '25 was the start of agentic ai when capabilities of model + api + scaffold reached autonomous state. Any study comapring events before that timeframe is comparing apples with oranges.

cowanon77 11 hours ago | parent | prev | next [-]

It seems to be true this time though; I have observed it myself and heard it from several experienced developers I personally know and respect. It feels like some threshold was crossed with Opus 4.5 and Gpt 5.3, where the models are now able to reliably solve certain classes of problems that were previously unreliable.

Time will tell of course, and it’s early, but inflection points do exist with progress.

TeriyakiBomb 3 hours ago | parent | next [-]

Thing is. You can find an extremely similar paragraph written about Claude 4.x or some equivalent gpt. And simultaneously, many people expressing their frustration and the shortcomings of <insert any model>

“But it’s different this time” - several people, several times over the last couple of years.

This is not at all a dig at you, I’m very sorry if it reads that way. My point is these things only get truly better in anecdotes. The ways in which they fail is yet to change. Just yesterday I had gpt 5.3 generate completely awful code for the Cinema 4D Python API. Also an anecdote. But for all of the people saying they are truly intelligent and truly reason, they still make obvious mistakes, write around problems, fail entirely at architectural decisions, fail at random, generate FAR too much code.

And no amount of harnesses, methodologies, loops make much of a difference. If you listen to people on the internet they say it’s all working. You listen to people on the job and they mostly say it’s creating tech debt and a review bottleneck. Also burnout, so much burnout.

I think LLMs are mediocre. I think it’s fine they’re mediocre. You can work with low expectations. But the hype cycles are so tiresome.

ModernMech 29 minutes ago | parent | prev | next [-]

If it's so evident, why can't someone prove it with something more than "it seems better and everyone agrees"?

bluefirebrand 11 hours ago | parent | prev [-]

I wonder how much of it is real and how much of it is people just being worn down by the hype to the point they can't fight it anymore

Very smart people aren't immune to being worn down over time

simonw 11 hours ago | parent [-]

I really don't think that's how it works. Smart, experienced developers who thought coding agents were junk for most of 2025 and think they're useful now in 2026 are not saying that because they got "worn down over time".

TeriyakiBomb 3 hours ago | parent | next [-]

It tends to be when the training data wanders into their area of expertise temporarily and they go “OMG, they hype is real. I was so wrong” and then a few releases later they’re on the train and furious that the skills in their domain space have not just stopped improving, but regressed. Cue someone else in a different part of the world starting the same cycle.

Meanwhile the guy who leaned in a year ago and gave up reading the output is beginning to see work grind to a halt and throwing more agents at it is increasingly not working.

You can see these tropes all over social media near constantly.

SpaceNoodled 11 hours ago | parent | prev | next [-]

As a smart, experienced developer who's getting worn down over time, I disagree.

gymbeaux 6 hours ago | parent | prev | next [-]

I didn’t start using Claude Code until late 2025. Prior to that I would use ChatGPT to give me snippets of code but I was still doing most of the actual code writing. Coworkers told me in late 2025 about how they hadn’t written a line of code in “months” and just use Claude Code/agentic “whatever” so I tried out Claude Code and was pleasantly surprised. It is passable to have entire apps written by LLMs (I’ve made several that I otherwise never would have had the time to create by hand), but I wouldn’t say maintainable or easily extendable. It’s hard to be specific, but there’s something about LLM code that doesn’t look “natural”, and I’m not talking about the excessive use of comments in code. The code itself is unnatural. Functional, but unnatural. I wouldn’t want to suddenly lose LLMs and have to read through and understand and continue enhancing a codebase created by an LLM.

fcatalan 4 hours ago | parent [-]

For me it feels a lot like generated images or video. I've made lots of things now, but those that are 100% LLM written "work" but are uncanny, weird and the details are wrong everywhere you care to look in detail.

trashface 9 hours ago | parent | prev [-]

I was getting useful coding work done with GPT 3.5. I think devs saying "the models are finally good enough" this year are just trying to save face from their own previous irrational denials.

ben_w 5 hours ago | parent [-]

Useful, yes, sometimes, but it wasn't fully automated "Here's our JIRA board URL, fix everything that's rated 1-3 story points and in the current sprint".

Now it is.

simonw 11 hours ago | parent | prev | next [-]

Nobody was saying coding agents started working in 2023 or 2024, because the category was defined by Claude Code which was first released in February 2025.

dgellow 26 minutes ago | parent [-]

I would say that Aider is what defined coding agents. That was at least multiple months before Claude code. I remember seeing a coworker use aider for a hackathon project Adeline November-December 2024 , and it was already decent and pretty close to the DX we consider coding agents to have

GolfPopper 4 hours ago | parent | prev | next [-]

Perhaps the LLM companies need to start hiring true Scotsmen?

ChrisMarshallNY 2 hours ago | parent [-]

University of Edinburgh is a good school.

the_gastropod 34 minutes ago | parent | prev | next [-]

I’ve been feeling gaslit about this too. Getting major “we’re still early!” crypto bro vibes from this constant goalpost moving.

dgellow 23 minutes ago | parent [-]

The „you will be left behind if you don’t fully embrace the whole thing right now“ is a 1:1 match with cryptocurrency hype

smrtinsert 6 hours ago | parent | prev | next [-]

Claude 4.5 was it (nov 2025?), without a doubt. It went from frequent hallucinations to highly usable with much less garbage output. If you were making demos of AI tools around this time your demo/pitch/product was saved and you probably looked like a genius.

bigstrat2003 11 hours ago | parent | prev | next [-]

Yep, the goalposts just keep shifting. In reality: they still don't work well, unless you're content with producing low quality work.

bathtub365 9 hours ago | parent | next [-]

This is the opposite of my experience since about February of this year.

gymbeaux 6 hours ago | parent [-]

The quality of the output is so variable. It depends on the model, “effort level”, prompting, probably even the programming language/app functionality, and libraries involved. For example, I find LLMs are best at making simple web apps. These web apps, while simple, would still take a senior engineer perhaps a week or two to create, but LLMs can spit them out inside of an hour. Conversely, LLMs struggle with things like Docker or local model stuff. Parallelization of code is a mixed bag. In these areas I think it often would have been faster for me to write the thing by hand.

throwaway7783 11 hours ago | parent | prev | next [-]

"unless you're content with producing low quality work." - With the right guiding hand, it is a productivity multiplier without compromising quality. As a fully autonomous developer, it is a disaster.

close04 an hour ago | parent | next [-]

How are junior devs becoming qualified “guiding hands” these days? If the expert with LLM assistance is multiplied, what’s a company’s incentive to pay for a junior, and how would they train to get good in these conditions?

Forgeties79 11 hours ago | parent | prev | next [-]

> With the right guiding hand, it is a productivity multiplier without compromising quality

This just reads like another variation of “it’s the user not the tool,” which is just endless runway for always blaming people and never acknowledging the limitations of LLM’s.

I’d be curious to hear how the recipients of your work enabled by the “productivity multiplier” feel about the quality.

inglor_cz 44 minutes ago | parent | next [-]

You can't play an entire orchestra's sheet music on a single guitar either, but your playing ability still matters a lot.

I would say that as of July 2026, with the right scaffolding, you can get reasonably good output out of a LLM, or better a combination of LLMs. For example, it pays off to prepare an implementation plan with one LLM and then let another LLM check it for flaws, then again. After several iterations like this, you will have a plan better than whatever you could come up with yourself.

It often is the user and not the tool. LLMs are complicated, have nontrivial failure modes, and the user needs to steer them carefully. They might be the most complicated tools on the planet right now.

Anecdotally, the recipients of my work have become visibly more happy in the last months. LLMs are great at diagnosing subtle problems which tend to appear at Friday night only, and this is the sort of problem that bugs actual people the most.

11 hours ago | parent | prev [-]
[deleted]
ffsm8 11 hours ago | parent | prev [-]

[flagged]

simonw 11 hours ago | parent [-]

I produce code that is significantly higher quality with the assistance of coding agents, because I no longer succumb to the temptation to cut corners due to lack of time.

One example: everything I do is properly tested and documented now, even the most trivial of changes. Previously I would have weighed those tradeoffs and sometimes decided not to bother with the tests because they weren't worth the time.

janussunaj 3 hours ago | parent [-]

> because I no longer succumb to the temptation to cut corners due to lack of time

No offense, but that says much more about the way you approach programming than about the quality of LLM outputs.

In my experience, LLMs are the ultimate corner-cutting tool. With LLMs, I now succumb to the temptation to cut corners, build something I haven't properly researched and don't fully understand, prioritize shipping quantity over quality.

Without LLMs, I have to understand the domain and the tools and ultimately my full solution (with all its warts and limitations). When I really care about the project and consider it "my baby", LLMs are out of the picture.

oceanplexian 11 hours ago | parent | prev [-]

What kind of work are you doing and what do you consider to be quality or not?

Of course don’t let me assume, maybe you have a higher quality disproof for the Jacobian conjecture you could share with the class.

IshKebab 3 hours ago | parent | prev [-]

I haven't. Around the start of 2026 is pretty widely mentioned as when they went from "this is broken slop" to "huh this is actually 90% what I would have written", which matches my experience.