| ▲ | efficax 6 hours ago |
| Reading the code does not mean you understand the code. One lesson that experience in software gave me: I never understood the code. You think it works a certain way, until you find out that it doesn't. What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand. If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way. Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code. |
|
| ▲ | layer8 5 hours ago | parent | next [-] |
| > I never understood the code. You think it works a certain way, until you find out that it doesn't.
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand. Testing isn’t the same as understanding the code, or proving (even informally) that it is correct. Having the LLM do all these things above doesn’t lead you or the LLM to understand the code, to logically reason about its behavior over all possible states and inputs. “Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know. |
| |
| ▲ | emtel 5 hours ago | parent | next [-] | | You’re technically correct, but the vast majority of software has never been built to the kinds of standards you are describing. LLMs are not displacing that kind of work! | | |
| ▲ | boomlinde 4 hours ago | parent | next [-] | | The person they responded to refers specifically to "healthcare, finance, automotive, defense, power plans, aviation, manufacturing", areas where I'd at least hope that we aspire to understand what the code does. | | |
| ▲ | leonidasrup 20 minutes ago | parent | next [-] | | Also software like SQLite: "
SQLite is built using a DO-178B-inspired process. The testing standards for SQLite are among the highest for commercial software. SQLite is open-source but it is not open-contribution. All the code in SQLite is written by a small team of experts. The project does not accept "pull requests" or patches from anonymous passers-by on the internet.
" https://sqlite.org/hirely.html https://sqlite.org/testing.html | |
| ▲ | kabes 3 hours ago | parent | prev | next [-] | | Having worked in software for healthcare, defense and finance I guarantee you we don't | | | |
| ▲ | Maxatar 3 hours ago | parent | prev | next [-] | | I worked briefly in heath care and the code is so brittle and so poorly understood that almost everyone is afraid to touch anything and instead it's just layers and layers of stuff trying to patch around existing code. | | |
| ▲ | yetihehe an hour ago | parent | next [-] | | I worked briefly in aviation and the part I've seen was very understandable and easy to extend and modify in understandable way. Some parts were hard, but by necessity. We also had very good tests. But maybe that one software was just a good exception. | |
| ▲ | mwwaters an hour ago | parent | prev [-] | | I know this is true. But I don’t think “Some parts of important codebases are black boxes. Therefore it’s fine if all of that code base becomes a far bigger black box” sounds like a good argument. Also, there was probably some human at some point that had some understanding of what they were trying to do and why. The black boxes generally get programmed around after they long left but at the time they had bugs ironed out over decades. (Yes I know sometimes true slop is done over a short period of time and the programmer leaves. But I’ve generally seen the black box built over decades instead). |
| |
| ▲ | ponector 32 minutes ago | parent | prev | next [-] | | I worked for a reinsurance company and their main pricing tool is a huge brittle excel file full of spaghetti VBS code. And yet, they manage to underwrite billions. | |
| ▲ | 3 hours ago | parent | prev [-] | | [deleted] |
| |
| ▲ | ffsm8 4 hours ago | parent | prev | next [-] | | yeah, the llm approach is incredibly wasteful wrt pretty much everything.
Performance, RAM, Development (Tokens). But it does give you surprisingly stasble rube-goldberg machines. And thats basically what 95-99% of enterprises want from their software. It annoyed me to no end when i began my career, but at this point ive accepted it and can definitely still have fun developing software with llms. As a matter of fact, as my perfectionism approach to software in my earlier years was never really appreciated... So i dont really mind the new MO. I still occasionally hand write though, esp. at the dayjob where ive got super small token budgets while continuously being told to use more AI. But that's normal, employers usually give off bipolar vibes with multiple stakeholders wanting to advance each of their bonus package KPI of any given quarter | | |
| ▲ | jonahx 3 hours ago | parent [-] | | > yeah, the llm approach is incredibly wasteful wrt pretty much everything. Everything except what matters most: human time. | | |
| |
| ▲ | skydhash 2 hours ago | parent | prev [-] | | The vast majority of software is not that important. I don’t really care about easytag (which I use for flac metadata), but I do care about xterm and tmux. |
| |
| ▲ | coldtea 4 hours ago | parent | prev | next [-] | | >“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know. We're not writing theorems, dude. Except in the equally pedantic sense that every program is a proof to a theorem... We're writing plain enterprise and web software, closer to CRUD than NASA. If you said that even before LLMs 0.1% of teams "checked all assumptions against what the code and underlying systems are actually guaranteeing" in any kind of formal way, you'd be overestimating it. | | |
| ▲ | layer8 4 hours ago | parent | next [-] | | I’m not talking about formal verification, but about diligent informal or semi-formal reasoning through the code, so that you can rightfully claim that you understand the code and will be unlikely to be surprised by its behavior. Having learned formal verification does train that form of exhaustive reasoning about properties of the program. This practice also has you structure the code such (and select your dependencies such) that you can reason about all relevant properties. This is perfectly applicable to what you’d call CRUD and enterprise applications (that’s half the projects I earn my living with). Testing and fuzzing are complementary, but not a substitute by any stretch. | | |
| ▲ | jonahx 4 hours ago | parent [-] | | GP is correct. Very few people were capable of even the informal analysis you are describing, and fewer did it. I'm not saying it's not valuable... just stating that, empirically, it rarely happened. | | |
| ▲ | IAmBroom 3 hours ago | parent [-] | | And layer8 is saying (two responses upwards by them) that this is a novel benefit of AI: it can do a particularly thorough and repetitive kind of fault analysis that is a real PITA for humans to do (by their nature, versus the nature of computers). |
|
| |
| ▲ | Silamoth 4 hours ago | parent | prev [-] | | Who’s “we” here? Formal verification isn’t common, sure. But you don’t speak for all programmers. You might work on “plain enterprise and web software”. But there’s still plenty of other software out there that many of us work on. And lots of code being written for internal use (e.g., data analysis code) that needs to be correct. Of course, even enterprise and web software benefits from a little rigorous thinking. It’s pretty wild that understanding your code and its assumptions and informally proving it works is controversial. But I guess that explains why most software I use has actively gotten worse over the years. | | |
| |
| ▲ | ModernMech 4 hours ago | parent | prev [-] | | Your code is only as good as what you can prove. Understanding the code is not the goal, it’s only important insofar as it helps you evolve the codebase predictably and without bugs or regressions, and understanding is not easily measurable or transferable. Moreover, when your codebase is hundreds of thousands to millions LOC, I question how much you can ever truly understand it at the level you’re saying. | | |
| ▲ | layer8 4 hours ago | parent [-] | | Regarding the last part, the strategy is to not have everything depend on everything, to instead modularize with succinct interfaces, so that you can reason locally. Of course beyond a certain project size, there is no single person who understands every part in detail. But for every part you can have someone who understands it, and can reason about it in terms of the interface contracts with the other parts. It’s also not essential that every detail is still understood at every point in time, as long as it’s sufficiently documented. What is essential is that for every part someone did reason through it with the necessary rigor at some point. | | |
| ▲ | ModernMech 3 hours ago | parent [-] | | > to instead modularize with succinct interfaces, so that you can reason locally Okay but how does AI change any of that? You can still do that with AI. > as long as it’s sufficiently documented. AI definitely helps with that. > What is essential is that for every part someone did reason through it with the necessary rigor at some point. Why is that essential though? What if the person who reasoned about it dies or leaves? Moreover, why is it imperative the reasoning happens at the source code level? | | |
| ▲ | discreteevent 3 hours ago | parent [-] | | >> to instead modularize with succinct interfaces, so that you can reason locally > Okay but how does AI change any of that? You can still do that with AI With your own code you reasoned about it which contributed to its stability. This meant that you could treat it like a black box. And if the abstraction leaked or was unstable, the code was still fresh enough in your head that you could evolve it and still preserve its invariants etc. With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI. | | |
| ▲ | ModernMech 2 hours ago | parent [-] | | > With your own code you reasoned about it which contributed to its stability. Okay but to what extent? People say this but there's no way to measure it really. Did you live through the 90s? People reasoned through all that code and it was very often quite unstable. I'm sure everyone involved with Windows ME reasoned about it quite a lot, probably elements of it locally were very sound, yet in totality it was an unstable mess. What fixed that situation wasn't that engineers today are reasoning better than engineers in the 90s, but IMO better tooling. Which brings me back to: your codebase is only as good as what it can prove. If there's any question, I just show you the proof rather than appealing to my reasoning being sound. > With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI. And? You haven't established reasoning about it is actually necessary and it certainly isn't sufficient. |
|
|
|
|
|
|
| ▲ | jdkoeck 5 hours ago | parent | prev | next [-] |
| > Reading the code does not mean you understand the code. Reading the code may not be enough to understand the behaviour of your program, but believing you can understand the behaviour of a program without at least reading the high level code is truly silly. (by high level, I mean the code living in the higher layers - of course we don't often read the code of the generated assembly, or the interpreter, or the browser, but that's because they're reliable abstractions, unlike prompts!) |
| |
| ▲ | rco8786 5 hours ago | parent [-] | | > believing you can understand the behaviour of a program without at least reading the high level code is truly silly. have you ever used a library after only reading the README and documentation, or do you always pull the source and read through it before you think you understand it? | | |
| ▲ | misterderpie 4 hours ago | parent | next [-] | | I would draw the difference here that a library is used by hundreds (of thousands) of people, and established across different scenarios. I don't read the boto3 library AWS provides, but I can trust them and the amount of customers enough to be certain enough that it behaves the way I expect it to.
The same can't be said with code we write in silos at our workplace or at home. It simply does not have the same test bench. Yes, libraries aren't bug-free, but they give me a reliable abstraction tested in the field. Not rarely you dig into library code if you notice unexpected behavior. If we could rely on our LLM or colleague written code, or own code, have run through the same amount of requests, sure I wouldn't need to review it, as my confidence can be north of 99.9999% it works correctly. But we can't. | |
| ▲ | ubertaco an hour ago | parent | prev | next [-] | | I've used libraries before, where I call the API surface that they expose based on method names and parameter types, without reading all the source. Generally those libraries don't implement my software's entire problem domain area; they tend to implement things like "CSV parser" or "HTTP server". The important code, that I'm actively reading and writing, tends to be the code around those library calls. This is different from building a product, which you only interact with via UI buttons/CLI/etc, without reading any code to understand how it conceptualizes that product's problem domain area. People do that latter thing, and we call them "users", not "developers". | |
| ▲ | 4 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | watermelon0 4 hours ago | parent | prev | next [-] | | You can use it without understanding, but this won't help you actually understand how it behaves. | | |
| ▲ | rco8786 4 hours ago | parent [-] | | I believe I can understand how a library behaves from the README and docs. I do this all the time. We all do. | | |
| ▲ | kevinh 4 hours ago | parent [-] | | You haven't run into cases where the documentation is missing or incomplete? You must be dealing with different libraries than I am. | | |
| ▲ | rco8786 4 hours ago | parent [-] | | Of course. And then we fix the problem. It's no different here. The point is that reading the code first is absolutely not a requirement when adopting a new library. | | |
| ▲ | skydhash 2 hours ago | parent [-] | | It is not. But there’s an element of trust being involved. Something like libflac or libcurl, I don’t read the code. I read the doc which does outline the behavior of each function and the conceptual model. If something break, it’s quite often my code. Why? because their code is battle tested. Which is quite different from AI generated code. | | |
| ▲ | rco8786 26 minutes ago | parent [-] | | That's because it's new. libflac and libcurl didn't come out as battle tested code. That takes time and...battles. And what are those battles if not people seeing issues with what those libraries are doing and fixing the code? There absolutely no reason that AI written code can't or won't become battle tested. | | |
| ▲ | skydhash 5 minutes ago | parent [-] | | The battle tested part is the reason I suspect my code rather than theirs first. The trust part is in their engineering process. New project from OpenBSD or Debian, I’m ready to try. Random guy on Github, bring out the 10 foot pole. |
|
|
|
|
|
| |
| ▲ | brabel 4 hours ago | parent | prev | next [-] | | I am sure OP cannot understand the Unix file API without actually reading every line of its implementations (on each different architecture)! Or any function for that matter , what does sort do?? Impossible to know without reading the source. And I’m sure after reading the source you will know every detail of how it works and will never forget it. | | |
| ▲ | discreteevent 3 hours ago | parent [-] | | It's hard for me to understand how you could work on software and not understand the qualitative difference between the Unix File API and some code that an AI spat out 5 minutes ago. |
| |
| ▲ | weakfish 4 hours ago | parent | prev [-] | | No, but the authors of $LIBRARY are accountable if it fails | | |
| ▲ | rco8786 4 hours ago | parent [-] | | Are they? I've never been able to blame library authors or hold them accountable for code running in my production environment. |
|
|
|
|
| ▲ | 12ag5a 6 hours ago | parent | prev | next [-] |
| Strange that the world worked before 2024 and software gets worse now. Your debit card transactions for example worked. This sounds like a typical testimonial whose mind has become captive to Claude. It is like Scientology. |
| |
| ▲ | bananaflag 6 hours ago | parent | next [-] | | Before 2024, I once went to an ATM to retrieve money and selected 50. Note that I selected it from a menu, not typed it. The ATM then told me that it cannot give me 50 because it is not a multiple of 5. | | |
| ▲ | anon7725 4 hours ago | parent [-] | | meanwhile for the other lim(n -> inf) times people tried this it mostly worked. |
| |
| ▲ | MattDamonSpace 6 hours ago | parent | prev | next [-] | | Yeah no one wrote buggy code before 2024 right | | |
| ▲ | lazystone 5 hours ago | parent | next [-] | | And after 2024 all bugs cease to exist, right. | | |
| ▲ | ofjcihen 5 hours ago | parent [-] | | No, but we’re spending insane amounts of money to essentially end up where we started. How can you not see the progress?! | | |
| ▲ | bluGill 5 hours ago | parent [-] | | Huh? We are spending a lot of money. (which we were doing before). However we are fixing a lot of bugs. 2024 was not that long ago, it is insane to think we might have fixed all the bugs in that time. I have personally used an LLM to fix a few long standing rare bugs that were hard to figure out. Those bugs are now gone, but there are still many more that we haven't discovered. LLMs are a great thing for bug fixing. However they are not a miracle. You still need to do all the other things about finding, testing and fixing bugs. You also need to care about bugs - vibe coding rarely cares about bugs. | | |
| ▲ | ofjcihen 5 hours ago | parent [-] | | Well sure, on the customer side. My money comment was regarding the incestuous spending circle going on that’s directly affecting the economy. |
|
|
| |
| ▲ | 4 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | quatotor 2 hours ago | parent | prev [-] | | [dead] |
| |
| ▲ | paimapi 5 hours ago | parent | prev | next [-] | | I mean, I think the problem isn't that the LLM doesn't know how to code, it's that companies are expecting 3-5x velocity with the bottleneck of code review and testing becoming much more severe than before if you're an MBA-brained exec who doesn't actively use LLMs to code and you just believe whatever slop it outputs at first without checking it, you're not going to realize how recklessly it can be used, how you need to be critical and skeptical of its outputs, that you need to explore it's reasoning and logic (which is still really easy compared to understanding legacy code and barely takes any time!) say you also believe all this marketing hype about 'how dangerous (ie capable) AI agents are.' LLMs can do anything you think so you just say 'ship it' without building out the tooling and capabilities to enable faster code review and better tests. and to keep the shareholders happy, you start cutting jobs that you can't directly connect to a KPI (ie the platform/SRE team who would be the ones who can trial, onboard, and maintain those capabilities for your teams) and from this, suddenly a lot of debit card stops working and the only one getting the blame are individual SWEs trying to hit their sprint velocity. the fact that you fucked up the whole SDLC real bad with your incompetence gets you a golden parachute and you job hop to a better paycheck. rinse and repeat | |
| ▲ | pprotas 5 hours ago | parent | prev | next [-] | | Debit card transactions still work | |
| ▲ | echelon 5 hours ago | parent | prev | next [-] | | > Your debit card transactions for example worked. I've built payment rails. Six nines SLA, high capacity, resilient distributed systems. I haven't written a single line of code since February, and I don't think I ever will again. These systems are incredibly good at replacing much of our work. They're only going to get better. Rather than debating if these models are good (they are), we should be trying to figure out if most of us will still be around in three years. You don't need a two pizza team anymore. "Look to the person to your left and to your right. Only one of you will remain by graduation" kind of energy. I'm not sure all of us is going to be in this career much longer. We'll have to see what the demand side looks like. | | |
| ▲ | bluGill 5 hours ago | parent | next [-] | | Most of developing good code is not code. I think the person to my left and right will both be here in 3 years despite us all using LLMs. We will spend even more time figuring out requirements, testing to ensure the code meet them and such. Those things were always most of the effort, and while LLMs help with that too there is so much work to be done that we will still be used. | |
| ▲ | whateveracct 2 hours ago | parent | prev | next [-] | | > You don't need a two pizza team anymore. on-call still exists. have fun round robin'ing that with 3 engineers. | | |
| ▲ | fragmede 9 minutes ago | parent [-] | | Does it? First line of defense is now an LLM with the runbook in its context and the human in the loop being woken up at 3 am, half asleep, just needs to make sure it doesn't rm -rf something. |
| |
| ▲ | wonnage 26 minutes ago | parent | prev | next [-] | | nothing completes a HN AI thread more than echelon jumping in from the top rail with a wild take | |
| ▲ | ModernMech 3 hours ago | parent | prev | next [-] | | On the one hand, you're exactly right. On the other, my local pool company is hiring a software engineer and hardware engineer because with AI, they can replace a 2 pizza team as you so succinctly put it. So no two pizza teams but that doesn't mean all the pizzas are gone, they're maybe going to be spread out and not concentrated in CA, between orgs you might not have thought as "tech" before. | |
| ▲ | nullsanity 4 hours ago | parent | prev | next [-] | | [dead] | |
| ▲ | 319286 5 hours ago | parent | prev [-] | | You can still find employment as an AI shill. |
| |
| ▲ | lbrito 3 hours ago | parent | prev | next [-] | | Yes, but Have You Tried the Latest Model? /s | |
| ▲ | rowanG077 5 hours ago | parent | prev | next [-] | | Why or how is software worse now? From my perspective we are entering a golden age of software, cheaper, better, faster. | | |
| ▲ | eureka7 2 hours ago | parent | next [-] | | This is just a personal anecdote, but for the past year or so, the software I use on the daily has never been more buggy. | |
| ▲ | cat-snatcher 5 hours ago | parent | prev | next [-] | | Because every time I see AI integration anywhere (and it's everywhere) I have an emotional breakdown and my day is ruined | | |
| ▲ | bluGill 5 hours ago | parent [-] | | AT is abused into many places it shouldn't be. That doesn't change that there are places where it is helpful. You should seek professional help about your emotional issues. | | |
| ▲ | cat-snatcher 5 hours ago | parent [-] | | I was being sarcastic. I’m actually a self-certified AI shill and personally paid by Dario and Sam for each and every post! |
|
| |
| ▲ | 5 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | anotha_one 5 hours ago | parent | prev [-] | | [dead] |
| |
| ▲ | InsideOutSanta 5 hours ago | parent | prev [-] | | > Strange that the world worked before 2024 It must have been a huge shock when you were suddenly transported from a working parallel universe into ours back in 2024. |
|
|
| ▲ | pu_pe 5 hours ago | parent | prev | next [-] |
| I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. All these mission critical industries listed in the article rely on extensive testing for quality assurance, with human code review being a layer on top of all that, but far from the most critical one. Interpretability is the same, our abilities to do that have increased rather than decreased. I think a codebase generated by AI is actually more understandable than one generated by humans at this point, and you can ask clarifying questions whenever you get stuck. TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it. |
| |
| ▲ | eithed 5 hours ago | parent | next [-] | | If you can quantify quality/reliability/understandability, you can tell LLM what kind of code do you expect. If not, you get whatever. At my current place we not only have automated tests, static analysis and static rector (linting, but also automatic pattern matcher for problematic code) but also:
- architecture tests that define relationships between application layers
- ADRs that guide developers (and agents as well) that communicate how new code should be written and how existing code should be treated I find that "how code should look like"/"what code should do" is an ambiguous idea that always is preached, but never defined = everyone's idea of quality is slightly different and only looking at existing code you tend to align. Everyone's idea of what the product does/should do is kept within their heads. If we define this knowledge in writing LLMs can not only write code according to the patterns that are thus defined, review existing code based on these documents, but also actually read acceptance criteria documents to check if the code does what it's intended to do (gherkin) Same goes for understandability - if LLM applies one pattern this time, another pattern another time, if you have multiple coding patterns then that hurts clarity. Sometimes LLMs work as common denominator thus achieving clarity, but I find that actually giving LLMs reference works. | |
| ▲ | munksbeer 4 hours ago | parent | prev | next [-] | | > I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop You're not going to get people to stop doing that by arguing on the internet, but in the end it won't matter, because it will stop, naturally. In the future, you'll just get left behind and not hired if you're building code by hand, it's that simple. Even traditional code reviews are going to go away. It'll be more about the scope and then verifying correctness. | | |
| ▲ | brazukadev 3 hours ago | parent [-] | | > In the future, you'll just get left behind and not hired if you're building code by hand, it's that simple. I expect the exactly opposite to happen. These are going to be the most requested developers as the last ones that understand how it work. They would then be convinced to use AI for speed, but vibecoders that just prompt AI are the ones that won't find jobs. | | |
| ▲ | sally_glance 3 hours ago | parent | next [-] | | Agreed for pure vibecoders, but I would expect vibecoding to just become a must have skill for other roles like product owners etc. Still expect some amount of developers to be retained for grooming the vibecoding environment, reviewing and incident response. | |
| ▲ | munksbeer 3 hours ago | parent | prev [-] | | I'm not sure I understand what you're saying? To keep it simple, how much code do you expect will be written by humans in say, two years time? |
|
| |
| ▲ | palmotea 5 hours ago | parent | prev | next [-] | | > I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. All these mission critical industries listed in the article rely on extensive testing for quality assurance, with human code review being a layer on top of all that, but far from the most critical one. At least some places are abolishing formal QA because LLMs. There's a cult of speed uber alles that has a big intersection with LLM enthusiasm. | | | |
| ▲ | suddenlybananas 5 hours ago | parent | prev | next [-] | | >TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it. If coding were solved, then this would be true no? | | |
| ▲ | pu_pe 5 hours ago | parent [-] | | There are far more moving pieces in deploying an application than coding. My impression is TFA is arguing that replacing humans with LLMs for coding would make other things like quality assurance more difficult, in which case I disagree. |
| |
| ▲ | westurner 5 hours ago | parent | prev | next [-] | | > I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop I agree. What does coverage-guided fuzzing fuzz if there is 100% test coverage? So, then, 100% branch test coverage is not a sufficient metric (because it doesn't indicate whether the code is fuzzed or formally verified for example). Would Branch coverage even be a sufficient software quality metric if we were to instead measure how many times each branch of code is covered by tests? How to verify that one test which executes 100% of the code and runs only one assertion on, say, a CLI utility exit code integer is actually sufficiently covering? > I think a codebase generated by AI is actually more understandable than one generated by humans at this point, From doing a larger port (of sphinx, docutils, myst-md-parser, pygments, to rust in westurner/dsport) with a lot of human in the loop and currently ~80% branch coverage,
this seems to be at least initially true but just like real life there's drift from even a good plan that you pay a more expensive model to prepare. I suppose it's the same challenge as architectural drift in open source non-LLM-assisted products and the solutions are pretty much the same: give better instructions (AGENTS.md,) and use better sufficiency criteria as an engineering manager (branch test coverage, fuzzing, formal methods, TLA+), and train and pay humans to do secure code review. Sometimes the agent doesn't notice that the code already solves for that and implements its own implementation with tests and it's wastefully redundant when the code should be refactored and the tests should be refactored so that we can delete code in order to minimize bloat. Unfortunately often, just like IRL software development, the response from the agent is not sufficient to close the issue. One proposed solution for this that is in retrospect obvious and also essential to success in "normal"/"traditional"/"legacy" (non-AI) engineering projects, is to always verify whether the candidate solution satisfies the criteria; From "Groundtruth – checks your AI coding agent's claims against the Git diff" https://news.ycombinator.com/item?id=48838209 : > "Follow up to verify that the work was actually satisfactorily completed" > Are there other sound management practices that aren't yet effectively implemented in current gen agents? Oh, and always write tests, docs, commit messages, and changelog entries; but don't waste tokens on documenting something that doesn't verifiably pass sufficient tests. | |
| ▲ | bdcravens 4 hours ago | parent | prev [-] | | > I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. It's not just a false dichotomy, it's intellectual dishonesty. It wasn't that long that conversations about code quality, technical debt, etc were on the front page of HN on the regular. Whether it was coding bootcamp grads who had just enough confidence to be dangerous, "just ship it!" cargo culters, or the product of management breathing down the necks of otherwise good developers, there's plenty of "human slop" running in production across servers worldwide. | | |
| ▲ | pona-a 4 hours ago | parent [-] | | But we shouldn't use one evil to justify another, or your argument regresses to whataboutism. |
|
|
|
| ▲ | flatline 6 hours ago | parent | prev | next [-] |
| I don't think the discrepancy is in LLM capability improvements over the past year. Correctness has never been a priority across an industry where rapid iteration and feature delivery drive sales. There's always some opportunity cost to doing things right, at the price of technical debt down the road. If AI is primarily used to produce fragile code, people will be wary of AI solutions. There's also ongoing public debate about AI safety and alignment. Deploying AI in safety critical applications feels riskier than ever in the current environment, even though it doesn't have to be. |
| |
| ▲ | ThrowawayR2 2 hours ago | parent [-] | | It would be deliciously ironic if AI was the straw that broke the camel's back where a deluge of bugs and anti-AI sentiment caused the public to vote for legal liability for software defects and licensure of developers. No more of this "no warranty, express or implied" business for us. |
|
|
| ▲ | clickety_clack 5 hours ago | parent | prev | next [-] |
| For me, coding is like writing. The act of doing it is how you reason out the problem. There’s a lot of magical thinking you can get away with in your head that doesn’t get properly tested until you write it down. For me, vibe coding is great and fast, but I’m not getting the same opportunity to think through the problem I’m trying to solve. |
|
| ▲ | jayd16 5 hours ago | parent | prev | next [-] |
| > The LLMS are very good at logic, by the way. It's wild to read this stuff and then also deal with the constant headaches of day to day hallucinations when interacting with Claude et al. |
| |
| ▲ | chimprich 4 hours ago | parent [-] | | I'm rather surprised to hear this. This feels like a post from about 18 months ago. I can't remember the last time I encountered a genuine code hallucination from a frontier model. They have other issues, but rarely this. What kind of domain are you working in? | | |
| ▲ | smrtinsert 9 minutes ago | parent | next [-] | | Fintech, easy to trigger some sort of failure mode or obvious gap once specs get detailed enough and the prompt intents become specific enough. Doesn't require exhausting available context. Opus 5 did feel like a regression, will hold out judgement on 5.5 which seems much more promising. | |
| ▲ | weakfish 4 hours ago | parent | prev | next [-] | | I see subtle ones at least daily, misunderstanding a component or hallucination of a spec for something. I’m in blockchain. | | |
| ▲ | daveguy 40 minutes ago | parent [-] | | But you should see how great it is at crud social media apps and 2D scrolling games! |
| |
| ▲ | jayd16 4 hours ago | parent | prev [-] | | C++ and Unreal Engine but it has full source access. If you're actually trying to deep dive on bugs, it's still confidently wrong a lot of the time. It's a bit better than 18 months ago but it's hard to say by how much. It just seems like the culture has moved to building up fixtures that let the LLMs brute force the problems. To my eyes that's the opposite of solving things logically. It has the added effect of hiding how the sausage is made, though. I mean, how can they possibly say they haven't written a line of code if they're actually going through it? I can only assume they're just looking at the results. So then how can they judge it's good at logic? If it was so good at not making mistakes, why even have tests? It's nonsensical on its face. |
|
|
|
| ▲ | asveikau 5 hours ago | parent | prev | next [-] |
| > build property tests Even for a narrow use like this, you need to audit the output and have the skills to know that it did the right thing. I've seen it before where you give an LLM what seems like a clear interface and ask it write a test and it writes something shallow that doesn't actually test anything, or has serious problems. |
|
| ▲ | allknowingfrog an hour ago | parent | prev | next [-] |
| You somehow squeezed "you're holding it wrong" and "LLMs are actually good now" into the same sentence. I wasn't sure it could be done. |
|
| ▲ | lolakutty 4 hours ago | parent | prev | next [-] |
| > find out all the ways this thing works. Mmm..aren't LLMs bad at exhaustively iterating all possibilities? So shouldn't the generated possibilities be manually checked? You can provide the list of possibilities and use LLMs to generate the tests. Then you have to review the generated tests... |
|
| ▲ | tshaddox 5 hours ago | parent | prev | next [-] |
| > I never understood the code. You think it works a certain way, until you find out that it doesn't. Those are two separate claims, unless by the former you mean “I never perfectly understood the code.” You can understand code imperfectly. And even with LLMs, you can’t get truly infallible guarantees about a system. |
|
| ▲ | geraneum 4 hours ago | parent | prev | next [-] |
| > find out all the ways this thing works. If you can’t understand the code, how do you know the LLM actually did what you asked it to do correctly? You wouldn’t know if it didn’t. |
|
| ▲ | captainbland 4 hours ago | parent | prev | next [-] |
| I think you can but you basically need to learn about a super simple and well characterised processor like the 8080 and write assembly for it. On x86/AMD64 there's no hope because they're out of order and have opaque instruction decoding. They could be doing anything! Performance and knowing what you're doing are sort of at odds with each other in that respect. |
| |
| ▲ | ThrowawayR2 2 hours ago | parent [-] | | > "out of order" Those who bring up OoO as if it caused unpredictable execution should ask their LLM what a reorder buffer is and why it exists. |
|
|
| ▲ | geauxvirtual 5 hours ago | parent | prev | next [-] |
| > I never understood the code. > I understand code Are you sure? |
| |
| ▲ | usewik 4 hours ago | parent [-] | | Article 'the' is doing a lot of work. I read it as author understands code but doesn't understand THE code. |
|
|
| ▲ | boomlinde 3 hours ago | parent | prev | next [-] |
| > What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand. How do you know that it's verifying that the system under test exhibits the properties you desire without either understanding or making blind assumptions about the code it generates to build a fuzzer or a property test? It seems to me that you have just shifted the problem of verification elsewhere and introduced another potential source of error. |
| |
| ▲ | mckn1ght 3 hours ago | parent | next [-] | | I came here to say something similar. We won’t need to understand the implementation. But we will need to understand the requirements. The tests, or some higher level DSL they’re (deterministically! not via LLMs) generated from, will still need to be human verified. | |
| ▲ | IAmBroom 3 hours ago | parent | prev [-] | | They aren't making that claim. They are saying AI can improve on and supplement error-checking. And some error-checking will in fact be made redundant by this tool - but obviously not all of it. | | |
| ▲ | boomlinde an hour ago | parent [-] | | > They aren't making that claim. "[...] without ever reading a single line of code" begs to differ. |
|
|
|
| ▲ | accidc 4 hours ago | parent | prev | next [-] |
| We could always build, more reliable software. The powers that be however (broadly generalizing) only care about solving the ‘problem’ superficially. Therefore, we will end up with requests to do more (at the current level of quality/ reliability), as opposed to building better software |
|
| ▲ | grumple 6 hours ago | parent | prev | next [-] |
| I’m using it a bunch. It saves a ton of time writing or reviewing code. It will catch things I won’t. But I’d express caution about the analysis or evaluation they do - LLMs will often confidently proclaim problems as solved or explain functionality and be wrong about it. Sometimes subtly, but sometimes just completely wrong. This is no different from humans, of course, except for the unabated confidence. |
|
| ▲ | jibalt 5 hours ago | parent | prev | next [-] |
| > I never understood the code. > I understand code Er, ok. |
|
| ▲ | bdcravens 5 hours ago | parent | prev | next [-] |
| > Reading the code does not mean you understand the code. Nor does writing it. |
|
| ▲ | slifin 5 hours ago | parent | prev | next [-] |
| I hate that in our industry we refuse to operationalise reverse debuggers "I never understood the code. You think it works a certain way, until you find out that it doesn't." Bret Victor made a talk called "seeing spaces" in 2014 that should have woken up this whole industry:
https://www.youtube.com/watch?v=klTjiXjqHrQ He emphasizes that without the ability to see inside what is being built, creators often fall into "non-scientific thinking" (14:42), moving away from deep understanding and instead "blindly following recipes, from superstitions and rules of thumb" (14:47-14:51). The worse is performance problems I've had engineers say some bizzaro things when discussing performance — we have the tools you can just measure the answer - we don't need to waste our time guessing |
| |
| ▲ | daveguy 23 minutes ago | parent | next [-] | | Thank you!! I saw this video about a decade ago, couldn't remember where it came from, and was having trouble finding it again! Bret Victor is an HCI juggernaut. From this info, here are some other links: his website: https://worrydream.com/ his page for learnable programming: https://worrydream.com/LearnableProgramming/ There was a programming environment closely related to his work that was a kind of visual database. But it may not have been developed by him / his lab. I don't see any obvious links to it on his website. | |
| ▲ | xpct 5 hours ago | parent | prev [-] | | Thank you for sharing :) good food for thought |
|
|
| ▲ | troupo 2 hours ago | parent | prev | next [-] |
| > I've written 100s of thousands of lines of difficult code. How do you know it's difficult if you say you don't understand it? > Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I run all the "latest and greatest" models the moment they become available to me. The amount of insanely bad code they produce remains largely the same, and largely in the same areas. And it cannot be caught by tests unless you know that bad code is there and end up with extremely bad tests anyway. I wrote about it here: https://dmitriid.com/adding-to-i-dont-read-ai-code-discourse Main one is, of course, "to get a record from a database read all records from it, and filter in memory". |
|
| ▲ | lirolero 6 hours ago | parent | prev | next [-] |
| [dead] |
|
| ▲ | nullsanity 4 hours ago | parent | prev [-] |
| [dead] |