| ▲ | layer8 5 hours ago |
| > I never understood the code. You think it works a certain way, until you find out that it doesn't.
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand. Testing isn’t the same as understanding the code, or proving (even informally) that it is correct. Having the LLM do all these things above doesn’t lead you or the LLM to understand the code, to logically reason about its behavior over all possible states and inputs. “Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know. |
|
| ▲ | emtel 5 hours ago | parent | next [-] |
| You’re technically correct, but the vast majority of software has never been built to the kinds of standards you are describing. LLMs are not displacing that kind of work! |
| |
| ▲ | boomlinde 4 hours ago | parent | next [-] | | The person they responded to refers specifically to "healthcare, finance, automotive, defense, power plans, aviation, manufacturing", areas where I'd at least hope that we aspire to understand what the code does. | | |
| ▲ | leonidasrup 19 minutes ago | parent | next [-] | | Also software like SQLite: "
SQLite is built using a DO-178B-inspired process. The testing standards for SQLite are among the highest for commercial software. SQLite is open-source but it is not open-contribution. All the code in SQLite is written by a small team of experts. The project does not accept "pull requests" or patches from anonymous passers-by on the internet.
" https://sqlite.org/hirely.html https://sqlite.org/testing.html | |
| ▲ | kabes 3 hours ago | parent | prev | next [-] | | Having worked in software for healthcare, defense and finance I guarantee you we don't | | | |
| ▲ | Maxatar 3 hours ago | parent | prev | next [-] | | I worked briefly in heath care and the code is so brittle and so poorly understood that almost everyone is afraid to touch anything and instead it's just layers and layers of stuff trying to patch around existing code. | | |
| ▲ | yetihehe an hour ago | parent | next [-] | | I worked briefly in aviation and the part I've seen was very understandable and easy to extend and modify in understandable way. Some parts were hard, but by necessity. We also had very good tests. But maybe that one software was just a good exception. | |
| ▲ | mwwaters an hour ago | parent | prev [-] | | I know this is true. But I don’t think “Some parts of important codebases are black boxes. Therefore it’s fine if all of that code base becomes a far bigger black box” sounds like a good argument. Also, there was probably some human at some point that had some understanding of what they were trying to do and why. The black boxes generally get programmed around after they long left but at the time they had bugs ironed out over decades. (Yes I know sometimes true slop is done over a short period of time and the programmer leaves. But I’ve generally seen the black box built over decades instead). |
| |
| ▲ | ponector 31 minutes ago | parent | prev | next [-] | | I worked for a reinsurance company and their main pricing tool is a huge brittle excel file full of spaghetti VBS code. And yet, they manage to underwrite billions. | |
| ▲ | 3 hours ago | parent | prev [-] | | [deleted] |
| |
| ▲ | ffsm8 4 hours ago | parent | prev | next [-] | | yeah, the llm approach is incredibly wasteful wrt pretty much everything.
Performance, RAM, Development (Tokens). But it does give you surprisingly stasble rube-goldberg machines. And thats basically what 95-99% of enterprises want from their software. It annoyed me to no end when i began my career, but at this point ive accepted it and can definitely still have fun developing software with llms. As a matter of fact, as my perfectionism approach to software in my earlier years was never really appreciated... So i dont really mind the new MO. I still occasionally hand write though, esp. at the dayjob where ive got super small token budgets while continuously being told to use more AI. But that's normal, employers usually give off bipolar vibes with multiple stakeholders wanting to advance each of their bonus package KPI of any given quarter | | |
| ▲ | jonahx 3 hours ago | parent [-] | | > yeah, the llm approach is incredibly wasteful wrt pretty much everything. Everything except what matters most: human time. | | |
| |
| ▲ | skydhash 2 hours ago | parent | prev [-] | | The vast majority of software is not that important. I don’t really care about easytag (which I use for flac metadata), but I do care about xterm and tmux. |
|
|
| ▲ | coldtea 4 hours ago | parent | prev | next [-] |
| >“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know. We're not writing theorems, dude. Except in the equally pedantic sense that every program is a proof to a theorem... We're writing plain enterprise and web software, closer to CRUD than NASA. If you said that even before LLMs 0.1% of teams "checked all assumptions against what the code and underlying systems are actually guaranteeing" in any kind of formal way, you'd be overestimating it. |
| |
| ▲ | layer8 4 hours ago | parent | next [-] | | I’m not talking about formal verification, but about diligent informal or semi-formal reasoning through the code, so that you can rightfully claim that you understand the code and will be unlikely to be surprised by its behavior. Having learned formal verification does train that form of exhaustive reasoning about properties of the program. This practice also has you structure the code such (and select your dependencies such) that you can reason about all relevant properties. This is perfectly applicable to what you’d call CRUD and enterprise applications (that’s half the projects I earn my living with). Testing and fuzzing are complementary, but not a substitute by any stretch. | | |
| ▲ | jonahx 4 hours ago | parent [-] | | GP is correct. Very few people were capable of even the informal analysis you are describing, and fewer did it. I'm not saying it's not valuable... just stating that, empirically, it rarely happened. | | |
| ▲ | IAmBroom 3 hours ago | parent [-] | | And layer8 is saying (two responses upwards by them) that this is a novel benefit of AI: it can do a particularly thorough and repetitive kind of fault analysis that is a real PITA for humans to do (by their nature, versus the nature of computers). |
|
| |
| ▲ | Silamoth 4 hours ago | parent | prev [-] | | Who’s “we” here? Formal verification isn’t common, sure. But you don’t speak for all programmers. You might work on “plain enterprise and web software”. But there’s still plenty of other software out there that many of us work on. And lots of code being written for internal use (e.g., data analysis code) that needs to be correct. Of course, even enterprise and web software benefits from a little rigorous thinking. It’s pretty wild that understanding your code and its assumptions and informally proving it works is controversial. But I guess that explains why most software I use has actively gotten worse over the years. | | |
|
|
| ▲ | ModernMech 4 hours ago | parent | prev [-] |
| Your code is only as good as what you can prove. Understanding the code is not the goal, it’s only important insofar as it helps you evolve the codebase predictably and without bugs or regressions, and understanding is not easily measurable or transferable. Moreover, when your codebase is hundreds of thousands to millions LOC, I question how much you can ever truly understand it at the level you’re saying. |
| |
| ▲ | layer8 4 hours ago | parent [-] | | Regarding the last part, the strategy is to not have everything depend on everything, to instead modularize with succinct interfaces, so that you can reason locally. Of course beyond a certain project size, there is no single person who understands every part in detail. But for every part you can have someone who understands it, and can reason about it in terms of the interface contracts with the other parts. It’s also not essential that every detail is still understood at every point in time, as long as it’s sufficiently documented. What is essential is that for every part someone did reason through it with the necessary rigor at some point. | | |
| ▲ | ModernMech 3 hours ago | parent [-] | | > to instead modularize with succinct interfaces, so that you can reason locally Okay but how does AI change any of that? You can still do that with AI. > as long as it’s sufficiently documented. AI definitely helps with that. > What is essential is that for every part someone did reason through it with the necessary rigor at some point. Why is that essential though? What if the person who reasoned about it dies or leaves? Moreover, why is it imperative the reasoning happens at the source code level? | | |
| ▲ | discreteevent 3 hours ago | parent [-] | | >> to instead modularize with succinct interfaces, so that you can reason locally > Okay but how does AI change any of that? You can still do that with AI With your own code you reasoned about it which contributed to its stability. This meant that you could treat it like a black box. And if the abstraction leaked or was unstable, the code was still fresh enough in your head that you could evolve it and still preserve its invariants etc. With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI. | | |
| ▲ | ModernMech 2 hours ago | parent [-] | | > With your own code you reasoned about it which contributed to its stability. Okay but to what extent? People say this but there's no way to measure it really. Did you live through the 90s? People reasoned through all that code and it was very often quite unstable. I'm sure everyone involved with Windows ME reasoned about it quite a lot, probably elements of it locally were very sound, yet in totality it was an unstable mess. What fixed that situation wasn't that engineers today are reasoning better than engineers in the 90s, but IMO better tooling. Which brings me back to: your codebase is only as good as what it can prove. If there's any question, I just show you the proof rather than appealing to my reasoning being sound. > With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI. And? You haven't established reasoning about it is actually necessary and it certainly isn't sufficient. |
|
|
|
|