| ▲ | askonomm 6 hours ago |
| What I've found is that AI allows lazy and incompetent developers to be more lazy and more incompetent. This then has the effect that product quality suffers more, faster. As a result of the sheer amount of code now being pushed out, code reviews, a thing that previously somewhat prevented lazy and incompetent developers from pushing out horrible code, is effectively dead in the water since no human can actually review such amounts of code realistically anymore. Some companies have adopted AI to review code, which, well ... you have AI make code, AI review code ... I hope you can see the stupidity here if you expect to see any deterministic results at all. I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers. Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent. |
|
| ▲ | rgoulter 6 hours ago | parent | next [-] |
| > a thing that previously somewhat prevented lazy and incompetent developers from pushing out horrible code Brings to mind this classification
https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Cl... """I distinguish four types. There are clever, hardworking, stupid, and lazy officers. Usually two characteristics are combined. Some are clever and hardworking; their place is the General Staff. The next ones are stupid and lazy; they make up 90 percent of every army and are suited to routine duties. Anyone who is both clever and lazy is qualified for the highest leadership duties, because he possesses the mental clarity and strength of nerve necessary for difficult decisions. One must beware of anyone who is both stupid and hardworking; he must not be entrusted with any responsibility because he will always only cause damage""" |
| |
| ▲ | banannaise 6 hours ago | parent | next [-] | | The problem here is that AI is consistently one of the four things: hardworking. This makes it very efficient at transforming "stupid and lazy" inputs into "stupid and hardworking" outputs. Now instead of 90% stupid and lazy (harmless, useful for grunt work) you have 90% stupid and hardworking (aggressively causing damage). | | |
| ▲ | djmips 5 hours ago | parent | next [-] | | And another corollary is the formerly golden lazy and clever are also transformed into lazy and productive because they no longer need to apply their cleverness to get results... | |
| ▲ | conmod278 5 hours ago | parent | prev | next [-] | | We developed languages that removed GOTO so that developers don't shoot themselves in the foot. We will surely develop harnesses that will ensure that majorly occurring problems are solved before they hit production. | | |
| ▲ | DanielHB 5 hours ago | parent | next [-] | | Since the output of human software work is code and AI software work is _also_ code they are both liable to shoot themselves in the foot in the same manner. You see this already, LLMs are a lot more reliable in statically typed languages with strong memory guarantees (like typescript or rust) than in weaker languages. IMO the only way LLM code can avoid most of the pitfalls of human code is if we make new programming languages targeted at being used by LLMs exclusively. Think of languages with very strong methods for formal proofing and stuff like that. The problem is that even if said language was invented, it would still fail catastrophically when integrated with systems not made in said language. We are very lucky that relational databases already provide a somewhat high level of formal proofing in this regard. Said language would be impossible to parse by humans, kinda like assembly where you can parse what an isolated piece of assembly code is doing, but if you can't comprehend a somewhat large pure-assembly codebase as a whole. | | |
| ▲ | monkpit 4 hours ago | parent | next [-] | | > You see this already, LLMs are a lot more reliable in statically typed languages with strong memory guarantees (like typescript or rust) than in weaker languages. It’s common advice to wire in deterministic feedback to your workflow with LLMs - static languages aren’t inherently better for LLMs, it’s that LLMs produce better code when given deterministic feedback, such as compiler results. | |
| ▲ | leptons 4 hours ago | parent | prev [-] | | I've written very large assembly codebases, it's no different than writing in any other language. You have functions you call with inputs and outputs - though usually those are pointers to memory locations. The program is not one long function, you can split it up into different files and folders and keep everything very well organized and easy to understand and reason about. |
| |
| ▲ | datsci_est_2015 5 hours ago | parent | prev | next [-] | | Good thing GOTO was the only footgun that was ever invented in a formal programming language. | |
| ▲ | danlitt 5 hours ago | parent | prev | next [-] | | > We solved [trivial problem]. We will surely solve [incomparably harder problem]. Based on what? This will not happen! | |
| ▲ | pphysch 5 hours ago | parent | prev [-] | | Goto is a syntax feature that can be trivially removed. Good luck removing "fundamental architectural flaws". |
| |
| ▲ | automatic6131 6 hours ago | parent | prev | next [-] | | I'm going to have to remember this, gold comment | |
| ▲ | esafak 4 hours ago | parent | prev [-] | | https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Cl... |
| |
| ▲ | Melkazt 5 hours ago | parent | prev [-] | | I'm both clever and stupid, depends on the day. |
|
|
| ▲ | topherPedersen 38 minutes ago | parent | prev | next [-] |
| I'm with you on this. At my employer, I feel like we are looked down upon if we don't take the lazy approach and let the ai attempt to one shot whatever it is we're working on. One reason I'm reluctant to hand over all of my work to the ai is I don't want to forget how to program or let my skills deteriorate. Another reason is I don't want to become dependent on ai and find myself in a situation where I'm not able to fly/navigate/land the airplane if my auto-pilot or ai malfunctions or fails. Then the last reason I don't want to take the lazy approach: When I've done "one shot tests" a lot of times the ai will try and take some lazy half-ass shortcut that we would not accept if it were a human doing the work. A lot of times it just doesn't do what you ask it to do. Where I've found ai extremely helpful though is asking questions about our codebase, or asking it to build me a function that takes in a, b, c arguments and spits out x, y, z. AI really is one of the greatest things mankind has ever produced, but I don't think it's so good yet that it can replace humans completely. Using it as a form of leverage though I think is what people should be doing. I suppose we'll see what happens to developers who let the ai take over completely. Some people are arguing that if you don't let the ai takeover completely your career is doomed, but personally I think you might be doomed if you forget how to fly the airplane by hand. |
| |
| ▲ | askonomm 18 minutes ago | parent [-] | | Like I mentioned here elsewhere: AI replaces the IDE/editor, not the thinking. We now just work one abstraction level higher, but your ability to architect solutions is as relevant, if not more so, than ever before. I design test cases, architectural plans, review the output, do verification. AI just writes the code and helps me with research. If your job can be entirely offloaded to AI then I’d question if your job was that needed to begin with, but writing code was never the job, engineering was. These are just tools to do the job with, they aren’t the job itself. We solve business problems through technology and to me it’s quite concerning how many developers think their job was knowing syntax of a particular language. Nobody besides themselves care about the syntax of a particular language, certainly the business doesn’t care. I question if retaining your knowledge of programming languages is that important anymore, but your ability to read code if needed (after all, most programming languages are quite similar so it’s not that hard), do architectural and systems thinking, yes. More than ever before. Honestly, the fear-mongering around AI replacing developers seems to really just expose the developers who never learned architectural and systems thinking, and were just translating Jira tickets to code. While I do not wish job loss upon anyone, I’m not very surprised if those types of jobs will disappear. |
|
|
| ▲ | whatever1 6 hours ago | parent | prev | next [-] |
| Even if you are competent I cannot review your 5,000 lines of code you produce per day vs the 100 you were producing before the LLM apocalypse. |
| |
| ▲ | rfgplk 6 hours ago | parent | next [-] | | 5,000 is the output velocity of someone not fully immersed in agentic coding. I've seen repos do ~100k to ~250k loc changes per week. | | |
| ▲ | 0c3ca83 6 hours ago | parent | next [-] | | Yes, they're certainly squeezing 500 lines of functionality into 250,000 lines of code. Agents are great at this. | | |
| ▲ | bitwize 5 hours ago | parent [-] | | Tell me you're not using a frontier model without telling me you're not using a frontier model | | |
| ▲ | necovek 2 hours ago | parent | next [-] | | I've asked Codex with GPT 6 Astra to *review" a one time benchmarking script for any mistakes (built by Claude Code using Opus 5.5) and it refactored the shit out of it claiming all sorts of stuff without even asking about the context in which it was developed. If I was to employ them to review the code without giving each the same baseline multi-page prompt, they go into endless loop of "improvement" with no end goal in sight. More and more frequently, I instruct frontier models to stop and go back to the task at hand. | |
| ▲ | whateveracct 2 hours ago | parent | prev | next [-] | | i have unlimited tokens and i throw Fable / Astra at everything. They suck ass still for anything nontrivial. I could commit that garbage but if I kept doing it, I will end up with a ball of mud only Fable / Astra can grok..convenient for Dario and SamA.. | |
| ▲ | paganel 5 hours ago | parent | prev | next [-] | | Where's the great software, then? I'm genuinely asking: where is it? Because I can't find it, and it's been close to a year since AI for programming has started to take off. | | |
| ▲ | dawnerd 5 hours ago | parent [-] | | In fact, a lot of the once great software that's switched to AI driven development has gotten worse. | | |
| ▲ | cat-snatcher 4 hours ago | parent [-] | | For example? | | |
| ▲ | dawnerd 3 hours ago | parent | next [-] | | Clickup. Windows. Github, VSCode... Could be argued they were going downhill before but it's a much faster decline since 2023-ish | | |
| ▲ | necovek an hour ago | parent [-] | | Somebody posted that GitHub status summary: in ten years, they averaged 10 incidents per month — this includes the last 12 months. In the last six months, average is 22. |
| |
| ▲ | 0c3ca83 4 hours ago | parent | prev [-] | | Google. | | |
| ▲ | cat-snatcher 3 hours ago | parent [-] | | It was going downhill way before AI. | | |
| ▲ | intrikate 3 hours ago | parent [-] | | That's true, but also doesn't prevent the accelerated decay that it seems to be undergoing since AI hit the scene en masse a few years ago. |
|
|
|
|
| |
| ▲ | 0c3ca83 5 hours ago | parent | prev [-] | | Tell me you never once bothered to look at the generated code without telling me you don't look at the generated code. Luajit is under 80,000 lines of code. |
|
| |
| ▲ | teiferer 5 hours ago | parent | prev | next [-] | | So where is all that new software? My laptop and phone run essentially the same software as 2 or 3 years ago. Yes there were some minor updates to some apps, but nothing faster than in the years prior. Where does all that supposed productivity go? | | |
| ▲ | al_borland 4 hours ago | parent | next [-] | | The only real updates I've seen to anything have been AI features... so all this AI is only being used to add AI to stuff. Most of which the average person doesn't seem to want or use. | |
| ▲ | Powdering7082 3 hours ago | parent | prev [-] | | Number of git pushes in GH is way up. 320M in Q1 this year vs 80M in Q1 or 2020. https://innovationgraph.github.com/global-metrics/git-pushes Just because you haven't installed new software doesn't mean that new software doesn't exist. | | |
| ▲ | legulere an hour ago | parent | next [-] | | To quote the article: "Don't confuse motion with progress." | |
| ▲ | californical 2 hours ago | parent | prev | next [-] | | But also, having significantly more code churn doesn’t necessarily mean there is more or better software. In fact, having more churn can lead to worse software due to diverging patterns and inconsistency | |
| ▲ | nemetroid 2 hours ago | parent | prev [-] | | The question was "where is all that new software?". | | |
| ▲ | necovek an hour ago | parent [-] | | I see somebody asking "where is that great software", but I think it's obvious to everyone that there is more (mostly bad) software. |
|
|
| |
| ▲ | whatever1 6 hours ago | parent | prev | next [-] | | I mean these guys are not even pretending to be reviewing the code. It just gets “reviewed” by an LLM, which will find a nitpick while ignoring the huge fire in the core of the design, force the planner to make even more sloppy code to cover for an irrelevant test case. Rinse old tokens and repeat until you hit limits. | | |
| ▲ | bunderbunder 5 hours ago | parent [-] | | This is exactly what I've seen. For example, I recently got brought in to help with quality on a large-scale system that had been ported to a new platform with the help of coding agents. The project was completed and declared operational in record time, but soon after the business discovered that: 1. The promised scalability improvements did not materialize. Instead, it got worse. 2. Observability had been lost. The telemetry was no longer trustworthy. 3. Users stopped trusting it because it was producing incorrect outputs. What I ended up discovering was that, while it scrupulously kept existing automated tests passing, any behavior that wasn't explicitly covered by a test was free to change any which way. And there were plenty of small things that weren't explicitly covered. Perhaps because the original authors thought they were so obvious and commonsense that they didn't need one, perhaps because mistakes happen. The why doesn't matter. The point is that reality is messy and imperfect, so giving someone a chance to look at things and think, "Huh, that's funny..." is an essential part of defense in depth. The real worst part was, this whole replatforming was a huge waste of time, anyway. The improvements they were looking for could easily have been accomplished with some controlled incremental changes to the original system. Mostly just removing a few basic and well-known performance antipatterns. But way back at the outset, the person in charge of the project asked their agent, "What's the best way to X," and the agent gave them a trendslop answer about how Y alternative technology is more scalable and we should just port to that. It was convincing and they were under intense time pressure to just ship some code because leadership is bought into the AI hype and now has the patience of a 4 year old, so they just went with it. |
| |
| ▲ | pretendscholar an hour ago | parent | prev | next [-] | | What kind of applications are people building that involve 250k loc a week? Genuinely trying to understand this. | | |
| ▲ | mstaoru an hour ago | parent [-] | | So far I mostly see a metric sh.. ton of meta- and meta-meta-projects re-wrapping AI wrapper tools, with sloppy slogans like "One Model, Five Harnesses. Combined." or "You run in the park. Rrrunnnrr.ai runs your AI." Looking inside, out of 250k it's often 200k of verbally incontinent self-explaining comments; or "smart" redesign of builtins. No new browser, no new iOS clone than runs on Android, no new easy to use DaVinci, no new CAD suite, no $5 SolidWorks clone, no redesigned K8s, no 10x performance speedup in Linux kernel. |
| |
| ▲ | Ambolia 6 hours ago | parent | prev [-] | | Can the users of the software even keep up at that point? We may have reached diminishing returns on software production, and not enough impact on the rest of the process. |
| |
| ▲ | bitwize 5 hours ago | parent | prev [-] | | That's okay. Reviewing the code will become the agents' job as well. A couple more step functions in model capability of the type we've seen in the past year, and there will pretty much be no reason for humans to be involved in the development process at all. All humans would need to do is communicate clearly what needs to be made and flag problems as they come up. | | |
| ▲ | weatherlite 3 hours ago | parent | next [-] | | > All humans would need to do is communicate clearly what needs to be made and flag problems as they come up. Kinda what i'm doing already, but for the young startup I'm at that's surprisingly tons of work. I miss the days we wrote code by hand boy those were fun 8.5 hours workdays. | |
| ▲ | xpct 5 hours ago | parent | prev | next [-] | | "A couple more step functions" is doing a lot of heavy lifting here | | |
| ▲ | pydry 4 hours ago | parent [-] | | It's the AI bro's mantra. It didnt go wrong And if it did, it was because you werent using the latest model. And if you were, it was because you didnt have the appropriate guardrails. And if you did, it's because you didnt have AGENTS.MD. And if you did, it's because you didnt prompt it properly. And if you did, it you're still going to be redundant soon because I'm sure the next model released will fix whatever went wrong. | | |
| ▲ | mywittyname 3 hours ago | parent [-] | | I've stopped calling out Claude mistakes on team meetings because this is so true. I mean, sure, I could have predicted in what ways an LLM would fuck up, but there's just so many ways I can't keep up. We just had a major production issue because someone's LLM wrote queries against dev databases. Which are very obviously dev databases because they are labelled with dev in the name, and in the table descriptions. AI reviewer didn't catch it, neither did the human reviewer for that matter. |
|
| |
| ▲ | necovek an hour ago | parent | prev [-] | | > All humans would need to do is communicate clearly what needs to be made and flag problems as they come up. Sounds like the easiest thing in the world: I wonder why did we not think of it earlier? |
|
|
|
| ▲ | patorjk 5 hours ago | parent | prev | next [-] |
| I'm seeing this too. I've worked with devs that would previously push PRs that wouldn't work or run correctly. Those PRs wouldn't get merged in. Now they're putting up PRs which seem to work at first glance, but have hidden problems. For example, one guy introduced a huge PR for a visualization and it seemed to work fine, though another dev mentioned to me that we already use recharts and it does 90% of what this guy's PR does (his code does all the drawing logic itself). Maybe AI will get good enough to clean up these kinds of messes, but in the near term I imagine there will be a lot of code bases that will be filling up with dragons. |
|
| ▲ | ben_w 6 hours ago | parent | prev | next [-] |
| Limitations of AI are a thing; but one rhetorical point keeps coming up (I don't think it's just you) and confusing me: > I hope you can see the stupidity here if you expect to see any deterministic results at all. Are you expecting humans to be deterministic in the code they produce? |
| |
| ▲ | Thanemate 6 hours ago | parent | next [-] | | Someone who knows that 1 + 1 = 2 will not decide that it's suddenly 3 unless we start accounting for health problems. Making mistakes is not the same as non-deterministic. | | |
| ▲ | InsideOutSanta 5 hours ago | parent | next [-] | | Neither will LLMs. That's not how their nondeterminism works. | | |
| ▲ | ofjcihen 5 hours ago | parent [-] | | I think that actually reinforces the distinction being made. An LLM’s nondeterminism is in the generation process: given the same prompt and model state, sampling can produce different outputs. That doesn’t mean the underlying fact itself becomes nondeterministic. A human who knows 1+1=2 can still say “3” because they misread the question, misspoke, were distracted, or made some other cognitive error. Likewise, an LLM can output “3” because the generation process selected an incorrect continuation. Those are both errors in producing an answer, not evidence that 1+1 somehow has multiple answers. So yes, human mistakes and LLM sampling are mechanistically different. If your argument is that LLMs and humans can both make mistakes, then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans. | | |
| ▲ | InsideOutSanta 5 hours ago | parent [-] | | > If your argument is that LLMs and humans can both make mistakes It's not, I'm just pointing out that LLMs won't make that mistake. You could ask an LLM what 1+1 is, and the number of times it says "3" is so small that it makes no sense to worry about it. It will phrase the response differently each time; that's the nondeterminism. But it won't say "3". > then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans. Yes, if we ignore everything else, that seems like a reasonable question. But let's not ignore everything else, like the fact that LLMs are much more productive than humans and likely already make fewer mistakes than the average programmer. | | |
| ▲ | lolakutty 4 hours ago | parent [-] | | >You could ask an LLM what 1+1 is, and the number of times it says "3" is so small that it makes no sense to worry about it... I think the disturbing fact is that you can take a frontier model with all the intelligence of humanity, and make it say 1 + 1 = 3, by specifically training for it... A human with that much knowledge will refuse that attempt. There in lies the difference.. | | |
| ▲ | thaumasiotes 3 hours ago | parent | next [-] | | > A human with that much knowledge will refuse that attempt. Well, that's not true. https://www.youtube.com/playlist?list=PLO3a3Ax6Yh6bbtKuxfYBP... | |
| ▲ | monkpit 4 hours ago | parent | prev [-] | | What’s your point though, really? “You can train a model to say things that are objectively wrong”? You can do the same with a human. | | |
| ▲ | lolakutty 4 hours ago | parent [-] | | Did you read what I wrote to the end? | | |
| ▲ | monkpit 4 hours ago | parent [-] | | Yes, I fail to see anything meaningful. If you move the goalposts and say “I have invented a human that cannot be convinced in any way to give a wrong answer” then what’s the point of that in this discussion, really? And you seem confused about how an LLM works and what it is - “the intelligence of all humanity” - not how it works. You’re debating using 2 imaginary things you created. | | |
| ▲ | lolakutty 3 hours ago | parent [-] | | Tell me how you convince a human with all knowledge we have, that 1 + 1 is 3. | | |
| ▲ | ben_w 2 hours ago | parent [-] | | Trivial. Hit them with a stick until they answer as you told them to. I think the post up-thread, https://news.ycombinator.com/item?id=49881653, was trying to make this point by linking to an episode of Star Trek TNG, with Picard being tortured until he said the "correct" (incorrect) number of lights. (I recommend against using fiction as evidence; in this case the general point happens to be valid, and is why torture is forbidden: we humans really do break, but breaking doesn't mean we tell the truth, it means we tell people what we think they want to hear). | | |
| ▲ | lolakutty 2 hours ago | parent [-] | | >Hit them with a stick until they answer as you told them to. Obviously, for this purpose, human should not have any feelings (because LLMs don't have), so can't feel pain. Or else the comparison can't work. | | |
| ▲ | ben_w 2 hours ago | parent [-] | | > Obviously, for this purpose, human should not have any feelings (because LLMs don't have), so can't feel pain. Or else the comparison can't work. Other than this forcing you to ignore the overwhelming majority of humans who have functioning pain nerves: LLMs have something functionally equivalent to pain, in this regard at least. During training, model weights are updated depending on if the feedback was positive or negative. It has a functional effect similar to that which pleasure and pain have with us. Not identical, so far as I know there's not been any reports of any machine learning model that is into BDSM, but for the most part functionally similar. There is also research which has found circuits in multiple LLM models, which serve similar roles at inference time and are distinct from other emotional representations: https://arxiv.org/abs/2609.16247 |
|
|
|
|
|
|
|
|
|
| |
| ▲ | ben_w 6 hours ago | parent | prev | next [-] | | > Someone who knows that 1 + 1 = 2 will not decide that it's suddenly 3 unless we start accounting for health problems. And? The p(that kind of error) is pretty small now. At what point does a probability coming out of an LLM look like "knowing", such that spitting out the wrong answer despite that probability looks like a health problem, a typo, or even just boredom? (Thinking of the Lizardman constant here: https://en.wiktionary.org/wiki/Lizardman%27s_Constant) It's a continuum for both them and us, even if the mechanism is wildly different. > Making mistakes is not the same as non-deterministic. i.e. when the dismissal is "non-deterministic" when it should be "Making mistakes", is itself a mistake. | |
| ▲ | p-e-w 6 hours ago | parent | prev | next [-] | | Lol, humans make such absurd mistakes (and worse) all the time through simple typos, which is effectively random. The key for 2 is right next to the key for 3, after all. | |
| ▲ | thaumasiotes 3 hours ago | parent | prev [-] | | > Someone who knows that 1 + 1 = 2 will not decide that it's suddenly 3 unless we start accounting for health problems. But this is plainly false. This kind of unforced error occurs all the time. For example, once when I was in high school I traced an error in my math homework to an intermediate calculation of "2 + 2" as being "3". There was no reason. What we can say about humans is that, if they know that 1 + 1 = 2, (a) they are unlikely to change their mind about this in any kind of lasting or permanent way, and (b) the rate at which they will mistakenly produce other values for 1 + 1 is very low. But it will happen occasionally, and when it does happen, "they just suddenly decided on the wrong value" is an extremely accurate description of what that looks like. |
| |
| ▲ | rhdunn 5 hours ago | parent | prev | next [-] | | By not reviewing, reading, or understanding the code generated by agentic LLMs the output is effectively like a compiler. However, a compiler has deterministic behaviour that can be repeated and verified. The behaviour/output of an LLM is not like that. Ask an LLM to create a dashboard to show games by genre and it will generate different results with each run, and each model/model version produces wildly different results. | |
| ▲ | lkjdsklf 5 hours ago | parent | prev | next [-] | | The difference is that with llms you have multiple levels of nondeterminism compounding each other | | | |
| ▲ | ex-aws-dude 4 hours ago | parent | prev [-] | | I've seen LLMs do something correct 98% of the time then randomly do something crazy that a human would never do because we have continual learning As humans we don't have our memory reset multiple times per day | | |
| ▲ | ben_w 3 hours ago | parent [-] | | > I've seen LLMs do something correct 98% of the time then randomly do something crazy that a human would never do because we have continual learning I've seen humans vote for Brexit, re-elect Trump, ask questions clearly already answered in an FAQ, try to pull on a door labelled "push", and insist on giving me homeopathic silicon dioxide pills* that cost £5** for a 10-12 gram packet. Continual learning is a difference, but not by itself a reason to care about "deterministic results". Nor, indeed, correct results. > As humans we don't have our memory reset multiple times per day Humans need sleep well before they can read a million tokens' worth of written text. We're more like 300k tokens if you're actually reading and not skimming for 16 hours straight. Again, different (in soooo many ways), but this isn't a relevant difference when the topic is "deterministic results". * yes, sand: https://dailymed.nlm.nih.gov/dailymed/fda/fdaDrugXsl.cfm?set... ** and that was what it cost in the 90s | | |
| ▲ | ex-aws-dude an hour ago | parent [-] | | If I do the same task 100 times I'm not going to suddenly do it crazily different at time 101 because I've built in the memory of how to do it There is no RNG involved when I decide to push vs pull the unlabeled door to my building every morning, it becomes deterministic because its baked into memory You can put stuff in context to deal with this but you can't do that for everything, its not practical and you would blow the context window |
|
|
|
|
| ▲ | Kuyawa 2 hours ago | parent | prev | next [-] |
| > forcing companies to increase the quality of their developers Just don't. Fire them! AI is better than a thousand devs. What you need is testers that know what to test that AI can't, not code or UX/UI (not talking about playwright here) but business intelligence if that is testable, the things that produce results (profits) and the reason it was asked for in the first place, to solve a problem If the problem was asked wrongly, the result will be wrong too. Fire devs, then PMs, then IT Managers if they really don't know how to outperform AI, and that's exactly the point, they won't be able to do it in code or tests or reviews, only in intelligence, for now... |
|
| ▲ | bwfan123 5 hours ago | parent | prev | next [-] |
| At a startup I worked, there was an engineer whose code was incoherent and buggy. So, we were literally better off if that engineer did nothing because their net output was negative. Engineers like that become weaponized with LLMs, and negative numbers become larger negative numbers when scaled up. |
| |
| ▲ | icedchai 4 hours ago | parent | next [-] | | I've seen similar. They wasted weeks of senior engineering time, between reviews, meetings, and follow up in Slack, only to have the PR closed without merge. The offending individual was eventually moved to another project. | |
| ▲ | ilaksh 4 hours ago | parent | prev [-] | | Is that the fault of AI or management for not firing them? | | |
| ▲ | bwfan123 4 hours ago | parent [-] | | > Is that the fault of AI or management for not firing them? How does the system behave in a variety of scenarios including failures and restarts. How is state maintained coherently. There are the kinds of systems problems that an engineer needs to reason through, and if there are bugs in such decisions, they end up becoming costly. I dont expect AI or LLMs to solve these problems at all, since each of them has nuances and tradeoffs which are specific to each system. In short, there is specification complexity in precisely describing system wide behaviors, and unfortunately, there is no lean/tla+ to meaningfully describe systems at scale. You could then ask: How can a system have guaranteed behaviors if they cannot be even stated or proved formally ? The answer to this is how protocols like raft/paxos initially convinced us of their behaviors which is in human review and understanding. That begs the question: How can human review and understanding be reliable, and the answer is that it is not reliable, but humans have ability and processes to continuously learn from experience in the real world. So, our understanding is grounded not only by whats out there in books etc, but also by our own interactions with the world. Long story short: The responsibility for system-wide behaviors of software systems relies on human review and understanding, which while imperfect can continuously learn. |
|
|
|
| ▲ | mjr00 6 hours ago | parent | prev | next [-] |
| > What I've found is that AI allows lazy and incompetent developers to be more lazy and more incompetent. This then has the effect that product quality suffers more, faster. Yeah. To me it seems very much like the "use dynamic typing for everything" fad. You had a bunch of junior and/or incompetent developers who went around insisting that type declarations are bad, static typing slows down development, you just code so much faster if everything is dynamically typed. And in the context of a new project, they were totally right. It took a few years for the debt to finally catch up, and people realized that these massive, untyped monoliths they had were unmaintainable. Now the two biggest dynamic languages (Python/JavaScript) are effectively typed languages, because nobody uses their untyped variants for serious work. Dynamic typing still has great uses -- interactive data exploration, putting together quick scripts (though less relevant with AI...), or even just simple prototypes -- but what we tried to do with it at the start, as an industry, was clearly dumb as hell. I suspect we'll look back in 5-10 years and realize that with some of the stuff we're doing with AI, too. It's already happened with things like Gastown. |
| |
| ▲ | paganel 5 hours ago | parent [-] | | The web wouldn't have taken off without dynamic typing, PHP first of all (and Python/JavaScript after that). People seem to forget how atrocious it was to write an .asp or .jsp (I think the extension was .jsp) page back in 2003-2005. | | |
| ▲ | bdangubic 5 hours ago | parent [-] | | hey man, don’t knock the JSP, I just edited a few :) it is alive and kicking in 2026 | | |
|
|
|
| ▲ | huijzer 6 hours ago | parent | prev | next [-] |
| > Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent. What I in general try to teach the other people about AI: It can be a great tool, but check the results! Especially in the case of engineering: Check and then double check. |
|
| ▲ | hanifbbz 6 hours ago | parent | prev | next [-] |
| In other words AI is a multiplier. |
| |
| ▲ | rfgplk 6 hours ago | parent [-] | | "optimize this code", "fix this code", "extend this code", "add this feature", "find errors and patch them", "find bugs and fix them", "rewrite this from python to rust". This is all that's needed to actually use LLMs nowadays. How is it a "multiplier" rather than an "equalizer"? | | |
| ▲ | askonomm 5 hours ago | parent | next [-] | | If that's how you create software then you belong to the lazy and incompetent group in my book. I provide AI with valuable context such as code coverage information, architecture analysis, test requirements, important "gotcha's" that a competent engineer would know about in their architecture or system etc. I'm still very much the person who comes up with the solutions. For me AI is replacing the code editor, it's not replacing the thinking. | |
| ▲ | swiftcoder 6 hours ago | parent | prev | next [-] | | > How is it a "multiplier" rather than an "equalizer"? Because without the responsible human engineer in the loop, it'll all gradually decay in a cascade of edge-cases. This happens with human written code as well (every "we'll replace this prototype before we ship" you've ever worked on), but with LLMs it happens at 10-100x the rate. | | |
| ▲ | bigfishrunning 5 hours ago | parent [-] | | > every "we'll replace this prototype before we ship" you've ever worked on These so rarely get replaced | | |
| ▲ | bcrosby95 5 hours ago | parent [-] | | This is why it's good to not keep your prototype a pile of shit as it grows to 5k, 10k, 50k, 100k lines of code. |
|
| |
| ▲ | CuriouslyC 4 hours ago | parent | prev | next [-] | | Not completely, just as a very personal example: In optimizing my game I noticed framerate hitching even after efficient algorithms were in place for expensive stuff, which was caused by shaders not being precompiled consistently or assets not being preloaded in time. The Agent who'd been profiling and optimizing had moved many preloads to a loading screen, which caused a long loading lock, and what it didn't move ahead was loaded and compiled at use, creating slow frames since work was being done on the main game loop. I instructed the agent to create a speculative pre-warming/compilation priority queue with a per-frame budget, with priority being determined by likelihood signals that the asset or shader will be used soon. Then I had the AI run fully headed games and hunt down causes for frames going over 16ms, and work through them until a batch of games had fewer than 1/1000 frames >16ms and no frames over 60ms after a short initial settling period. The approach, the metrics, the validation system and the loop were "prompt engineering" above and beyond what I would expect from someone who was merely "vibe coding a game." | |
| ▲ | sortoflog 5 hours ago | parent | prev | next [-] | | The skill floor has definitely been lowered, but if this were actually true then firms would be replacing senior software positions with entry level ones, not the other way around. | |
| ▲ | 5 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | dnikolovv 6 hours ago | parent | prev | next [-] | | Do you use the word "equalizer" in this context to mean that AI has made the playing field equal for both competent developers and laypeople? Do you reckon that competence plays no role these days? | |
| ▲ | zxor 4 hours ago | parent | prev [-] | | If this is how you use LLMs, you are the problem. |
|
|
|
| ▲ | chanux 5 hours ago | parent | prev | next [-] |
| > AI allows lazy and incompetent developers to be more lazy and more incompetent. I like to put this as "LLMS give lazy and incompetent developers more runway." |
|
| ▲ | Zardoz84 6 hours ago | parent | prev | next [-] |
| > you have AI make code, AI review code ... I hope you can see the stupidity here... You will be surprised how many times, catches errores made by the AI coding agent. However,as you point, isn't deterministic. And you can guarantee the end results is 100% fine code |
|
| ▲ | empath75 6 hours ago | parent | prev | next [-] |
| Honestly, I would still rather commit claude written code from lazy and incompetent developers than code that they wrote. |
|
| ▲ | ls-a 6 hours ago | parent | prev [-] |
| [dead] |