Remix.run Logo
▲ saghm a day ago

> It may surprise some people here to see that Andrew is warming up to using LLMs to discover bugs (inspired by results from SQLlite) and considers it a tool on the path to getting to bug free software.

Someone in the thread below says their bug was closed because of mentioning that they use AI to confirm the bug they had encountered. Is he going to go back and reopen all of those now that he learned what pretty much everyone else already knew?

▲idle_zealot a day ago | parent | next [-]

> Someone in the thread below says their bug was closed because of mentioning that they use AI to confirm the bug they had encountered

I don't know about this particular case, but if I saw someone report a bug and as evidence claim they had X Y and Z LLMs verify it I would be pretty upset. If you're going to use an LLM to make a replication, just do that and give me the replication, don't point to your notoriously error-prone tools as though they lend your report credence.

It's in a similar vein to people who reply to questions with "well Claude says: <chat transcript dump>"

▲manwe150 a day ago | parent | next [-]

I feel very torn as a maintainer on this, since the only way to respond to the increased noise from AI has been to have AI do the research for me to extract all the links and line numbers I used to have to find by hand to explain why the PR needs more effort to be completed. But I also will be highly dismissive of any submitter who just posts AI text without cleaning it up first. It feels hypocritical, but the alternative is that I just can’t respond to most people instead due to limited bandwidth. I mark if something comes directly from the LLM though and to try to express my degree of confidence in its claims.

▲childintime 20 hours ago | parent [-]

Seems like we might need to go to different issue tracking model for AI reports. Such reports must not only provide a bug report/issue but must provide a fix as well. That fix must then be reasoned back to why it solves the bug, why it is minimal, why it is unlikely to silently cause new bugs. It can't introduce new capabilities/side-effects. The platform verifies the claims in the report using another LLM. Then the tests are run, which are no longer publicly accessible to avoid slipping malicious code through the cracks. When okay the patch can go into staging. Perhaps a special version containing only accumulated AI fixes.

▲xdavidliu a day ago | parent | prev | next [-]

> It's in a similar vein to people who reply to questions with "well Claude says: <chat transcript dump>"

or substantially worse: "<chat transcript dump>"

▲kingcauchy a day ago | parent | prev [-]

Human reproductions are notoriously error prone whether ai assisted or not right? I’m not sure what the analogy is between “llms are error prone” and “thoughtlessly copy-pasting something from Claude” is.

▲dwattttt a day ago | parent [-]

Human reproductions don't have to be bad. It feels just as justified to push back on a bad bug report whether human or AI and say "I don't have enough to go on here".

If a report is improved and becomes actionable, that's great.

▲tech_hutch a day ago | parent [-]

Hey, we're all the result of human reproductions.

▲kingcauchy a day ago | parent [-]

Might be my favorite comment ever now.

▲B4uler5 a day ago | parent | prev | next [-]

He specifically says in that presentation that they are open to LLMs helping them get to bug free, but because the language is still in flux, they would rather prioritise bugs users actually find rather than those found by LLMs, essentially with the intent of unblocking people rather than wasting time fixing things that may need to be fixed again or be wasted work come the next update.

▲saghm a day ago | parent [-]

My understanding of the situation is that the user did find a bug themselves, because it literally broke on their own code, and then they just had the LLM help them verify their theory of the bug that they already had.

▲dnautics a day ago | parent | prev | next [-]

> Is he going to go back and reopen all of those now that he learned what pretty much everyone else already knew?

In the state of the tagged video he says still not accepting AI submissions until a certain set of preconditions is met. So... No?

▲saghm a day ago | parent [-]

So "warming up to" is moving at a glacial pace, in both senses of the word

▲dnautics a day ago | parent [-]

Fine but you could have checked the primary source.

▲itishappy a day ago | parent | prev | next [-]

Shouldn't everybody already know that while using AI to find bugs for oneself is amazingly efficient, using AI to submit bug reports for others is quite the opposite? Burden of verification and all.

▲dev-in a day ago | parent [-]

Exactly, the main issue isn't the LLM creating or confirming the report. Rather, the maintainer has no idea how it was prompted, and the results may be completely wrong. Some LLMs also have a bad habit of trying to please the user, confirming their biases.

▲nvlled a day ago | parent | prev | next [-]

Note the distinction:

- using the LLM to find (possible) bugs and a human confirms it by testing, reviewing, etc.

- using the LLM to find and confirm the bug without the human confirming it

▲saghm a day ago | parent [-]

Neither of those are my understanding of what happened: a human found a bug, and then had an LLM verify their theory of the issue. To me, that's just extra due diligence and pretty weird to use as grounds to ignore.

▲nvlled a day ago | parent [-]

You are referring to this comment right? https://news.ycombinator.com/item?id=49938732

> that I had every LLM check it to confirm it's a bug

I'm primarily objecting the way he phrased it, as opposed to just saying "I've tested and confirmed and reproduced the bugs". Instead, it sounds uncertain and detached, like he asked the LLMs if it's indeed a bug without further verification.

▲osigurdson a day ago | parent | prev | next [-]

I suspect they can't easily tell the difference between a high quality LLM based contribution and a low quality one.

▲rererereferred 14 hours ago | parent [-]

They want to focus first on bugs hitting people's actual code to unblock them, not theoretical bugs that LLMs can hit by coming up with some contrived code, regardless of the quality of the bug report.

▲mjburgess a day ago | parent | prev | next [-]

There's a difference between Opus 4.5 and Astra 6

▲f33d5173 a day ago | parent [-]

With regards to whether when it finds a legitimate bug, it should be given weight?

▲mjburgess a day ago | parent [-]

With regards to whether a blanket ban has a payoff

▲throwaway27448 a day ago | parent | next [-]

Why would the difference matter? A bug report is valid or it isn't.

▲manwe150 a day ago | parent [-]

A lot of bug reports aren’t valid, or handle a case that can’t realistically happen or realistically be handled (eg what do you do if you detect a crash while the prior crash is crashing and the logging pipe is throwing errors)

▲throwaway27448 a day ago | parent [-]

Sure but this is true regardless of the model. Either it found a bug or it didn't—who cares about the intermediary steps or tools used so long as the reporter can reproduce it?

▲manwe150 a day ago | parent [-]

That’s assuming the reporter reproduced it. I’ve seen so much slop these days. Sometimes people post whole conversations in which the final answer from the AI is that the initial report was impossible to fit to the data after reproducing it, then continues in circles arguing that proves the original reported issue was real and the refuted cause isn’t actually refuted (eg it said there was a bug in function A, in code which never used function A — and that example was with the top tier models just released last month)

▲physicallyIllfr a day ago | parent | prev [-]

[dead]

▲vitaminCPP a day ago | parent | prev | next [-]

citation needed

▲saghm a day ago | parent [-]

Scroll down: https://news.ycombinator.com/item?id=49938732

▲nelox a day ago | parent | prev [-]

Nobody refuses penicillin because it came from mold in a dish. If an AI found a cure for a disease, people would ask one question: does it work?

Bug fixes should get the same treatment. A patch is either correct or it isn't. Projects that ban AI-written fixes outright are asking "who wrote this?" instead of "is this right?", and users live with the bug in the meantime.

I get why maintainers are fed up. Review time is scarce, and they're drowning in plausible-looking garbage. But that's a problem with low-quality submissions, not with AI as such. Require tests, require a human who vouches for the patch and will answer for it, and ban repeat offenders. Then hold every patch to that same bar, whoever or whatever wrote it.

▲tstenner 14 hours ago | parent [-]

But if I called you in the middle of the night to offer you something that will make every 9th sneeze on every second tuesday of March in a leap year less irritating your first reaction wouldn't be to thank me.

▲saghm 5 hours ago | parent [-]

If I were making a product to try to eliminate sneezes, and someone said "I noticed an issue with it, and after I looked into your manufacturing process, I think I figured out what's causing it, so I double checked using an AI tool and it added further evidence that the defect I found is legitimate and would be solved by the solution I had come up with", it would be bizarre for me to say "I would have accepted this if you didn't spend the extra time at the end to verify your existing findings with an AI"