Remix.run Logo
▲ chamomeal 2 hours ago

I agree with many perspectives in this thread. I empathize with OP. I’m scared about the future of my job and already feel it less fulfilling, despite being more productive than ever.

But I have one singular counterexample to most scary narratives and essays about LLMs writing software, which is my coworker who hardly uses LLMs at all. At least, he hardly uses them compared to me. He still writes most of his code by hand. He still greps around the codebase without claude code. Like he’s definitely using claude code, but only for like one-off specific tasks.

He’s a totally average developer, like everybody else on my team (including me). But he’s clearly more helpful than the rest of us. When people from other teams have questions about how something works, he’s always the first to respond. He’s always the one providing useful context in planning calls. He catches stuff in code review that I didn’t catch, and claude/copilot didn’t catch.

I think I’ve leaned too far into AI, partly because I’m a lazybones and am kind of burned out already. But there’s a stark contrast between me and my coworker that there didn’t used to be. I think the context rot is really setting in, and getting worse. So I think there’s still value in caring about the details

▲enraged_camel an hour ago | parent [-]

>> He catches stuff in code review that I didn’t catch, and claude/copilot didn’t catch.

This, I find completely unbelievable. Because we had (emphasis on had) seasoned engineers on the team who similarly eschewed AI tools and insisted on doing everything by hand in the manner you describe, well after the rest of the team adopted AI. Yet when it came to code reviews, even in parts of the codebase they were familiar with, the bugs they found often came down to nits, bike-shedding and opinion-based feedback (that they usually tried to frame as objective fact). They would also disagree with almost every AI finding, arguing that it was an unrealistic scenario or an edge case not worth worrying about.

Fundamentally, I don't think humans are going to be capable of providing high quality feedback on PRs authored by AI agents unless those PRs are fairly small in lines of code and volume. It's just way too much information and context for one person to keep in their head. I read a statistic that said the average lines of code a senior engineer can read and provide good feedback on is about 400 per hour, and that number goes down the more time they spend doing code reviews. So, to anyone who insists on trying to keep up with AI, I say: good luck.

▲majormajor 27 minutes ago | parent | next [-]

Claude's code review skill, in particular, can find some good stuff. But it has some big blind spots around certain types of code. And it likes to come up with a lot of nits too—I think it's really really trained to try to always find between 2 and 8 things or somesuch. Good news is that it is very receptive to "nah" on the bikeshed ones and doesn't stick with them, but will stick with big issues. It'll probably bring up a few more nits though that it didn't bring up the first time!

But I can completely believe that someone who knows the code by heart would have a better signal to noise ratio on their reviews.

I'm trying to find the sweet spot because I've found some NASTY bugs Claude missed, and also had Claude find some nasty ones for me. And this is in codebases with tens-of-thousands of AI-generated lines of code + AI-driven reviews. So I want to bring both to the table.

The existence of some of these major "oh man that changes a lot of our assumptions" bugs that were only found because someone poked on the agent and said "I don't think you're paying enough attention to this" justifies that, IME.

And the better you are at pointing the agent at the truly-important parts, the better the agent's gonna be at finding shit you missed.

▲geraneum an hour ago | parent | prev [-]

> This, I find completely unbelievable.

You both have anecdotes. Anecdotes don’t “cancel” each other out.

Here’s a third one. In some of the code reviews I’ve encountered that AI gives a lot of feedback, it’s just providing noise. Things that should be ignored or when following the feedback causes more harm which requires more token to “fix” later on. That can also happen. Sometimes the thing it spits out goes against the common sense, and sometimes it works very well.

> So, to anyone who insists on trying to keep up with AI, I say: good luck.

This I agree with, for a different reason. It’s like trying to swim in a sea of honey and trash mix. It’s exhausting.

▲abuani 37 minutes ago | parent | next [-]

> Here’s a third one. In some of the code reviews I’ve encountered that AI gives a lot of feedback, it’s just providing noise

Something I've found fun is seeing how long it takes for an llm review tool to come back satisfied with a PR. Think 100 lines of code changed, nothing terribly significant, but also not trivial. I'll have a local Claude session setup to babysit the PR and wait for feedback, accept all the recommendations, push the change up and request a review. I cap the number of iterations at 10 just so I'm not blowing a stupid amount of money. I've yet to come up with a PR where the llm reviewer is satisfied with the changes and has _no feedback_.

So where's the reasonable cutoff point for llm based reviews?

▲enraged_camel 19 minutes ago | parent | prev [-]

>> Here’s a third one. In some of the code reviews I’ve encountered that AI gives a lot of feedback, it’s just providing noise.

This is completely normal if you haven't written your own custom skill with instructions on what types of issues the AI should emphasize, which ones would be considered nits and which ones aren't a problem at all.

In our repos we use a classification system: blocker, should-fix and nit. Each one has specific definitions, criteria and examples encoded in the skill file. When the time comes to review a PR, agents invoke the skill, and frankly do a stellar job. A human then reads each finding, asks the AI follow-up questions and makes the final decision in terms of whether the finding goes in the PR review.

The reason I know this works is that we have one guy on the team who does not use this skill, and blindly throws his agents at PRs. And the results are exactly as you describe.