Remix.run Logo
▲ moshegramovsky 2 hours ago

I work on a very large code base (millions of LOC) and I've had lackluster results with autonomous work and 1 shotting. AI is definitely fantastic at working on many programming problems but I am not seeing amazing results at refactoring. In fact, I am seeing very poor results, even with Astra, even with extensive planning docs. All the recent models I've used can definitely get that refactor done, but not autonomously. It needs to be small slices. I've yet to see it 1 shot anything really complicated.

Here's a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That's a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn't. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn't follow existing API patterns. The AI could invent 10 different, new patterns when there is already 1 existing pattern they should use.

I feel like a lot of this involves a lot of babysitting prompts. Not that there's anything wrong with that of course.