| ▲ | qsort 9 hours ago | ||||||||||||||||
Very hard to say anything definitive on this because it's a moving target, but last time I tried models still had a distinct sense of "consistently good, sometimes great at micro, bad at macro". Similar to how, even for relatively pedestrian CRUD, they'll do code that's objectively fine at the function/file/class level but can still make a mess if you don't supervise them at least at a high-level. | |||||||||||||||||
| ▲ | bonoboTP 3 hours ago | parent | next [-] | ||||||||||||||||
It is moving so fast that your experience with older models is irrelevant today. Sorry. Claude Code, Fable 5, xhigh reasoning, allow it to run the full CI, end to end and benchmark, it will not make silly mistakes (or only occasionally). Also, be able to state what you desire. Have any docs or materials in the same directory so the model can reference it. For even better results: turn on speech recognition and braindump what you know about the system, its goals, its context, history anything relevant, any gotchas you'd explain to a new employee or intern. Talk for 5-10 minutes. This is optional, "make it faster" can already get a large part of the job done. And if it doesn't work well, describe what you dislike in its solution and tell it. Even just one extra iteration can make things work. (I guess GPT-5.6 can be similarly good, I use Claude) I feel like some people are emotionally invested in it not working and subconsciously sabotage their own effective use of the tool. | |||||||||||||||||
| |||||||||||||||||
| ▲ | sigmoid10 4 hours ago | parent | prev [-] | ||||||||||||||||
Which model and harness did you try? | |||||||||||||||||
| |||||||||||||||||