Remix.run Logo
▲ maherbeg 2 hours ago

There's lots more you can do! Use the model to monitor your deployments after they get deployed. Have them fix and watch CI issues for you. Run adverserial review. Automatically watch metrics every day and highlight performance regressions. Start reviewing your previous sessions to find ways to statically reject different failure modes and have the agent have more success earlier on etc.

Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?

▲klardotsh 31 minutes ago | parent | next [-]

The thing with watching CI in an agent loop is that it burns tons of tokens. At work I ended up writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code, and then updated my `/glab-ci-feedback` skill to use that. Saved a ton of token churn, and now I have a runbook a human could just as easily use if they don’t want to (or can’t) use an agent loop.

… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.

▲unddoch 10 minutes ago | parent [-]

I think they are trying now to to bake CI awareness into Claude Desktop, didn't use it yet.

But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already

▲Hauthorn 41 minutes ago | parent | prev | next [-]

> Another thing to think about is, what would it take for you to care less about the understanding.

Could you explain why it would be a goal to understand the system less, rather than more?

It seems harder to know if you have good tests while lowering your expertise in the system.

▲ 26 minutes ago | parent [-]
[deleted]
▲datadrivenangel 2 hours ago | parent | prev | next [-]

Opus 5.5 on Low seems smarter, cheaper, and faster than sonnet on medium, so what's the point of sonnet?

▲xgb84j 44 minutes ago | parent [-]

Claude Code has the issue that sub agents inherit the thinking level. This means that to use a smarter or dumber sub agent you need a different model. That's not a particularly good reason, but that's my one use case for Sonnet.

▲copperx 37 minutes ago | parent [-]

Or just use a better harness.

▲crooked-v 27 minutes ago | parent | prev [-]

> Run adversarial review.

Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".

I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.