Remix.run Logo
▲ Jev-Driven SRE Diagnosis: What Worked and What Failed(sregym.com)
31 points by matt_d 7 hours ago | 10 comments
▲clintonb 6 hours ago | parent | next [-]

> Ultimately, we believe that incorporating Jev-driven diagnosis into an SRE agent’s workflow is a significant step toward effectively combining System One and System Two models.

Why? What’s the end goal? Lower cost? Faster analysis?

I configured an agent to respond to pages via Slack. It has access to ClickStack (for telemetry), Kubernetes, and GitHub. It runs one of the Sonnet models. The cost is so low, responding to less than a dozen pages a day (mostly from a very sensitive error count alarm) that swapping got Jev makes zero sense. This is especially true if the results are less trustworthy.

▲victor9000 28 minutes ago | parent [-]

using jev is a significant step towards using jev

▲beebmam 5 hours ago | parent | prev | next [-]

It seems to me that Jev is designed for when many quick decisions, with low input context, need to be made. SRE work is probably best suited for very few critical decisions that need to be made, with high input context.

▲olgava an hour ago | parent | prev | next [-]

> SRE work is probably best suited for very few critical decisions

Yeah, and the article puts the cost of all 105 diagnoses at about $0.15 in Jev calls. At a dozen pages a day, I wouldn't worry much about the bill for either model. I'd be more interested in the pass rate: 76.2% vs 77.8% for GPT-5.6 Sol (medium).

I can see trying a cheaper model first if you're handling lots of requests and it can resolve most of them without escalating. A dozen pages a day doesn't seem like a reason to add that extra step.

▲amne 2 hours ago | parent | prev | next [-]

I switched to a decision model for my local "AI" voice control thingy. faster-whisper produces the text then onnx scores all my HA entities to determine which one I'm talking about .. then it lists all its actions and a second scoring round to determine what action I want. some regex to extract numbers if I'm saying things like "water the lawn for 10 minutes". It's pretty dumb when compared to what an LLM can do but considering you now have semantic scoring "at your finger tips" without resorting to levenshteins or other crude algorithms like that it's insane how smart it can look.

And to answer the obvious question: because latency. This thing turns on the light or water valve in under 100ms on a laptop from 2001 with 8gb ram .. all running locally on the laptop.

▲soltanov 4 hours ago | parent | prev | next [-]

Premise is flawed: SRE is not a throughput problem, it is an accuracy problem.

▲sdcfgy an hour ago | parent | next [-]

Yes and with an incredibly large context and incredibly large domain specific knowledge.

We had an LLM based SRE product from a vendor I won't mention because we had an NDA as our management is sucking them off by trading whitepapers for discounts. It was like a drunk monkey with a wrecking ball. Think we had to pull it in under 2 weeks because it took out multiple production systems and lead to an entire cluster failover.

The meat sacks now know they have job security.

▲jasonjmcghee 2 hours ago | parent | prev [-]

I'm guessing you're getting downvoted due to dismissive tone, but I think there's truth to this.

Labeling something as "don't escalate" that should have been isn't great. Paying 5x and having to wait 10s instead of 100ms or whatever for a reasoning model (with tools?) is very likely worth it.

▲endangeredhuman 2 hours ago | parent | prev | next [-]

This is an interesting experiment. Curious why you used LLM-as-a-Judge (gpt-6-astra). Did we have a ground truth of actual RCA done by a human to compare against?

▲N_Lens 4 hours ago | parent | prev [-]

Lots of hype-mongering around Jev atm. Color me sceptical.