Remix.run Logo
StellaSearch 2 days ago

Fully agree. I found for most of my work with LLMs and Finance, ~90% of the time a high or low confidence score was accurate. There's the occasional ambiguous case, and that'll happen, but the engineering work that comes with building a classifier makes it not practical for my usecases.

inspectorSlap 2 days ago | parent [-]

I'm working on a universal one. My belief is that as agentic capabilities increase, behavioral attribution will be increasingly needed to maintain quality.