Remix.run Logo
mritchie712 7 hours ago

in short: it's faster, cheaper, smart structured output.

each "question" is answered in parallel instead of a sequential (like an LLM). so if you have an input like:

    {"is_it_hotdog": noul, "is_it_apple", noul}

it answers is_it_hotdog and is_it_apple in parallel and gives a probability.
satvikpendem 6 hours ago | parent [-]

Can't I just parallelize my LLM calls myself for each question?

orbital-decay 6 hours ago | parent | next [-]

You can. It will be expensive, slow, and less reliable than a specialized model.

zwily 5 hours ago | parent | prev [-]

Anything you can do in Jev can be done with an LLM at much greater cost and latency.

mmnfrdmcx 5 hours ago | parent [-]

Agree, except the probabilities for outcomes in the structured output. I don't think you can get those for most frontier LLMs (logprobas). You can get it for open source models but not frontier LLMs.

esafak 4 hours ago | parent [-]

That number is a big deal, assuming it is well calibrated. Did they talk about calibration?

Matticus_Rex 3 hours ago | parent [-]

I've seen them talk about it a bit on Twitter -- it seems to be fairly well-calibrated in general, but obviously you need to test it on your use case and dial it in comparison with known data for best results.