| ▲ | Jeeves. Reasoning improves Jev-like decision models(github.com) |
| 78 points by nicowaltz 2 hours ago | 34 comments |
| |
|
| ▲ | sharih 17 minutes ago | parent | next [-] |
| What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast. |
| |
| ▲ | zihotki 8 minutes ago | parent | next [-] | | I would hold your horses to paint it as dirt cheap.. In my cases for spam detection Luna was 20% cheaper due to prompt caching, although not as fast. | |
| ▲ | esafak 4 minutes ago | parent | prev [-] | | Jev ought to offer a flex mode that uses their spare capacity for a discount. |
|
|
| ▲ | thm 43 minutes ago | parent | prev | next [-] |
| Ask Jeeves - Only took us 30 years to come full circle. |
| |
| ▲ | tmnstr85 34 minutes ago | parent [-] | | this was the comment i came here for | | |
| ▲ | onaclov2000 11 minutes ago | parent [-] | | My bots are all named Jeeves lol. I have a CLI tool I use that connects up to a LLM I made and I call it Jeeves too ...so funny. I really didn't use Jeeves all that much I tended to use...I think it was called Web crawler pre-google era |
|
|
|
| ▲ | RamblingCTO 43 minutes ago | parent | prev | next [-] |
| Super dope. If it would ship as prod ready code supporting mps as well that would be even doper. But funny that jev is getting its lunch eaten apparently in under two weeks? |
| |
| ▲ | pavlov 32 minutes ago | parent [-] | | It’s ok, one week of AI hype is now enough to close a billion-dollar term sheet with VCs. |
|
|
| ▲ | swader999 an hour ago | parent | prev | next [-] |
| Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work. |
|
| ▲ | Naitik88 37 minutes ago | parent | prev | next [-] |
| what about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions. |
|
| ▲ | zerop an hour ago | parent | prev | next [-] |
| Are there "good" Open source Decision models built on Gemma-4 and also trainiable on own data? |
| |
|
| ▲ | woadwarrior01 an hour ago | parent | prev | next [-] |
| This isn't really surprising. LLM reasoning and before that, chain of thought prompting are essentially forms of test-time compute scaling. |
|
| ▲ | mxkuzn 42 minutes ago | parent | prev | next [-] |
| interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions. |
|
| ▲ | captainbland 40 minutes ago | parent | prev | next [-] |
| See if it can beat Jev's Pokémon benchmark |
|
| ▲ | AnodicElegy 26 minutes ago | parent | prev | next [-] |
| I'm surprised we haven't seen a "Jehovah" yet. |
| |
|
| ▲ | alienbaby 2 hours ago | parent | prev | next [-] |
| Just curious, where has this term 'noul' come from for yes/no ansers? /a bit more digging and.. A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true. I hate it :) |
| |
| ▲ | LudwigNagasena 3 minutes ago | parent | next [-] | | In Bayesian statistics that’s called credence. Weird that they felt the need to invent a new term. | |
| ▲ | doginasuit an hour ago | parent | prev | next [-] | | I like it. It is short and distinct which is a good fit for a primitive. It describes its fundamental meaning and draws a connotation with Boolean. | |
| ▲ | k__ an hour ago | parent | prev | next [-] | | The whole "no hallucinations" premise is based on that. Like, yeah, you don't hallucinate, but only because you force the user to decide in the end. | | |
| ▲ | kjs3 40 minutes ago | parent | next [-] | | force the user to decide in the end And that's...bad? | |
| ▲ | doginasuit 43 minutes ago | parent | prev | next [-] | | That seems like the only possible way to eliminate hallucination, short of a model that is never wrong. | |
| ▲ | rusk an hour ago | parent | prev [-] | | Wait til you hear about how digital circuits work at die level |
| |
| ▲ | keepitwiel 2 hours ago | parent | prev | next [-] | | Bernoulli | |
| ▲ | user3939382 2 hours ago | parent | prev [-] | | If you want to get super pedantic about what’s happening in a transistor every digital Boolean is actually this | | |
| ▲ | kevindamm an hour ago | parent [-] | | Not quite.. that boolean is about whether the voltage exceeds some threshold. It's not about how close the voltage is to the circuit's maximum possible threshold, or how much it exceeds the threshold. In an analog circuit, maybe. |
|
|
|
| ▲ | raverbashing an hour ago | parent | prev | next [-] |
| Jeeves, that's a name I haven't heard in a long time... |
| |
| ▲ | gizajob an hour ago | parent | next [-] | | Personally I’m happy that after a 30 year effort and hundreds of billions spent, AskJeeves finally works as intended. | |
| ▲ | fishfasell an hour ago | parent | prev | next [-] | | If Jeeves returned as an AI chat bot it would be the most brilliant resurgence of nostalgia | | |
| ▲ | grokkedit an hour ago | parent [-] | | jeeves is currently the name of my local hosted assistant, in its context there are rules that tell it to behave like good old jeeves. soon I'll make sure that my home assistant pod answers to "Hey jeeves" |
| |
| ▲ | kjs3 42 minutes ago | parent | prev | next [-] | | We locked him in the basement with Clippy, Bob and BonziBuddy. Who opened the damn basement door??? | |
| ▲ | lherron 15 minutes ago | parent | prev [-] | | …a long time. |
|
|
| ▲ | phplovesong an hour ago | parent | prev [-] |
| So "askjeeves" has been resurrected? |