| ▲ | Build your own decision model(nishtahir.com) |
| 207 points by softwaredoug 7 hours ago | 46 comments |
| |
|
| ▲ | kekebo 7 minutes ago | parent | next [-] |
| Just for the author: On any of my iOS 26 browsers (orion, brave, safari), a page reload occurs when the model download completes, which resets the state, so I never get to interact with the model. |
| |
|
| ▲ | sachaa an hour ago | parent | prev | next [-] |
| This is a great way to start but the perf of LLM models like Qwen are not ideal for local execution. I have adapted Laya (pure decision model) to run in a browser and I am able to get responses under 200ms. Give it a try: https://wexare-ai.github.io/browser-laya/ |
| |
| ▲ | make3 33 minutes ago | parent [-] | | Laya is just ModernBert fine-tuned. It's still a language model (just a bidirectional one). Let's call it what it is instead of this weird Decision Model mysticism. |
|
|
| ▲ | sn0n 4 hours ago | parent | prev | next [-] |
| I saw Jev and without understanding it said to myself, “I can build that!” And brainstormed some weird Alex trebek jeopardy generator I called trebek, a bun ran typescript app that you give it an input, it decides if it’s a category, question or answer then generates what’s missing. Trained a SQLite-vec database on the English language for a few days with a small qwen embedding model to add vec embeds to the db then ripped the cord on the embeddings before I started feeding my llm the proper specs for Jev and now i have a cool jeopardy generator that now doubles as a Jev clone classifier with a custom v1 endpoint for system 0 or whatever it is. Level understanding here, fun experiment though ^>^ and produces useful outputs |
| |
|
| ▲ | mehar_pro 27 minutes ago | parent | prev | next [-] |
| What if Jev is something more than a classifier? There's so much potential here waiting to be seen. I built my own Jev, then realized it could literally run on a webpage. So then I turned it into an npm library https://genclass.dev |
|
| ▲ | nico 5 hours ago | parent | prev | next [-] |
| This is very cool. If you are looking for something similar but more lightweight, that you can run (and train) on CPU, try out Jeffy: https://jeffyclassify.com/ On GitHub: https://github.com/nicobrenner/jeffy |
| |
| ▲ | mrkn1 5 hours ago | parent [-] | | Cool project too. If you are looking for a 500MB instead of gigabytes, with evals on Jevbench that you can run fast on CPU check out gutsy. https://github.com/kouhxp/gutsy | | |
| ▲ | nico 4 hours ago | parent [-] | | Very cool, thank you for sharing The banking77 numbers called my attention. Using a local classifier you can get 94%+ accuracy: https://playground.jeffyclassify.com/#model/banking77 I think Jev-like models are amazing for exploration and finding the right workflows, but the moment you have fixed classification tasks, it’s often more efficient to use an adhoc classifier, which you can quickly and easily train on CPU with not that much data (you can get an email classifier to 95% accuracy/f1 with 50-100 emails) Edit: Would love to somehow mix both approaches automatically and have a general model which can take novel tasks, but then switch to a classifier after it gets enough data for training an adhoc model |
|
|
|
| ▲ | fuddle 2 hours ago | parent | prev | next [-] |
| I'm looking forward to a "Build a Decision Model (From Scratch)" book. |
| |
| ▲ | make3 28 minutes ago | parent [-] | | Take ModernBert, add a fully connected layer with 255 outputs, take a bunch of classification datasets from Huggingface, write the code to have the datasets fit the jev format on these 255 outputs, do supervised fine tuning on the datasets with that format. then use a confidence loss of some kind |
|
|
| ▲ | ford 5 hours ago | parent | prev | next [-] |
| These are neat - and the source of the many Jev clones we've seen. I think their recent funding round is in part because of their algorithms/data. It remains to be seen if that's a big enough edge to be worth 1 billion+ dollars |
|
| ▲ | sva_ 5 hours ago | parent | prev | next [-] |
| Does someone have examples of interesting stuff that has been built utilizing Jev/decision models? The way this is hyped up surely there must be some good stuff? |
| |
| ▲ | yawnxyz 23 minutes ago | parent | next [-] | | use it to classify a bunch of research papers (by feeding it section by section, or summaries of sections if too long) works like a charm (not a product, so not much to share) | |
| ▲ | zeroq 4 hours ago | parent | prev | next [-] | | I see Jev as a major step towards commoditizing current LLM paradigm. One thing would be to further optimize this particular route to work purely on CPU. This will grant an option to embed this feature into any application, from MS Office to games. The other is integrating this into agent workflow to vastly minimize token consumption. | | |
| ▲ | Lukas_Skywalker 3 minutes ago | parent [-] | | So, what exactly is the > interesting stuff that has been built utilizing Jev/decision models in that case? |
| |
| ▲ | JLO64 4 hours ago | parent | prev | next [-] | | Not in a serious manner but I created a testing harness for a Nintendo 3DS game I'm making that uses the OpenAI Decisions API. The main advantage is the speed (~2-300ms per input) which I really need for this purpose. | |
| ▲ | JKCalhoun 3 hours ago | parent | prev | next [-] | | Also like to see a "layman harness" like LM Studio integrate an open decision model in its workflow. | |
| ▲ | nico 5 hours ago | parent | prev | next [-] | | Not with Jev, but you can use classifiers for a lot of use cases, here’s a few: https://playground.jeffyclassify.com/ | |
| ▲ | UltraSane 4 hours ago | parent | prev [-] | | I use Jev in a Claude code hooks to detect dangerous commands. |
|
|
| ▲ | demibabs 6 hours ago | parent | prev | next [-] |
| Is simply changing the temperature so that the model appears calibrated over a particular benchmark after the fact “allowed”? Feels p-hacking esque. |
| |
| ▲ | JMKH42 5 hours ago | parent [-] | | if it works it works!
as long as the test set is reasonably large and diverse its better than nothing. You could characterize how robust it is by throwing dozens of different types of work at it and see how much the confidence varies |
|
|
| ▲ | howunfortunate 7 hours ago | parent | prev | next [-] |
| As an MLE who has been failing to get anyone interested in classifiers for many years, the hype around Jev makes me scream internally. Yes, I get that a zero-shot classifier is more convenient than the traditional kind, it's very cool. Kind of. But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming. |
| |
| ▲ | firasd 6 hours ago | parent | next [-] | | I think part of what made Jev catch on is that the API is like an if(...) or switch statement People are just so used to the chat style APIs that they didn't even consider doing things like sending a bunch of emojis to a chat model and then asking for the optimal one in this context etc. Also chat models are pricier for the same behavior and can also output something random like a refusal But yeah ironically I think in the initial breakthrough LLM paper on GPT-3 in 2020 some of the multiple choice questions were answered by comparing token probabilities of specific continuations rather than fill in the blank | | |
| ▲ | howunfortunate 6 hours ago | parent [-] | | You know, it's a good point. Maybe I should have spent less time pitching to PMs and more time pitching ground up to devs, who have the right foundation to intuitively understand the usefulness. |
| |
| ▲ | jofzar 6 hours ago | parent | prev | next [-] | | > But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming. It's the scale of "perfectly cheap", jev (specifically) is so dirt cheap and fast that you can throw it at things that should not be justifiable in the past and you barely have to do any work other then quick testing. | | |
| ▲ | jofzar 4 hours ago | parent [-] | | I will also say, not having to get a team to build this for you, having to get business justification from your team for that teams hours, then spending time revising and testing that out and then you have to "prove" that it's worth having in your feature as a AI cost Versus "Let's put jev here and see how it works, if it works, then fantastic let's build a business case" | | |
| |
| ▲ | laybak 5 hours ago | parent | prev | next [-] | | on the bright side, maybe this zero-shot classifier wave could be the back bone for more specialized classifiers (with more mindful selection of data and training) | |
| ▲ | esafak 5 hours ago | parent | prev | next [-] | | Jev is classification for normies. Look ma, no training. | |
| ▲ | copperx 6 hours ago | parent | prev | next [-] | | I want to share the rage. Can you expound on what makes you scream? | | |
| ▲ | howunfortunate 6 hours ago | parent [-] | | Idk, imagine you worked on Skype's B2B sales team for years and then COVID happens and Zoom blows up. Is it a better thing? Yeah. Does it affect me in any tangible way? No. But come on, really people? All you needed was like one tiny bell & whistle to take this from nothing to the hottest thing of all time? | | |
| ▲ | JMKH42 5 hours ago | parent [-] | | I bet the timing was the key, people in the last year have been furiously building things that use LLMs as an API and as we work on this stuff we have systems with N LLM steps and M of them are frustrating because you want a specific choice picked or list of things ranked and sometimes the llm will just output something else entirely! So along comes this thing you can graft in that is more reliable, faster, and cheaper for that, and I get it immediately. A year ago I'd be like "kinda cool but what it for?" |
|
| |
| ▲ | win311fwg 5 hours ago | parent | prev | next [-] | | > But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now I'm no prompt engineer, but I've found them to be too slow any time I wanted to use them that way. I have never tried Jev, but apparently it is supposed to be fast, so it seems like, according to the marketing, it could become usable where LLMs haven't been. | |
| ▲ | BoorishBears 5 hours ago | parent | prev | next [-] | | Maybe instead of screaming you can take this as a chance to level up your engineering. Good engineers don't treat approach as A == B or even A like B, when extremely integral parts of their applications differ. Zero-shot isn't just "more convenient", in a low data regime: it's the only workable solution, and 100x so if your plan involves the acornym "BERT" (because even the largest of those models has the world knowledge of a fart to draw priors from) Better ergonomics while being faster and cheaper as the existing things really is enough to justify callling what you've done a new thing, in a world of finite resources and time. It's actually making me scream how many people don't get that. | | |
| ▲ | howunfortunate 4 hours ago | parent [-] | | > in a low data regime: it's the only workable solution Low data regimes no longer exist in the age of LLMs, one can trivially generate a training and eval set and distill a good classifier on any domain within a day. But I do acknowledge zero-shot is more convenient. Personally I don't think Jev has any moat so I won't bother with their model specifically, but yes I do anticipate using this type of thing more in the future. | | |
| ▲ | BoorishBears 4 hours ago | parent [-] | | I'm going to crash out at the stupidity and earth-shattering banality of that first paragraph if I try to respond faithfully, so I won't and agree to disagree. | | |
| ▲ | howunfortunate 2 hours ago | parent [-] | | I don't think you actually know what the word 'banal' means, which is quite amusing given your condescension. |
|
|
| |
| ▲ | dominotw 5 hours ago | parent | prev [-] | | > As an MLE Yea but those models you were building were lame and inaccessible to play with for common devs. Just because they have the same api doesnt mean you were building the same thing. |
|
|
| ▲ | ursaguild 6 hours ago | parent | prev | next [-] |
| This is really cool to see. Being able to play the token generation was awesome. Amazing job with breaking down how to think about these models. This made the idea of Jev/decision models really easy to grasp for me. The idea of calibrating the model was helpful. I thought this was a great overview. |
|
| ▲ | vivzkestrel an hour ago | parent | prev | next [-] |
| - i have to keep scrolling down on your home page https://nishtahir.com/ to see what posts you have - could you kindly put all that in a /blog page with pagination and not infinite scroll? |
|
| ▲ | bellajbadr 7 hours ago | parent | prev [-] |
| Is this only about getting fix json output? |
| |