Remix.run Logo
▲ Ollaya – Ollama for open-source, Jev-style decision models(ollaya.dev)
107 points by Ardakilic an hour ago | 33 comments
▲george_max an hour ago | parent | next [-]

Has anyone actually seen better or the same results with Laya compared to Jev? From my experience, Laya performs significantly worse. It's less confident and often makes wrong decisions with more complex queries.

▲scronkfinkle an hour ago | parent | next [-]

Yes. JEV generalizes better because they probably have an enormous corpus and trained on it for a long time. Laya's out of the box model is much weaker. However, in the age of LLM's it's incredibly easy and cheap to generate large datasets to fine tune laya for your task, and the training loop is pretty quick and cheap too.

It's so easy that I question why I would ever pay for JEV when eventually I'll have done enough random things that I will also have a large corpus and likely a general model as well.

▲mtkd 20 minutes ago | parent [-]

Isn't the point of Jev that it generalises better?

It's a fast classifier you can use out-the-box, ~1.5bn tokens is about $40 (I've been hammering it)

It just works ... a whole bunch of low-level/low-importance workflow stuff that was getting farmed out to small/fast LLM models now has a competitive alternative ... and bits that hadn't even been considered to go into some external descision/classifier service can be tested/deployed at ~$0.00003/req

I don't get this wall of negativity on it, it's genuinely innovative/useful tech ... would expect HN to be more positive, regardless of whether it's the absolute best execution

▲not_a_bot_4sho 7 minutes ago | parent | next [-]

I didn't see any negativity in the post you replied to.

I think the point being made is that Jev is great but it has no competitive moat, and open source versions will very soon catch up if their secret sauce is just synthetic data.

(Whether or not that is true, I don't know.)

▲shepardrtc 16 minutes ago | parent | prev [-]

It really does just work. And it works so well I already integrated it into my product. Saves me about 75% of costs for the section its working in, which isn't a small amount. I see a lot of negativity and I don't really get it either. Its so cheap and so fast, why not give it a try?

▲jonmagic 16 minutes ago | parent | prev | next [-]

I've been following jevbench twice a day for the past week and that's been a lot of fun. Latest update:

Rank System Score Public / sealed accuracy Evidence

1 decider-4b v2 64.13 83.5% / 34.7% Evaluator-run, offline

2 Jev 1.13 63.29 86.6% / 36.7% Evaluator-run API

3 JevK5 v0.2 62.04 85.3% / 33.1% Evaluator-run

4 Cygnet 12B 61.76 87.9% / 33.8% Evaluator-run, offline

5 Hopper 59.43 82.3% / 34.1% Evaluator-run

28 Kev 4B 36.14 66.2% / 22.4% Evaluator-run

41 Laya 421M 30.25 58.4% / 30.8% Evaluator-run

▲cobanov an hour ago | parent | prev | next [-]

Developer here. You're right, Laya is a lot weaker than Jev, especially on harder queries. It's a small model, so it's fast, but that's the trade-off. The open models that get close to Jev are much bigger, and running those is what I'm working on next.

▲mikodin 26 minutes ago | parent [-]

What are the models? I am super curious in these as well

▲iamflimflam1 31 minutes ago | parent | prev [-]

Nothing yet. Unfortunately it sometimes feels like our industry has been overrun by grifters and chancers.

I’m sure this has been a gradual and long decline. Maybe it even started with the dot com boom and accelerated with crypto. With AI it seems to have got worse.

▲ranyume an hour ago | parent | prev | next [-]

>Run decision models locally.

>example is a text classification task instead of a decision

▲hbrn 35 minutes ago | parent | next [-]

"Decision model" is just marketing jargon.

decision model = classifier

system one model = small non-reasoning LLM

noul = boolean

confidence = f(probabilities)

It's sad to see how gullible engineers are today.

▲OgAstorga 43 minutes ago | parent | prev | next [-]

text classification is equivalente to decision. This is exactly the same thing Jev does.

▲ranyume 36 minutes ago | parent | next [-]

If it has four legs, a tail and barks why not call it a dog?

▲gchamonlive 28 minutes ago | parent [-]

Because this specific dog only barks in structured text

▲ricardobeat 37 minutes ago | parent | prev | next [-]

It is not. In a benchmark with actual decisions - navigation, traffic, waypoints - laya does slightly better than a small classifier, with very low correlation to state changes.

▲abirch 39 minutes ago | parent | prev [-]

Jev does it more efficiently because it doesn't use an LLM https://typesafe.ai/blog/introducing-system-one-models-and-j...

▲cobanov an hour ago | parent | prev [-]

Fair point, that example is basically classification. I'll change it to something that looks more like a real decision.

▲gauravsapkotanp 11 minutes ago | parent | prev | next [-]

I have also tried this and its really awesome

▲mococa 23 minutes ago | parent | prev | next [-]

It would be really cool to have LLMs and System One in a single tool - in this case, if Ollama implemented it.

▲datadrivenangel an hour ago | parent | prev | next [-]

Are there many models that are comparable to Jev for generic decision making?

Smarter move if you have an eval set is to just train a classifier and call it a day.

▲rgbrgb an hour ago | parent | next [-]

there's this thing with a bunch of similar models https://huggingface.co/spaces/multimodalart/jev-decision-ind...

top open one is trained by perplexity cto for $3k, kinda cool https://x.com/denisyarats/status/2102252088067850507

▲physicallyIllfr 37 minutes ago | parent [-]

<<<"i was curious to see if i could train a competitive Jev-like model completely autonomously with a swarm of agents using our internal system."

Bro is writing off the H200 lol

On a sidenote I really can't stand the term "swarm" and definately plays into AI doomerism.

▲cobanov an hour ago | parent | prev [-]

The link rgbrgb posted is a good overview. The best open ones are close to Jev now, but they're big models. And I agree, if you have an eval set for a fixed task, a trained classifier is the better choice.

▲emmettbt an hour ago | parent | prev | next [-]

Cool... but this does seem undermined by the fact that Ollama can add support for decision models at any time.

▲cobanov an hour ago | parent | next [-]

Fair, and I'd be happy if they did. Ollaya uses the same API as Jev, so your code isn't tied to it either way

▲accountrequired an hour ago | parent | prev [-]

and that ollama is go-llama and not rust, so it's not really the ollama of anything

▲handfuloflight an hour ago | parent | prev | next [-]

Sounds good on latency but how is its actual decision quality vs. Jev?

▲cobanov an hour ago | parent [-]

Depends on the model. The small ones I support today are well below Jev on harder queries, but fine for simple, well-defined questions. The open models that get close to Jev are bigger, and I'm adding support for those next.

▲george_max an hour ago | parent | prev | next [-]

I am fairly confident if Jev-style decision models are seen as prominent (which, they seem to be), Ollama will support them. Surprised the team hasn't implemented this already.

▲eserozvataf an hour ago | parent | prev | next [-]

great project for empowering open-source alternatives.

▲rkovashikawa an hour ago | parent | next [-]

open-source is the only way for safe AI development. whoever doesn’t share the weights/code will lag behind.

▲cobanov an hour ago | parent | prev [-]

Thanks!

▲0xbadcafebee 7 minutes ago | parent | prev [-]

[delayed]