Remix.run Logo
▲ armcat 3 hours ago

It's weird because two days after Jev was released there were a dozen decision models, a week later there are several dozen, mostly open source, OpenAI's own Decisions API [1] beats it, and you can easily finetune your own [2]. But as others have pointed out, this doesn't matter.

EDIT: As I wrote this Microsoft just released their own Decision-1 model [3].

[1] https://developers.openai.com/api/docs/guides/decisions

[2] https://unsloth.ai/docs/basics/train-your-own-decision-model...

[3] https://commandline.microsoft.com/microsoft-decision-1-model...

▲nico a minute ago | parent | next [-]

Yup, I also released an open source classifiers tool, Jeffy. It comes with 68 pre trained classifiers which run and train on CPU alone. They run locally and are faster than Jev/Laya/Decisions. And they can do things like label email, all the way to even playing Doom

* https://jeffyclassify.com/

* https://playground.jeffyclassify.com/#doom

* https://github.com/nicobrenner/jeffy

▲827a 2 hours ago | parent | prev | next [-]

Decision models have the potential to have an even larger impact on the Real World than LLMs have to this point (which is obviously quite large). But the model itself matters less than the product experiences you build around the model, and its very likely that the incumbent labs are treating the area as something more like "oh yeah I guess we can ship that and then forget about it" rather than investing in what building business processes on decision models looks like. Unlike full language models, I don't think the primary business of Typesafe will be serving Jev at API pricing; it'll look a lot more like putting Jev at the center of a much more expensive suite of software.

There's the potential for an inverse LLM play. In contrast with LLMs, all that seems to matter is the model, and the products the labs build around the models are all really samey and boring; the same left panel list of agents, main view agent conversation, right hand extra context, and we're now in the era of everyone creating the same cutesey furry friend on top of all this tech.

▲pantelisk 2 hours ago | parent | next [-]

Yes, the best way to think of a general classifier like this is like a smart switch statement. Essentially a "JEV" like thing becomes a sort of programming primitive. Once you see it, it's hard to not get excited.

But even if others surpass them and make better solutions, the fact that nobody was able to see it before typesafe is a testament of what they might be able to come up with next.

I sound like a fanboy but I swear I 'm unaffiliated with typesafe. I was building my own version of this way before they announced JEV (mine was ALE and it was mentioned here on HN for a bit), in use for VR gaming (so one can give commands to NPCs with voice and supports multiple commands in sequence in a single pass), but I missed the "killer usecase" of being a new primitive, like everyone else.

TLDR: There's a lot of value in thinking ahead and seeing the future. The clones are nice and exciting but they give me "I could have built this first, yes but you didn't" vibes. I hope they manage to keep it up and push the space forward again

▲user43928 8 minutes ago | parent | next [-]

Honestly, I don't feel the least bit of excitement here and I'm normally enthusiastic about AI.

Yes, this generates probabilities over a set of given options instead of the whole token vocabulary.

I don't see what's so exciting about it, compared to regular LLMs which support structured output options in the API.

▲overfeed 43 minutes ago | parent | prev | next [-]

> [...]the fact that nobody was able to see it before typesafe is a testament of what they might be able to come up with next.

Counterpoint: there are a lot of one-hit wonders, and they vastly outnumber the idea-factory people. This is not to minimize those people, a single idea can be very successful (see Zuckerberg), but it doesn't mean your subsequent ideas will also be great (see Zuckerberg)

▲aranelsurion 2 hours ago | parent | prev | next [-]

I remember your blog post! Thanks for writing it, was pretty cool and a practical application.

Here if anyone is interested: https://pantel.is/projects/ai-gaming-companion/

▲skeeter2020 17 minutes ago | parent | prev [-]

>> TLDR: There's a lot of value in thinking ahead and seeing the future. The clones are nice and exciting but they give me "I could have built this first, yes but you didn't" vibes. I hope they manage to keep it up and push the space forward again

This all sounds intelligent and likely, and yet we can come up with countless counter examples where the first mover is not the big winner, and nobody cares about who did it first. There is typically way more value in nailing the execution of a big idea someone else came up with, rather than "seeing the future".

▲nlpnerd 2 hours ago | parent | prev | next [-]

Yeah, agreed. A model or primitive on its own has no moat and frankly limited value. The paradigm behind "System One" models on the other hand is potentially huge.

https://seldon-ai.com/blog/fronter-llms-are-semantic-interpr...

▲TeMPOraL an hour ago | parent [-]

Skimming the post, it seems to argue for reconstructing the very rigidity that LLMs let us escape, and that very aspect of LLMs is what made them useful and explode in popularity so much.

▲porridgeraisin 2 hours ago | parent | prev [-]

> Unlike full language models, I don't think the primary business of Typesafe will be serving Jev at API pricing; it'll look a lot more like putting Jev at the center of a much more expensive suite of software.

Precisely this. Should be top comment.

Also, the model moat is understated as training data for these purposes also accrues to the winner, which due to the first mover advantage as well as the distribution advantage you speak of, is typesafe. In contrast to relatively open coding data. Openai anthropic also have that, but like you say its a different business.

▲outofpaper 25 minutes ago | parent [-]

The are still just 1tok output of pretty standard llms just along with the logprobs converted to some json

▲cheesecakegood 18 minutes ago | parent [-]

At least in theory (TypeSafe has been pretty close-lipped about the details so this might just be hot air, and I think the evidence is a bit spotty) this is false, since they use a different reinforcement training method.

If the word “calibration” in probability doesn’t mean anything to you, the difference isn’t very apparent, but that doesn’t mean it doesn’t exist.

▲ttul 3 hours ago | parent | prev | next [-]

But Jev established the branding and investors are betting that Jev will be acquired by one of the big labs soon - and if they aren't, the money itself can create a positive outcome by allowing Jev to hire incredible talent and scale the company rapidly.

▲Oras 2 hours ago | parent | next [-]

Rapidly? Their 2 years of stealth was replicated in 2 weeks.

I give it to them for creating the hype (good marketing), and for making a useful classifier. Not sure what they would scale rapidly though.

▲clickety_clack 21 minutes ago | parent | next [-]

It’s like openclaw. There was a bunch of technically better ones that came along afterwards, but nobody remembers what any of them were called.

▲jeremyjh 4 minutes ago | parent [-]

You say this, while Hermes Agent has been at top of the openrouter.ai leaderboard for several months and currently has 3X the token usage of OpenClaw.

▲rsalus 2 hours ago | parent | prev | next [-]

I feel like half the game now is marketing though, so I can see why they'd be attractive to an investor. Maybe if they scale they can come up with something.

▲tiborsaas an hour ago | parent | prev | next [-]

You can replicate any fitness app in no time and you will make close to $0. Brand recognition matters a lot.

▲throwaway7783 an hour ago | parent [-]

100%. Product finesse + Marketing is now the "moat".

▲ModernMech 2 hours ago | parent | prev | next [-]

It’s the classic SV flip. You scale your investors, your executive team, your sales people, hire a bunch of engineering you don’t need, then sell the company. The company’s product doesn’t matter, the company is the product.

▲brink 2 hours ago | parent | prev | next [-]

Either the investors know something we don't, or the market is irrational.

▲ 2 hours ago | parent [-]
[deleted]
▲mococa 2 hours ago | parent | prev [-]

> Their 2 years of stealth was replicated in 2 weeks.

Actually they stolen the idea from a paper.

▲shdh an hour ago | parent [-]

So did Oracle with relational databases by that logic

▲qsod 2 hours ago | parent | prev [-]

[dead]

▲dkersten 3 hours ago | parent | prev | next [-]

Most of them appear to be small LLM’s fine tuned for the role.

That’s a different set of properties in terms of size, cost, and latency. Jev (apparently, not like I’ve seen its insides) is extremely cheap, extremely fast, doesn’t cost any output tokens as it speaks the output natively, can’t get the output wrong because it speaks the format natively, and (presumably based on the docs), the context is separate from the question, meaning it should be immune (or at least highly resistant) to prompt injection attacks.

It’s not just about the accuracy of the result, it’s a collection of all the properties that make Jev interesting.

Jev took years to develop, I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties. Even if fine tuned LLMs can outperform it on raw accuracy.

▲devin 3 hours ago | parent | next [-]

Jev did not take years to develop. What it does was published in arxiv back in 2025. TypeSafe just marketed it.

▲baobabKoodaa 2 hours ago | parent [-]

What specific arxiv paper are you referencing here?

▲homarp 2 hours ago | parent [-]

https://arxiv.org/abs/2503.23303

and https://arxiv.org/abs/2510.01237

▲baobabKoodaa 2 hours ago | parent | next [-]

Yeah, I figured it was gonna be this one.

"SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization"

Jev is a general-purpose thing. That is a specific-purpose thing. General-purpose thing is not the same as specific-purpose thing. What makes people think these are the same thing? I don't get it.

▲devin 2 hours ago | parent [-]

What do you mean? Jev is trivially different from what is described in this paper.

▲baobabKoodaa 2 hours ago | parent [-]

Bullshit. Below is copypaste from the paper in the section that outlines the "key contributions" of the paper. As you can see, it is focused on one specific problem: predicting sales conversions. So if you were to take this system and use it for some other task ("evaluate customer mood" for example), it would not work. Because, again, it is not describing a general purpose solution. It is describing a solution that is specific to one problem: sales conversions.

Copypasta:

• A reinforcement learning architecture specifically designed for sales conversation analysis and conversion prediction

• A synthetic data generation pipeline leveraging GPT-4O to create diverse and realistic sales conversations

• Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific features

• A meta-learning approach enabling the system to express confidence in its predictions based on conversation similarity to training data

• Integration mechanisms providing real-time guidance within existing sales platforms

• Extensive comparative evaluation demonstrating significant performance improvements over LLM-based approaches

▲rpdillon 6 minutes ago | parent [-]

> Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific features

Jev is basically the embeddings side of an LLM. Yes, it's a good idea, but the moat is non-existent.

▲zwaps 2 hours ago | parent | prev [-]

This is a fine-tuned model. The author even states that the model is competitive with Jev only if fine-tuned on the evaluation at hand.

Literally misses the point of Jev, which you don't need to fine-tune to get accuracy nor - and no other model has this - some sort of out of sample calibration

▲TeMPOraL an hour ago | parent | prev | next [-]

Jev is something your favorite LLM could zero-shot months ago, if you pointed it to the right arXiv paper (some of which are linked in this thread).

▲mrinterweb 15 minutes ago | parent [-]

That probably explains why there were so many competitors around withing days of the Jev announcement. They are not starting with a moat, and there doesn't seem to be any moat in sight. Just buzzword recognition because everything is comparing to "jev".

▲hbrn 2 hours ago | parent | prev [-]

> I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties

But why? If the simplest way to achieve Jev's capabilities (accuracy, cost, latency) is by fine tuning a small model, what makes you think that this isn't exactly what Typesafe did?

And even if they did something different - what makes you think it was a good idea in the first place, given how easy their results were replicated without any "secret sauce"?

▲dkersten an hour ago | parent [-]

My point is that it’s not replicated. You replicate the accuracy, but not the other properties. The Jev-competitors only proved that you can get or beat the accuracy, nothing about the other properties. Especially the “zero hallucination” output and the (if it works how the documentation make it sound) prompt injection resistant architecture. You can’t get that with a fine tuned LLM.

▲hbrn an hour ago | parent [-]

> zero hallucination

Plenty has been said about this claim. If you're still falling for this, I feel sorry for you.

If you remove wheels from your car, your car will get a "no speeding ticket" property, and yet there's nothing exciting about it.

> You can’t get that with a fine tuned LLM

Of course you can. All these claims are nothing but marketing.

▲baobabKoodaa 3 hours ago | parent | prev | next [-]

> you can easily finetune your own

no, you can't, and it's unclear why you would think this.

▲ricericerice 2 hours ago | parent [-]

you can easily finetune your own*

*if you have a sufficiently sized and quality dataset for the specific classifications you're targeting

▲baobabKoodaa 2 hours ago | parent [-]

And even if you do have that, you haven't made your own Jev, because Jev is a general-purpose thing, whereas what you have built is a specific-purpose thing.

▲girvo an hour ago | parent | prev | next [-]

Counterpoint: my work has already allowed us to call and test Jev. Those others? Who knows when, if ever.

▲jgilias an hour ago | parent | prev | next [-]

Isn’t the OpenAI decisions API basically just Luna cosplaying a decisions model and pretending the confidence score isn’t just a hallucination?

▲hbrn 41 minutes ago | parent | next [-]

And what do you think Jev confidence score is?

Here's a hint: confidence is not generated by a model.

▲jgilias 12 minutes ago | parent [-]

Thanks, fixed my understanding!

Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs?

▲phalangion an hour ago | parent | prev [-]

What’s the difference?

▲scottyah 3 hours ago | parent | prev | next [-]

> OpenAI's own Decisions API [1] beats it

Have you heard that from a different source than OpenAI? From what I'd heard other models haven't gotten close, and the open source ones are like running gemma4 E2B against Opus 5.5- sure, the API calls go in and are returned the same but the quality isn't close.

▲throwaw12 2 hours ago | parent | prev | next [-]

You are right in terms of how fast competition created alternatives.

But, for OpenAI this is not a primary business, for open source models as well, so they will not be chasing the market and customers to buy their product and promise them to maintain it.

TypeSafe will do all this, they will try to understand your use cases and then solve your pain point, while others are providing raw material.

▲bushbaba 2 hours ago | parent | prev | next [-]

A major VC could type safe ai money, then head to a larger AI company looking to raise their series E+ and demand they acquire typesafe as part of their funding allotment.

such an arrangement can end up beneficial to the VC firm

▲mlmonkey 3 hours ago | parent | prev | next [-]

OpenAI's "Decisions" library has this in requirements:

To run the SDK examples below, use these OpenAI SDK versions or later: Python 3.26.0,

I thought Pythin 3.15.0 just came out, 3.26.0 must be really far off?

▲bayesianbot 3 hours ago | parent | next [-]

That is their Python SDK version, not Python version

▲ 3 hours ago | parent | prev [-]
[deleted]
▲ 2 hours ago | parent | prev | next [-]
[deleted]
▲user3939382 an hour ago | parent | prev | next [-]

Investments aren’t made because the product is amazing, they’re made because there’s a compelling exit scenario. Engineers don’t want to hear this but more generally, the critical success factors for a business aren’t product or engineering they’re relationships i.e. sales and team dynamics. If technical excellence dictated business outcomes in tech Salesforce wouldn’t exist for example.

▲hkalbasi 2 hours ago | parent | prev | next [-]

> OpenAI's own Decisions API [1] beats it

Jev is 42$/B but OpenAI is 100$/B token.

▲amelius 2 hours ago | parent | prev | next [-]

I mean I'm already ditching my Apple stocks because soon AI will be able to replicate iOS and MacOS.

▲doctorpangloss 2 hours ago | parent | prev | next [-]

"It doesn't matter"

By all means, become an A16Z LP.

▲nlpnerd 3 hours ago | parent | prev | next [-]

You are assuming that the VCs have done their due diligence. For a "hot" company like Typesafe AI, most likely little due diligence was done. That's the way it's played.

▲moralestapia 2 hours ago | parent | prev [-]

Nothing beats nepo, brother.