| ▲ | dkersten 3 hours ago |
| Most of them appear to be small LLM’s fine tuned for the role. That’s a different set of properties in terms of size, cost, and latency. Jev (apparently, not like I’ve seen its insides) is extremely cheap, extremely fast, doesn’t cost any output tokens as it speaks the output natively, can’t get the output wrong because it speaks the format natively, and (presumably based on the docs), the context is separate from the question, meaning it should be immune (or at least highly resistant) to prompt injection attacks. It’s not just about the accuracy of the result, it’s a collection of all the properties that make Jev interesting. Jev took years to develop, I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties. Even if fine tuned LLMs can outperform it on raw accuracy. |
|
| ▲ | devin 3 hours ago | parent | next [-] |
| Jev did not take years to develop. What it does was published in arxiv back in 2025. TypeSafe just marketed it. |
| |
| ▲ | baobabKoodaa 2 hours ago | parent [-] | | What specific arxiv paper are you referencing here? | | |
| ▲ | homarp 2 hours ago | parent [-] | | https://arxiv.org/abs/2503.23303 and https://arxiv.org/abs/2510.01237 | | |
| ▲ | baobabKoodaa 2 hours ago | parent | next [-] | | Yeah, I figured it was gonna be this one. "SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization" Jev is a general-purpose thing. That is a specific-purpose thing. General-purpose thing is not the same as specific-purpose thing. What makes people think these are the same thing? I don't get it. | | |
| ▲ | devin 2 hours ago | parent [-] | | What do you mean? Jev is trivially different from what is described in this paper. | | |
| ▲ | baobabKoodaa 2 hours ago | parent [-] | | Bullshit. Below is copypaste from the paper in the section that outlines the "key contributions" of the paper. As you can see, it is focused on one specific problem: predicting sales conversions. So if you were to take this system and use it for some other task ("evaluate customer mood" for example), it would not work. Because, again, it is not describing a general purpose solution. It is describing a solution that is specific to one problem: sales conversions. Copypasta: • A reinforcement learning architecture specifically designed for sales conversation analysis and conversion
prediction • A synthetic data generation pipeline leveraging GPT-4O to create diverse and realistic sales conversations • Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific
features • A meta-learning approach enabling the system to express confidence in its predictions based on conversation similarity to training data • Integration mechanisms providing real-time guidance within existing sales platforms • Extensive comparative evaluation demonstrating significant performance improvements over LLM-based approaches | | |
| ▲ | rpdillon 6 minutes ago | parent [-] | | > Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific features Jev is basically the embeddings side of an LLM. Yes, it's a good idea, but the moat is non-existent. |
|
|
| |
| ▲ | zwaps 2 hours ago | parent | prev [-] | | This is a fine-tuned model. The author even states that the model is competitive with Jev only if fine-tuned on the evaluation at hand. Literally misses the point of Jev, which you don't need to fine-tune to get accuracy nor - and no other model has this - some sort of out of sample calibration |
|
|
|
|
| ▲ | TeMPOraL an hour ago | parent | prev | next [-] |
| Jev is something your favorite LLM could zero-shot months ago, if you pointed it to the right arXiv paper (some of which are linked in this thread). |
| |
| ▲ | mrinterweb 14 minutes ago | parent [-] | | That probably explains why there were so many competitors around withing days of the Jev announcement. They are not starting with a moat, and there doesn't seem to be any moat in sight. Just buzzword recognition because everything is comparing to "jev". |
|
|
| ▲ | hbrn 2 hours ago | parent | prev [-] |
| > I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties But why? If the simplest way to achieve Jev's capabilities (accuracy, cost, latency) is by fine tuning a small model, what makes you think that this isn't exactly what Typesafe did? And even if they did something different - what makes you think it was a good idea in the first place, given how easy their results were replicated without any "secret sauce"? |
| |
| ▲ | dkersten an hour ago | parent [-] | | My point is that it’s not replicated. You replicate the accuracy, but not the other properties. The Jev-competitors only proved that you can get or beat the accuracy, nothing about the other properties. Especially the “zero hallucination” output and the (if it works how the documentation make it sound) prompt injection resistant architecture. You can’t get that with a fine tuned LLM. | | |
| ▲ | hbrn an hour ago | parent [-] | | > zero hallucination Plenty has been said about this claim. If you're still falling for this, I feel sorry for you. If you remove wheels from your car, your car will get a "no speeding ticket" property, and yet there's nothing exciting about it. > You can’t get that with a fine tuned LLM Of course you can. All these claims are nothing but marketing. |
|
|