| ▲ | tbeseda an hour ago | |
For _your_ classification it's unacceptable. The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that? I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars. Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know. | ||
| ▲ | nico 34 minutes ago | parent | next [-] | |
Not sure the task at hand here. But if it doesn’t require any reasoning/thinking and it’s just a classification task, it’s worth a shot to look into training your own classifier I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes The classifiers also run in <1ms, so they can be very fast and precise at the same time But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one) For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data | ||
| ▲ | clhodapp 39 minutes ago | parent | prev [-] | |
Needing fine-tuning for the use-case completely changes the product category | ||