| ▲ | sachaa 2 hours ago | |||||||
This is a great way to start but the perf of LLM models like Qwen are not ideal for local execution. I have adapted Laya (pure decision model) to run in a browser and I am able to get responses under 200ms. Give it a try: https://wexare-ai.github.io/browser-laya/ | ||||||||
| ▲ | Tostino an hour ago | parent | next [-] | |||||||
Why would a Qwen 1.7b be too slow? You should get well under 200ms there. | ||||||||
| ▲ | make3 an hour ago | parent | prev [-] | |||||||
Laya is just ModernBert fine-tuned. It's still a language model (just a bidirectional one). Let's call it what it is instead of this weird Decision Model mysticism. | ||||||||
| ||||||||