Remix.run Logo
▲ sachaa 2 hours ago

This is a great way to start but the perf of LLM models like Qwen are not ideal for local execution. I have adapted Laya (pure decision model) to run in a browser and I am able to get responses under 200ms. Give it a try: https://wexare-ai.github.io/browser-laya/

▲Tostino an hour ago | parent | next [-]

Why would a Qwen 1.7b be too slow? You should get well under 200ms there.

▲make3 an hour ago | parent | prev [-]

Laya is just ModernBert fine-tuned. It's still a language model (just a bidirectional one). Let's call it what it is instead of this weird Decision Model mysticism.

▲arendtio 6 minutes ago | parent [-]

I am also not a big fan of everybody calling them system one models. AFAIK, system one also refers to activities like driving a car, but I don't want to have those decision models driving cars.

So I think I understand what is meant, but I don't like the comparison.