Remix.run Logo
▲ Roark66 43 minutes ago

You know what, considering there was a recent "small" open weights LLM released recently that meets 90% of my coding needs I'm inclined to agree.

Qwen3.8-Flash-Next - relatively small, it runs on 6 6 year old GPUs on my home PC happily running 5 simultaneous 262k sessions with additional 10 cached in RAM (bought back when you didn't have to remortgage your house for Ram) and it has been the first local model that is not a toy.

But there is a class of problems where I still reach for Anthropic's fable...

However, I have a hunch bordering with certainty Anthropic is achieving such great results by doing a lot of harness tricks.

For example opus 4.8, is not much better on coding than before mentioned Qwen model, but gets amazing results on factual knowledge stuff (the knowing all works of Shakespeare thing). How hard would it be to add a general knowledge RAG to requests that contain relevant questions and beat all benchmarks like that? Not very hard.

So I think there is big innovation to be had in harnesses, routers, inference and so on.

As to money spent on AI per developer my current client (a fortune 200 software company) spends $500 per month. That is $6k a year. A lot more than your examples. And many people run out of their quota pretty quickly.