Remix.run Logo
▲ jorblumesea 6 hours ago

This is literally the plan, open weight models are something like 60% of token spend, and it will get worse. many companies now have model gateways where you can slot in cheaper models via cli for cheaper. we've been using glm 5.x and it's pretty close to SOTA frontier models.

it's also why there have been so many calls for regulation and slowdowns.

▲LeBit 5 hours ago | parent | next [-]

Yup.

I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.

I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.

▲nozzlegear 5 hours ago | parent [-]

Exactly what I've been doing. I don't need the all-powerful GPT-6 Math Scoopa, or Opus T-1000, just to write react, svelte and C# for me; my local Qwen3.8 is more than capable, and I can switch to Deepseek and GLM on OpenRouter when I need speed. I just pop in to read the comments on HN for the latest drama and navel gazing, then I click the Hide button and move on. Couldn't give a wooden nickel what their latest and greatest models are capable of anymore, it's just PR buzz.

▲0cf8612b2e1e 5 hours ago | parent | prev [-]

There is already tooling to automatically pick models within an organization. Eventually it could be as easy as flipping a switch in group policy that forces everyone to switch to the cheaper models.

Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.