Remix.run Logo
user43928 2 hours ago

OpenAI decreased prices with the 5.6 model family.

And later they further cut Sol and Terra pricing by 20% (maybe only in the API) and Luna by 80%.

In fact Luna still outperformed DeepSeek Flash 4.1 in cost per task on Artificial Analysis when I last checked.

However, Luna is slightly less intelligent. I have a feeling that it's pretty dumb and prone to hallucination unless running at xhigh or max effort, where it somehow manages to work quite well.

I did not personally test the open weight models beyond the old Qwen 3.6 27B, which produced unusably bad results for me.

The competition is great, and I hope Chinese models will continue to force leading US labs to offer models at a low price point.

That said, I don't think the Chinese labs have anything over OpenAI and Anthropic when it comes to capability or efficiency - I have no reason not to believe the US labs have even lower cost to serve the models.

tacomagick 2 hours ago | parent | next [-]

OpenAI had to cut costs because of Anthropic. I also do not trust the benchmarks when it comes to models anymore. I have tried both Claude and OpenAI models and while it is true that the 5.6 series is smarter than Deepseek (at the time i tested it against 4.0) at that price it is still not worth it and sometimes randomly refuses to do tasks or stops midway etc.

Do also remember China is this far in the AI race despite all chip restrictions from America. If they were in equal standards I truly think Chinese models would have long surpassed American ones. Also would like to remind how Anthropic CEO is being hostile and blaming Chinese models with distilling meanwhile their own models claimed to be Qwen¹ and their stance against open models is negative² and they still keep blaming China for it.

1- https://news.ycombinator.com/item?id=48671252

2-https://www.anthropic.com/news/position-open-weights-models

goosejuice 2 hours ago | parent | next [-]

> Also would like to remind how Anthropic CEO is being hostile and blaming Chinese models with distilling

Why wouldn't he? If there really was 25,000 accounts breaking ToS any CEO would at minimum be upset. Evidence of Claude distilling qwen would be damning but that a) makes no sense b) doesn't exist afaik.

user43928 2 hours ago | parent | prev [-]

Not sure about that.

Given the difference in compute, it seems plausible.

However, the researchers at the US labs are surely no less talented, and they have better access to hire talent globally.

They too have to serve their models efficiently at a large scale, and with current capacity constraints this must be a top priority.

Implicated an hour ago | parent | prev [-]

> I did not personally test the open weight models beyond the old Qwen 3.6 27B, which produced unusably bad results for me.

So you don't have much perspective on things, it seems. Let me introduce you to the GLM 5.2 and then 5.3/5.3 flash series of... "oh, wow, I should have bought some RTX PRO 6000's while they were 'cheap'" stage of progression.

As someone carrying multiple max subscriptions to both claude and codex - primary workhorse is glm 5.3 flash running on rented GPUs for less than a latte/hr.

I also found qwen 3.6 27B nearly useless for my own needs. DS4 flash 0731 and then 4.1 have been nearly as eye opening as glm 5.3 flash, but have their own warts.

user43928 2 minutes ago | parent | next [-]

Why use GLM 5.3 Flash when you also have access to Astra, Sol, Fable?

Or I guess the other way around, if GLM 5.3 Flash is so good, why Claude and Codex?

CamperBob2 a minute ago | parent | prev [-]

Try DS4.1 Flash. It's another eye-opener. If you run it in Claude Code, it's easy to forget you're not actually talking to a high-end Opus model.