Remix.run Logo
nsoonhui 14 hours ago

I did try to use Chinese open models, but for my production work they simply couldn't cope at all; both GLM 5.3 and Deepseek v4 went into infinite loop and wasted my tokens until my OpenRouter wallet reached 0; good thing I didn't enable the auto topup. US models, by contrast, breezed past them.

Even for simpler tasks, Chinese models took long time to complete, and I needed to supervise closely. The price , in the end, didn't come cheap, mainly because too much time wasted on thinking.

So maybe one day Chinese models will squeeze out the American ones, but today is not that day.

As far as consumers are concerned, I feel blindly shilling for anyone purely for ideological reasons are quite meaningless, especially when it comes to open/close source and US/China rivalry. I have no obligation to support "open source/weight" or the "underdogs" just because they are so. We only want things that work, and at a cheap price.

seanmcdirmid 14 hours ago | parent | next [-]

I’ve been using deepseek and it works great for my problems. An expensive day is when I spend $7 in tokens, and that takes lots of queries. Openrouter doesn’t give you the cache discount I think, which is really important.

irthomasthomas 14 hours ago | parent | next [-]

Why on earth would you use openrouter for this? The cache discount for deepseek is the highest by far, it is the cache that makes the official API so cheap, even after the recent price rise.

hiq 9 hours ago | parent [-]

I didn't know about this limitation from openrouter, and I thought plenty of users were using it to quickly switch as needed without incurring such a penalty. Why does there seem to be so many users then? Are some folks just fine paying multiples of what they could?

irthomasthomas 8 hours ago | parent [-]

I really don't know. It was built before prompt caching was common, and the switching cost was much lower.

nsoonhui 13 hours ago | parent | prev [-]

That's strange. For Deepseek I burnt through USD 5 on a relatively simple task, in one afternoon. That the simple task took a whole afternoon, a lot of baby sitting, the slowness, and so much money relatively, really shook me to the core.

seanmcdirmid 5 hours ago | parent [-]

What was the tax? You were using v4 flash right? What thinking setting were you using?

I’m pretty happy with it, it’s about as good as Gemini flash 3.5, the only other model I have deep experience with.

weiran 9 hours ago | parent | prev [-]

The only way the Chinese models financially make sense is if you use their subs or host it yourself. DSV4 is especially token hungry.