| ▲ | EbNar 5 hours ago |
| Since a few months, I almost exclusively use the Chinese "flash" models for my needs. They are a joy and they cost pennies per answer. Great job. |
|
| ▲ | ActionHank 4 hours ago | parent | next [-] |
| I am legitimately more excited for this release than any frontier models at this point. I don't need a model that can invent new mathematics. I need something that is fast, cheap, and consistent. Give me that and I can build and scale. |
| |
| ▲ | Oras 2 hours ago | parent [-] | | LLMs are not consistent | | |
| ▲ | ActionHank an hour ago | parent [-] | | True, make them cheap and fast enough and you can scope and stack agents sufficiently that the error rate tends close enough to zero to be meaningfully useful. |
|
|
|
| ▲ | pimeys 4 hours ago | parent | prev | next [-] |
| Yes. I'm working in the agent industry and my god are we excited on new versions of Chinese flash models. The direct competition is Gemini Flash, and these models are much better on agentic tasks with fraction of the task price compared to Gemini. Things like oh here's a set of simple instructions for you to follow, call these tools, return this report. 20-30% of the price per task. And especially Deepseek Flash produces better quality than Gemini does. Where Gemini still wins is non-text input what Deepseek cannot do, yet, and Deepseek Flash has this thing of cheaper models where a failing tool call can derail your agent to a retry loop if you're not careful on instructions in the error message. If they fix and make the tool calls to work better in non-optimal situations, it's much easier to switch from Gemini without a few weeks of evals and bugfixing. |
| |
| ▲ | bitexploder 2 hours ago | parent | next [-] | | Which versions of flash and at what thinking levels? Which chinese flash models and at what thinking levels? What tasks? What completion rates? How was quality evaluated? | | |
| ▲ | pimeys 2 hours ago | parent [-] | | - Which versions: 3.6 vs 3.7 vs. 3.8 for Gemini Flash, and v4 0731 for Deepseek v4 Flash, and GLM 5.3 Flash - Medium for Gemini, high for Deepseek. - Things like find information, then understand something about it, then send a slack message or email etc. - Completion rates somewhere in 80-90%, Deepseek a bit better than Gemini - Quality evaluated by Fable 5.1 and Astra 6.0 acting as a rubric judge. Gemini quality would probably be better with high thinking level, but that would be 40% more expensive. And Deepseek is already third the price of Gemini. |
| |
| ▲ | urieiejr 3 hours ago | parent | prev [-] | | translation I make vaporware that doesnt do shit reliably and this chinese crap spouts plausible demos and spam calls more cheaply than the competition saaar | | |
| ▲ | pimeys 2 hours ago | parent [-] | | Well, it's much more than that. In general everybody's building agents now. You see these things that can help you to do things like adding things like OCR an appointment from a picture of a hand-written paper and add it to your calendar, search things from the internet, find that email with a PDF and add it to your local paperless instance. Building an agent like this by yourself is really easy. Now, we have Gemini's subscription, OpenAI's ChatGPT subscription and all those, 20 bucks a month right? What if you can spend that 20 bucks in tokens to do your own. And you pay 15 bucks _a year_ in tokens to run that? And you own the data, you own your code and integrations. It's really easy to do, and these flash models are _more than enough_ for simple agentic tasks. |
|
|
|
| ▲ | hgoel an hour ago | parent | prev | next [-] |
| Yeah, the latest batch of <256GB Chinese models are really nice. They're far less cryptic than Claude, and competent enough to feel almost near Opus. I canceled all my subscriptions and switched to running the Chinese models locally (not as a cost saving measure). |
|
| ▲ | darkoob12 4 hours ago | parent | prev | next [-] |
| My mental bias always kept me away from Chinese models. Because i know that china is a surveillance state and all the things we know about CCP. But after what we learned about OpenAI and how they most likely used user data to basically cheat in an open competition i think it does not matter which AI provider you use all of them will own your data and all of them can spy on you. So I am willing to switch to Chinese models. This way we help them develop and improve models some day we can run them locally. |
| |
| ▲ | kzrdude 39 minutes ago | parent | next [-] | | The Chinese labs have released interesting papers to accompany their releases too, especially DeepSeek and Kimi. This improves their standing among a few of us, who really like to see and read the papers with details about what they have changed and how their models work. | |
| ▲ | efficax 19 minutes ago | parent | prev | next [-] | | the US is also a surveillance state except about 80% of the surveillance is private companies (that are closely tied to the state) | |
| ▲ | ricardobeat 4 hours ago | parent | prev | next [-] | | These models are open-weights. Anyone can host them, you don’t have to use chinese servers even though most of them offer zero data-retention policies. | | |
| ▲ | m00dy 3 hours ago | parent [-] | | >>zero data-retention policies Yeah, that’s basically an industry-wide scam. | | |
| |
| ▲ | jsw97 3 hours ago | parent | prev | next [-] | | Even you think both cheat, you can't possibly think they both cheat the same amount. | |
| ▲ | Mashimo 4 hours ago | parent | prev | next [-] | | The new meta model is fast and very cheap as well, and when used through OpenCode you get quite a lot of free tokens. But meta is also THE surveillance company, so probably also not a good choice in your case. | | |
| ▲ | miroljub 4 hours ago | parent [-] | | If you work on open source projects, I don't care about surveilance. It's right there on Github with full history anyways. | | |
| ▲ | hn8726 3 hours ago | parent [-] | | Is it? You're still putting a lot of thought and guidance into the agent's harness, the final code is just a tiny bit of that. It's like giving a junior developer final code vs explaining the whys and nuance. Which I'm not sure I want to give Meta |
|
| |
| ▲ | el_io 4 hours ago | parent | prev | next [-] | | You can use those models from Openrouter, they have many Non-Chinese providers. | |
| ▲ | 4 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | epolanski 4 hours ago | parent | prev [-] | | You're naive if you're thinking the scumbags running the US companies aren't using your data. In any case old rules apply: if privacy is a concern don't share the data. I share all my work-related code because it's worthless, but I don't and would never share company business and process details, access to production/user data, etc. Meanwhile I know of people connecting all the kind of MCPs for datadog/sentry/jira/concluce/production databases to their harnessess..lol. |
|
|
| ▲ | XzAeRosho 4 hours ago | parent | prev | next [-] |
| Same for me. DeepSeek models are incredibly good at implementation and light planning. I still default to Opus models for feature planning, but for most simple features the Pro models suffice. Incredible good value and product they have built. |
|
| ▲ | serf 4 hours ago | parent | prev [-] |
| I recently had to config my harness to watch for cybersecurity flags from astra and funnel requests to flash when they occur because Astra gets queezy when you talk to it about UDP packets in games. Works fantastic. Glad there is a more 'uncensored' thing to fall back to when the frontier folk are too sensitive. |
| |
| ▲ | nicce 3 hours ago | parent [-] | | The only positive side is that it is harder for students to feed university exercises to the agent in cybersecurity and expect it to make them all. |
|