Remix.run Logo
▲ lmf4lol 3 hours ago

Oh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.

Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.

But as a main driver. I love flash. And it brought our bill down by A LOT :D

▲PcChip 2 hours ago | parent | next [-]

>We run all our Personal Assistants now on flash

are you worried about sending all your data to third parties, especially if they're in different countries?

▲techmunky 2 hours ago | parent | next [-]

Not shilling for them but Ollama cloud hosts domestically with ZDR afaik. I run 95% of my open weight inference through them. The rest goes through Opencode Go $10 plan (which is enough to run 3 hermes agents on DSF 4.1 and leave plenty of left to experiment with when new models drop).

▲octoberfranklin 2 hours ago | parent [-]

Just a reminder that any API using a Cloudflare TLS certificate isn't ZDR.

The model engine provider might be ZDR, but the service as a whole isn't.

▲techmunky an hour ago | parent [-]

did not know. ty! have my updoot as thanks

▲figmert 2 hours ago | parent | prev | next [-]

I use it through OpenRouter, which has ZDR enforcement.

▲crossroadsguy 2 hours ago | parent | prev | next [-]

I have asked OP that question but I think there are providers who are not in China and they just host the model/inference.

▲yieldcrv 2 hours ago | parent | prev [-]

Just use a provider hosting it in your country especially if your country has major data centers then its the same as using Anthropic or GPT of GCP Model Garden or AWS Bedrock

nobody here is talking about running frontier level intelligence locally so if you’re Chinaphobic and prefer layers of corporations siphoning your data in between you and the party there are plenty of options instead of directly to the party

▲crossroadsguy 2 hours ago | parent | prev | next [-]

What is the cost of access like for DeepSeek-v4.1-flash, compared to GLM-5.3-flash via ZAI's Coding Plan? Because that's what I use; and often hit the "wait". I wouldn't mind trying a new model subscription or even API access which hits around glm-5.3-flash level weight class (which seem to be enough for me; with quite some human suprvision and nudging) but gives muuuuuuch moooore tokens for the same price.

(And what are the preferred providers?)

▲HKCM852 37 minutes ago | parent | prev | next [-]

What personal assistants are you using?

▲aftbit 3 hours ago | parent | prev [-]

Have you compared it against actual SOTA models like latest Fable or Astra?

▲mtrovo an hour ago | parent | next [-]

The author explains this very well tbh:

> Today's models are now good enough for high-quality unattended tasks. Chasing the latest and greatest is silly. It is fun to see the new Fable capabilities, but the tasks we throw at them are usually ridiculous (maybe even insulting) if you believe in LLM sentience. It's like asking a math PhD to organize the files on your desktop.

I'm using DS V4.1 Flash as my main model since their release and it works great for all my coding tasks. My setup is OpenCode Go subscription and obra/superpowers skill.

The only times I try to change models are on general planning tasks (like research this codebase for tech debt mitigation opportunities) or if I need deep research which would benefit from searching the web, in which I still think Gemini is still the best because of the speed and access to google search index. But these are not even 20% of my daily tasks.

▲sneurlax 3 hours ago | parent | prev [-]

Of course there's still a huge performance gap

but DS 4.1 Flash is good enough for most tasks