Remix.run Logo
k__ 6 hours ago

Around 5 percentage points better. (E.g., 87% instead of 82%)

Gecko4072 6 hours ago | parent | next [-]

So not worth it over flash? Even at ~7x the size it isn't worth the price hike. Flash may be a monster of a model due to all the RL it received from free usage everywhere.

networked 6 hours ago | parent | next [-]

I haven't tried DeepSeek V4 Pro 0813 yet. Recent experience tells me that larger models are worth it in non-obvious ways. MiMo-V2.5-Pro solved problems that DeepSeek V4 Flash 0731 couldn't solve for me: for example, adding a live counter for elided reasoning lines to a terminal-based coding harness. You wouldn't be able to tell from the scores on their respective Artifical Analysis page (https://artificialanalysis.ai/models/mimo-v2-5-pro, https://artificialanalysis.ai/models/deepseek-v4-flash). I like the DeepSeek V4 models, though. They critiqued my engineering decisions better than MiMo, and they seem to have a distinct aesthetic in the SVGs they write.

trollbridge 6 hours ago | parent [-]

Interesting - I've been dropping into MiMo-V2.5-Pro-UltraSpeed whenever Flash seems to be "stuck" and it usually figures it out. I use UltraSpeed just because I'm so frustrated by then that I'm impatient.

I still find 5.6-Sol can solve some things neither of those can, but it's so slow (and it's so hard to trace / debug the reasoning) that I just let it run overnight.

networked 5 hours ago | parent [-]

What about 5.6 Terra and especially Luna? Luna scores pretty high on benchmarks and seems to have different habits (like a denser pattern of tool use) and blind spots.

I'm trying out a development workflow where I generate mundane code with MiMo and Luna (and soon V4 Pro 0813?) and have Opus 5, which is running on only a Pro subscription, review and refactor it. I'm not sure it will justify the context switching, but it's an interesting exercise.

trollbridge 5 hours ago | parent [-]

Terra and Luna are fine, but they’re quite slow (OAI seems to be really slow lately) and don’t have the reasoning traces. My workflow really depends on them or I can’t switch models effectively.

saaga 6 hours ago | parent | prev | next [-]

Yea that's what I was thinking. Flash is nuts. I find I have to be a more precise and specific with it but damn. It's crossed a threshold of production grade coding for sure.

I was running a session over a couple days and it didnt cross a dollar lol.

npn 6 hours ago | parent | prev | next [-]

I still believe this is not the full potential of pro models. I expect they will release another checkpoint later this year.

k__ 6 hours ago | parent | prev | next [-]

I tried the previous Pro model and in the end it was 50% more expensive than the previous Flash.

Wasn't worth it.

eli 5 hours ago | parent | prev [-]

Opus 5 medium to Opus 5 max is only 3 points, if that puts it in context

sparkling 6 hours ago | parent | prev [-]

deepseek-v4-flash feels so fast and snappy, i'm loving it. Happy to trade speed for the the 5% degraded benchmarking performance.

saaga 6 hours ago | parent | next [-]

I feel the same too. I like the speed. I'm also a big fan of glm 5.2 fast. I can't wait for like 2000 t/s on these haha.

k__ 6 hours ago | parent | prev [-]

I wouldn't exactly call it snappy, but faster than Pro, yes.

ericd 6 hours ago | parent [-]

Single request depth on vllm with dspark, I'm getting ~200 tps, I'd say it's pretty snappy.

JacobAsmuth 6 hours ago | parent | next [-]

Well sure but you're running on tens of thousands of dollars of hardware.

ericd 4 hours ago | parent [-]

It's much faster than other models on that same hardware in the same size class. I've tested a few, it's by far the fastest I've tested.

And it wasn't tens* until recently. Didn't expect this to be one of my best performing assets this year.

k__ 4 hours ago | parent | prev [-]

I get like 80.