Remix.run Logo
jklmnopqrstuvw 5 hours ago

Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project.

Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug.

Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.

bigmadshoe 3 hours ago | parent | next [-]

Why are people giving these n=1 comparisons like they mean anything? The worst offender is that pelican guy. These are non-deterministic systems and a single trial should not update your priors much at all.

Of course it's significant that your response had a bug and took four times longer, but if you're only going to try once, this isn't real science, it's just vibes.

jklmnopqrstuvw 2 hours ago | parent [-]

Months ago I start making this kind of test for my own reference. At beginning I I test each model multiple times, and results always same(pass or fail). Later I test only once for new models, I trust the results.

computerex 5 hours ago | parent | prev | next [-]

Repeat the test like 5 times for each model and see the results.

epolanski 5 hours ago | parent [-]

+1, a single test means little.

jklmnopqrstuvw 5 hours ago | parent [-]

I don't think so. I specifically kept this PR to test model capabilities, and I've already tested a bunch of models. Current test results show that the more advanced the model is, the easier it passes. For example, GPT-5.5 Medium fails the test(has bug), but High passed.

computerex 4 hours ago | parent | next [-]

They are causal autoregressive models, the output is sensitive even to the implementation nuances in inference. Even 1 token that's badly selected could throw off the entire answer.

segmondy 3 hours ago | parent [-]

you're thinking of one shot. if they are running an agentic loop then they don't need multiple passes. an agentic loop is multiple passes with tool calls and tools could fail and agent would correct from seeing the failure. a bad model will compound on error and fail, a good model will correct. 1 test is fine to gauge the quality of the model.

computerex 2 hours ago | parent [-]

An agent doing a task even with multiple back to back calls like normal without an example is zero shot. An agent doing a task with 1 example is one shot. An agent doing a task with a few examples is few shot. I don't think you are correctly using these terms.

The multiple back to back LLM calls are done on accumulating context, so if there is a sampling error it could throw the entire session out of whack, because LLM's build on the previous context.

It's actually meaningless to argue, one could simply sample more than 1 times and let the numbers speak for themselves.

gpt5 an hour ago | parent [-]

That's not true. An agent in a loop can test itself, review, verify and iterate as much as needed. That's one of the primary reasons more capable models tend to have a higher success rate.

I don't disagree that multiple tests increase confidence, but it's not correct to argue that an agent in a loop harness is equivalent to oneshotting

seunosewa 5 hours ago | parent | prev [-]

Do it a second time at least.

5 hours ago | parent | prev | next [-]
[deleted]
Zetaphor 5 hours ago | parent | prev | next [-]

It's the third link on the front page right now?

hugmynutus 3 hours ago | parent | prev | next [-]

Nullius in verba

ferongr 5 hours ago | parent | prev | next [-]

[flagged]

nozzlegear 5 hours ago | parent | next [-]

This but unironically

5 hours ago | parent | next [-]
[deleted]
SV_BubbleTime 5 hours ago | parent | prev [-]

[flagged]

nozzlegear 4 hours ago | parent | next [-]

Is this bait?

> So.. flesh this out and don’t be a coward about it.

No.

aliasxneo 5 hours ago | parent | prev [-]

There are still sane people here, we just don't talk about "rocket man" because it enrages the particular subgroup on display here and usually goes no where actually productive (and has like a 50% of getting flagged to death anyways). I'm not pro Elon by any means, but the standard HN profile of him is pretty bat shit crazy.

sergiotapia 4 hours ago | parent | next [-]

Correct

dgellow 3 hours ago | parent | prev [-]

He is honestly batshit insane

gafferongames 2 hours ago | parent | prev | next [-]

Yes.

numpad0 5 hours ago | parent | prev | next [-]

no he and his stuffs are now considered transparent, no pun intended. I think he deserves it since his minions were persistent with usage of "this ___ has hateful bias against ___" canned response.

5 hours ago | parent | prev [-]
[deleted]
NooneAtAll3 5 hours ago | parent | prev [-]

I thought it was impossible to downvote posts?

benjiro29 3 hours ago | parent | next [-]

I thought it was impossible to downvote posts?

User Posts can be downvoted but you need over 500 karma to have access to the downvote button. A Submission can not be downvoted.

numpad0 4 hours ago | parent | prev [-]

Maybe a tug of war between flags and vouches might work like downvotes?