Remix.run Logo
dumberquestions 5 hours ago

"..and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token."

"3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%)"

So which one is it? 65% or 49%?

petu 5 hours ago | parent [-]

First sentence is about token efficiency.

dumberquestions 5 hours ago | parent [-]

You're right, should've gotten some LLM to summarize it instead of skimming.

semilin 5 hours ago | parent [-]

Or you could have read it more closely before posting a comment saying it didn't make sense. You know, the old school way.

dumberquestions 5 hours ago | parent [-]

If I'm going to read it wrong might as well have an LLM to blame.