Remix.run Logo
pu_pe 2 days ago

How do you explain the fact that Qwen3.8 27B performs vastly better than any open model from even one year ago, if using the same test-time compute and harness?

abixb 2 days ago | parent [-]

"Vastly better" in what ways? Benchmarks? You know Benchmarks can be optimized for and benchmaxxed for, right?

pu_pe 2 days ago | parent [-]

It's obviously more capable in any task I tried (coding, translation, summarizing, etc). Benchmarks are not the only way to tell if a model is better or not.

huurtehoog 2 days ago | parent | next [-]

I wanna see numbers showing companies and countries having excess growth due to these tools. Where are these data?

It's all vibes, and the numbers contradict the vibes. There's 30 years of literature trying to explain the "productivity paradox" where we can't see any excess productivity driven by computer technology. Lots of FOMO, no hard data. For an entire generation. And people come here every day and say stuff like you just said and they really seem to think that "this time is different".

butlike a day ago | parent | prev [-]

You gotta define 'obviously' here.