Remix.run Logo
weird-eye-issue 2 hours ago

Then why are they (US frontier models) still so far ahead whenever I test them against the latest Chinese models? No bias here, I'd love them to be better for my own personal gain, but I haven't seen it

regularfry an hour ago | parent | next [-]

Behind on architecture, ahead on training? It seemed pretty obvious to me that the opus 4.7 and 4.8 releases were more about trying to retain 4.6-level capabilities while being cheaper to run, which would fit. And they can burn so much money on training.

weird-eye-issue an hour ago | parent [-]

I don't know I just care about the end result. And yeah what you're mentioning here is a pretty common conspiracy theory but you don't actually have any insight into that do you?

embedding-shape 11 minutes ago | parent | prev [-]

There is so much misinformation in the ecosystem, parrots just hitting "Reply" without thinking one iota, you really cannot trust "human" opinions on the internet anymore, anywhere.

Same with local LLMs, I'd love to use them for my day-to-day software engineering, and I'm not exactly GPU poor, then people with 12GB VRAM try to convince me their local setup is perfectly fine running latest Qwen and it does real engineering but whenever I try, they're a far cry from what Codex+GPT 5.x would do.

Only way to be sure is creating your own private benchmarks and use those, and the difference in quality becomes very apparent, very quickly, for your specific use cases.