| ▲ | cameronh90 a day ago | |||||||||||||
Maybe others have found otherwise, but I find the benchmarks drastically different to real world "feel" of a model, even within the same harness. I'm not sure if this just reflects personal interaction styles, or if it is indicative of benchmaxxing or unrealistic automated benchmarking methodology. Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium tends to be rated pretty poor on agentic tasks compared to the Chinese lab open models, but I find the latter much more likely to lose track of their own behaviour during a long-horizon task or get stuck in a doom loop. (This isn't a comment on GLM-5.3 Flash as I've not used it!) | ||||||||||||||
| ▲ | tyre a day ago | parent | next [-] | |||||||||||||
Opus is a yap god. I've found it much, much better with Claude Code's output style set to `Concise` and this: https://news.ycombinator.com/item?id=49413456 We shouldn't have to resort to this, but it can be mitigated enough that it stays as my daily worker agent. Although, I mostly use Fable to farm out to Opus agents so I don't have as much exposure to what kind of blathering is going on in there. | ||||||||||||||
| ||||||||||||||
| ▲ | joshheitzman 6 hours ago | parent | prev | next [-] | |||||||||||||
Yeah, I just ignore the benchmarks at this point. For open-weight models the provider's setup impacts performance so you can have different experience's with the same model at the same quantization from different provider's. Just have to use them on real tasks with your actual harness to really know how they will perform and hope the provider doesn't do something to degrade performance (e.g. update the middleware to a new version with a defect that impairs performance). | ||||||||||||||
| ▲ | crossroadsguy 12 hours ago | parent | prev | next [-] | |||||||||||||
For me it's been this: Opus (at least in my experience) is unbeatable at "planning the work". That includes a lot of things, including getting arch. sorted, a chassis/skeleton done. Filling that up and doing the actual "coding," though, I've noticed no real difference between the Claude model and GLM. So it'll be interesting to see how this flash model compares cost-wise to what I'm currently using, which is GLM 5.3, for coding. Looks like it will be reduced even further and might be great if it's faster (and better?) than 5.3 in my real world/personal experience. (I'm someone who doesn't really care about delays of a few seconds, or even more than few seconds. But if I am trying to notice then sure Claude is definitely faster as well). | ||||||||||||||
| ▲ | trey-jones a day ago | parent | prev [-] | |||||||||||||
I've been using 5.3 since they initially announced it and my gut feel is that it's not as good as 5.2 for agentic tasks. I'm still using it - I don't think it's bad. I'm just not convinced it's better. | ||||||||||||||
| ||||||||||||||