| ▲ | disiplus 3 hours ago | |||||||
Was this post generated with LLM, did he properly mention anywhere why exactly did it fail with example or i have trouble reading. | ||||||||
| ▲ | desmondl an hour ago | parent | next [-] | |||||||
Yeah, I clicked expecting a review of GLM 5.3 Flash, but the article was more a retrospective about what he learned during his September challenge: "Only use GLM 5.3 Flash for one month" He said his experiment was a failure because: 1. He accidentally spent 450M tokens vibe coding with the wrong model, instead of GLM 5.3 Flash. 2. When he used GLM 5.3 Flash, it was sometimes slow. So he switched to other models (Deepseek / Qwen) instead. His guess to why it was slow: GLM 5.3 Flash was so good that the providers were congested. 3. He still needed to use other models besides GLM 5.3 Flash, for R&D and benchmarking. His takeaways from doing the experiment were: 1. Measure local usage more. 2. Experiment with agent orchestration, with bounded goals. 3. Don't count other models that are used for R&D. 4. Play with Jev. 5. Include experiments with flagship models to compare with cheap open models. His conclusion about GLM 5.3 Flash: Probably viable for day to day work, but he'll have more thoughts next month. | ||||||||
| ▲ | raggi an hour ago | parent | prev [-] | |||||||
Some vague commentary about performance with what appears to be assumptions about GPU availability, but no clarity about which inference provider is being used. If ZAI is assumed, I believe they aren't subject to the assumptions in the post based on what they've said publicly, but if they were using some other provider, perhaps. The second reason appeared to be simply "because we chose not to". The post seems to be pretty much content-less in any practical sense. I clicked on it because I do quite like this models average performance and I was hoping to see some kind of review content. | ||||||||
| ||||||||