| ▲ | jdthedisciple 3 hours ago | ||||||||||||||||||||||
How does your comparison work? It places Gemini 3.6 Flash Medium above GPT 5.6 Sol High and Fable 5 Medium, which makes me skeptical because that... would be making headlines that I'm not seeing right now. | |||||||||||||||||||||||
| ▲ | XCSme 3 hours ago | parent [-] | ||||||||||||||||||||||
I have created various questions/tests and put the models through the same tests. I record whether the answers are correct, and the generation stats (costs, latencies, tokens used, etc.). I have no idea why the Gemini models do so well. I have recently added new tests, whose sole purpose was to find some cases on which Gemini 3 Flash fails (I don't like cherry-picking models or tests, but I also find it strange Gemini Flash models leading in accuracy). I made a more complex coding/tool-usage test, that I expected it to fail, it did fail it once locally in my debug tests, but when I finalized the test and ran the entire testing suite for all models, somehow Gemini 3 Flash still got it right... Gemini models are REALLY intelligent (and they are actually my favorite model to use via the chat app to ask questions), but they somehow fail in real-word coding tasks where they have to modify files, check results, debug, etc. My tests harness provides a lot of mock data, and limits the number of actions a model can choose from. I am starting to think that maybe the models are not bad, just that the coding harness are not optimized for those type of models, and Google doesn't really provide their own "Codex". | |||||||||||||||||||||||
| |||||||||||||||||||||||