Remix.run Logo
m_w_ 5 hours ago

It's a bit disheartening to see no comparison to other models here - and I'm not sure this pushes the curve anywhere. 3.6 flash is more expensive than GLM 5.2 - but seemingly worse, although this post is really light (lite?) on details.

It seemed for a time that Google had finally gotten the ball rolling, but I'm doubting that more and more as time passes. We'll see what happens with 3.5 pro I suppose.

XCSme 4 hours ago | parent | next [-]

Here, my comparison of 3.6 Flash vs Sol vs Luna vs Terra: https://aibenchy.com/compare/google-gemini-3-6-flash-medium/...

jdthedisciple 3 hours ago | parent [-]

How does your comparison work? It places Gemini 3.6 Flash Medium above GPT 5.6 Sol High and Fable 5 Medium, which makes me skeptical because that... would be making headlines that I'm not seeing right now.

XCSme 3 hours ago | parent [-]

I have created various questions/tests and put the models through the same tests.

I record whether the answers are correct, and the generation stats (costs, latencies, tokens used, etc.).

I have no idea why the Gemini models do so well.

I have recently added new tests, whose sole purpose was to find some cases on which Gemini 3 Flash fails (I don't like cherry-picking models or tests, but I also find it strange Gemini Flash models leading in accuracy). I made a more complex coding/tool-usage test, that I expected it to fail, it did fail it once locally in my debug tests, but when I finalized the test and ran the entire testing suite for all models, somehow Gemini 3 Flash still got it right...

Gemini models are REALLY intelligent (and they are actually my favorite model to use via the chat app to ask questions), but they somehow fail in real-word coding tasks where they have to modify files, check results, debug, etc.

My tests harness provides a lot of mock data, and limits the number of actions a model can choose from. I am starting to think that maybe the models are not bad, just that the coding harness are not optimized for those type of models, and Google doesn't really provide their own "Codex".

jdthedisciple 2 hours ago | parent | next [-]

Interesting, well it'd be interesting to check out some individual examples where Gemini beat the others.

Also, would be great if you could add GPT 5.6 Sol XHigh and Fable 5 High as well, just to see if at least those beat Gemini which is currently your #1.

XCSme 2 hours ago | parent [-]

I don't like to divulge tests, but one of them is a chess puzzle.

> would be great if you could add GPT 5.6 Sol XHigh and Fable 5 High as well

I would like too, but I avoided them for several reasons:

1) Cost - this is a hobby project, those models would cost tens of dollars for each benchmark run, multiply this by tens or hundreds of models and ...

2) Time - the high models are already taking a really long answer to respond (5-10minutes per question). I run each question with 3 repeats (run the same test three times), so it would take 30 minutes per test. If I change my tests, methodology, or add a new test, it would take a really long time to run the benchmark. Also, I like having results immediately when a new model is released, now I can post within 30 minutes of a model's release the benchmark results.

3) High reasoning usually does WORSE on most tests - if you look at the leaderboard, it's sometimes counter-intuitive, but models with high or max reasoning usually do worse than medium and low. This is because the questions are quite targeted/direct, and the models overthink the question and miss the solution. Or the long thinking context makes them perform poorly. The generation tasks (SVGs/HTML animation) are usually better with longer reasoning, but short code fixes, trivia questions, puzzles, etc. are answered by low/med reasoning with more accuracy in general

Also, Fable is borderline un-testable, it refuses to answer many questions, so it scores poorly anyway.

Gemini scores 21/22 because it answers all tests, and it does them correctly, consistently. The only failed test is I think because it miscounted the lines in a file, when responding on which line the bug was in a code snippet.

XCSme 2 hours ago | parent | prev [-]

Oh, and I've also added weights to different categories, so Coding and Tool usage categories influence the score more. This done both to better account for how most people are being used, and also to reduce Gemini's dominance in general/domain specific knowledge.

So yes, Gemini models are at the top, even if I actually (not proud of it) tried to make tests that actually favour other coding-focused models.

florakel 3 hours ago | parent | prev | next [-]

It’s really surprising. When Apple announced the multi-billion dollar deal with Google to power Apple Intelligence I thought great things were coming. Instead we are getting more and more bad news: delayed Pro models and AI leadership leaving. I wonder if Apple know something the rest of us don’t know or if they are already regretting their decision.

winstonp 3 minutes ago | parent | next [-]

Apple isn't counting on their model to be a frontier coding and cowork model. Gemini is perfectly fine for the tasks that new Siri is supposed to be doing.

WarmWash 2 hours ago | parent | prev | next [-]

Besides Apple apparently making Siri AI model agnostic, the choice to go with Google was almost certainly for practical reasons. Google is a low-risk established player that already has a long work history with Apple. Google also isn't in an existential battle to establish themselves, Gemini still amounts to just another project at Google. There is tangible non-zero risk that either OAI or Anthropic will be gone in 5 years, or will be forced to leave Apple high and dry to save themselves. There is almost no risk Google will be in either such position. And worst case scenario, Google has incredibly deep pockets should Apple pursue a "refund."

revolvingthrow 3 hours ago | parent | prev | next [-]

What Apple wants out of Google is Siri that runs at 8gb ram and isn’t a horrible embarrassment that feels like a primitive markov chain. Given how good Gemma 4 is, Google can squeeze some serious performance in small models. Whether they can make bleeding edge models is irrelevant to Apple.

zarzavat 3 hours ago | parent [-]

"Siri, please solve the Jacobian conjecture, and also set an alarm for 8am tomorrow"

Petersipoi 3 hours ago | parent | prev [-]

As someone on the Apple beta.. the model is almost completely irrelevant to the experience. Apple has gone and done Apple things by nerfing the experience so completely that almost any model in the past year would be fine. I still reach for ChatGPT/Claude/Grok constantly instead of the AI toy that Apple calls the new Siri.

armarr 5 hours ago | parent | prev | next [-]

GLM was twice as verbose running the Artificial Analysis benchmark. So it ends up being more expensive

Havoc 4 hours ago | parent | next [-]

>verbose

GLM defaults to max effort btw

https://docs.together.ai/docs/glm-5.2-quickstart#reasoning-e...

maxloh 4 hours ago | parent | prev [-]

Not really. Gemini 3.6 Flash actually cost $0.01 more per task, compared to GLM 5.2.

https://artificialanalysis.ai/models/gemini-3-6-flash

CSMastermind 4 hours ago | parent | prev | next [-]

All the benchmarks I see put it around the capabilities of Opus 4.8 Medium or Sonnet 5 High.

As far as I can tell it's slightly better than GLM 5.2.

sczi 3 hours ago | parent | prev | next [-]

The one thing I've found google's models to be the best at is proofreading text in non-english languages. Probably because I imagine they have the most training data for it as Google probably has the most complete archive of the internet.

ur-whale an hour ago | parent | prev | next [-]

> It's a bit disheartening to see no comparison to other models here

Disheartening, but not surprising: the comparison would not be very flattering for Google.

dd8601fn 5 hours ago | parent | prev | next [-]

> really light (lite?) on

Light. Lite is product marketing seepage.

crab_galaxy 5 hours ago | parent [-]

Yeah that’s the joke :p

kzrdude an hour ago | parent | next [-]

I've never questioned the word lite before because it's existed my whole life.. So does it make sense? Why does it exist and where does it come from? More than coming from "light".

dd8601fn 5 hours ago | parent | prev [-]

Sorry, it went right over my head!

lopatin 5 hours ago | parent | prev [-]

[dead]