Remix.run Logo
eis 6 hours ago

3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets...

3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash

Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...

WASDx 5 hours ago | parent [-]

3.7 high and 3.8 medium are essentially the same on AA intelligence and cost. Output tokens on DeepSWE gives the same picture. So there might be something to it but they have done other things as well. At least the tokens are really fast.

zuzululu 2 hours ago | parent [-]

i find deepswe not very reliable for instance it puts grok 4.6 xhigh over sol medium