Remix.run Logo
ConfusedDog 5 hours ago

Why would 3.6 flash perform a little worse than 3.5 flash on Artificial Analysis Coding Index...

https://artificialanalysis.ai/models/gemini-3-6-flash?intell...

sosodev 5 hours ago | parent | next [-]

Because AA Coding "Index" consists only of two benchmarks (Terminal-Bench v2.1, SciCode) and generally fails to be meaningfully representative of agentic coding capabilities.

Alifatisk 5 hours ago | parent [-]

Whats a better option for AA Coding Index?

WASDx 4 hours ago | parent [-]

DeepSWE and FrontierCode are more realistic if you read up on what they actually measure. But the most realistic is to try it yourself. Benchmarks can only vaguely represent typical usage, and how you judge the result. Giving the same real task you have to a few models will make you understand them better than chasing benchmarks.

wmedrano 5 hours ago | parent | prev [-]

Could be a good tradeoff for the flash model though. 3.5 -> 3.6 is a tiny bit cheaper and maybe faster?

artificialanalysis.ai has it going from 165 tps -> 304 tps. openrouter.ai needs more data but it has it going from ~100 tps -> ~150 tps, though at peak 3.5 has reached 156tps.