Remix.run Logo
hmate9 7 hours ago

It is more expensive per task than 5.6-sol high: https://artificialanalysis.ai/models/gemini-3-8-flash#price-...

HJain13 6 hours ago | parent | next [-]

Cheaper at medium level while still being same score as Sol medium

radicalriddler 6 hours ago | parent | prev | next [-]

Huh, according to some of those charts, it's both dumber, and more expensive to run against their benchmarking tasks than Fable??? Seems crazy to me.

sejje 4 hours ago | parent [-]

Perhaps the model is able to evaluate that it's not done, and to keep pressing on in the face of mounting failures, until it eventually arrives at a solution. Where Fable can skip that.

jdthedisciple 6 hours ago | parent | prev [-]

Sol is still underrated imo, especially for the current discounted price