| ▲ | torginus 5 hours ago | |
You can see the breakdown here on what subtasks it outperforms and underperforms Fable. For example it trails in GPDVal which is a collection of everyday office tasks apparently, and r3 banking, which is a fintech related practical problem solving benchmark. https://artificialanalysis.ai/models/gpt-6-astra Edit: Just looking at the charts Gemini 3.8 looks like an absolute banger. Not much worse than SOTA, cheap, and fast too. | ||