| ▲ | andai an hour ago | |
Their benchmark used to show other metrics, like output tokens and time, but now only shows cost: https://cognition.com/frontiercode Which is too bad, since all of the gains here appear to be from massively reduced output tokens? The model SWE-2 is based on, Kimi K3, is cheaper per token than Sol, but costs more per task (ArtificialAnalysis) due to using way more tokens. Whereas, based on the graphs, SWE-2 appears even more token-efficient than Sol! That might have been worth showing off, if true. | ||