| ▲ | sroerick 2 hours ago | |||||||
Not to mention - If you switch the view to "coding tasks" on this website:
So it's pretty dang cheap lol. Nobody is using frontier inference for "office tasks". | ||||||||
| ▲ | ainch an hour ago | parent | next [-] | |||||||
Agreed, Kimi is cheaper for coding - I say that explicitly in the post too. However I'd have to disagree with you on the "office task" front. General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack or Salesforce. If you're Anthropic pitching Claude to a bank (taking an example I'm familiar with), coding probably accounts for ~20% tops of the workforce, and it doesn't drive direct revenues. The 'agentic coding bump', but for all your analysts, traders, and wealth managers, would be a much more attractive prospect. I don't disagree that coding is the most successful use case so far (and probably more relevant to a HN audience). But I think the future of the labs is also contingent on them making progress on more general white collar work. I suspect that's why the Opus 5 release blog lists 3 coding benchmarks (FrontierBench, DeepSWE and FrontierCode) to 3 or 4 more general ones applicable to office work - depending on how you slice it (GDPVal, AutomationBench, Legal Agent Benchmark, BrowseComp). | ||||||||
| ||||||||
| ▲ | ofjcihen 2 hours ago | parent | prev [-] | |||||||
Right? The availability of this being in the article that’s pushing the opposite narrative is like…what? | ||||||||
| ||||||||