| ▲ | ainch an hour ago | |
Agreed, Kimi is cheaper for coding - I say that explicitly in the post too. However I'd have to disagree with you on the "office task" front. General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack or Salesforce. If you're Anthropic pitching Claude to a bank (taking an example I'm familiar with), coding probably accounts for ~20% tops of the workforce, and it doesn't drive direct revenues. The 'agentic coding bump', but for all your analysts, traders, and wealth managers, would be a much more attractive prospect. I don't disagree that coding is the most successful use case so far (and probably more relevant to a HN audience). But I think the future of the labs is also contingent on them making progress on more general white collar work. I suspect that's why the Opus 5 release blog lists 3 coding benchmarks (FrontierBench, DeepSWE and FrontierCode) to 3 or 4 more general ones applicable to office work - depending on how you slice it (GDPVal, AutomationBench, Legal Agent Benchmark, BrowseComp). | ||
| ▲ | sroerick 34 minutes ago | parent [-] | |
Okay, I think that's fair, but I'm not convinced there's anybody actually doing large amounts of compute on office tasks? Do you know anybody? Can you point to anybody publicly documenting this? Can you can you point to any specific workflows where fable is being used in lieu of more basic models? Even if you provide exceptional answers for all of these I still think it is disingenuous at best to ignore coding tasks in writing this. I have to assume coding is 90% of the use cases for the frontier. You've made a case that the labs need these customers. You haven't made a case that the labs have these customers. | ||