Remix.run Logo
ofjcihen 2 hours ago

I mean at this point their very existence depends on it so I’m not sure if I’d be surprised

sroerick 2 hours ago | parent [-]

Not to mention -

If you switch the view to "coding tasks" on this website:

  Kimi K3: $3.18 per task
  GLM 5.2: $6.51 per task
  GPT 5.6 Sol: $7.02 per task
  Opus 5: 8.23 per task
  Fable: 11.70 per task
So it's pretty dang cheap lol. Nobody is using frontier inference for "office tasks".
ainch an hour ago | parent | next [-]

Agreed, Kimi is cheaper for coding - I say that explicitly in the post too. However I'd have to disagree with you on the "office task" front.

General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack or Salesforce. If you're Anthropic pitching Claude to a bank (taking an example I'm familiar with), coding probably accounts for ~20% tops of the workforce, and it doesn't drive direct revenues. The 'agentic coding bump', but for all your analysts, traders, and wealth managers, would be a much more attractive prospect.

I don't disagree that coding is the most successful use case so far (and probably more relevant to a HN audience). But I think the future of the labs is also contingent on them making progress on more general white collar work. I suspect that's why the Opus 5 release blog lists 3 coding benchmarks (FrontierBench, DeepSWE and FrontierCode) to 3 or 4 more general ones applicable to office work - depending on how you slice it (GDPVal, AutomationBench, Legal Agent Benchmark, BrowseComp).

sroerick 37 minutes ago | parent [-]

Okay, I think that's fair, but I'm not convinced there's anybody actually doing large amounts of compute on office tasks? Do you know anybody? Can you point to anybody publicly documenting this? Can you can you point to any specific workflows where fable is being used in lieu of more basic models?

Even if you provide exceptional answers for all of these I still think it is disingenuous at best to ignore coding tasks in writing this. I have to assume coding is 90% of the use cases for the frontier.

You've made a case that the labs need these customers. You haven't made a case that the labs have these customers.

ofjcihen 2 hours ago | parent | prev [-]

Right? The availability of this being in the article that’s pushing the opposite narrative is like…what?

sroerick 2 hours ago | parent [-]

I was actually shocked to see that much improvement on GLM 5.2. I am getting pretty great rates in GLM5.2 right now and I'm extremely happy with the output. I have found Kimi to be generally a little slower but noticeably better at architecture and structuring things. I would have thought is for sure currently more expensive than GLM5.2, particularly with subscriptions etc, but I'm excited for this to decrease further.