| ▲ | sroerick an hour ago | |
Okay, I think that's fair, but I'm not convinced there's anybody actually doing large amounts of compute on office tasks? Do you know anybody? Can you point to anybody publicly documenting this? Can you can you point to any specific workflows where fable is being used in lieu of more basic models? Even if you provide exceptional answers for all of these I still think it is disingenuous at best to ignore coding tasks in writing this. I have to assume coding is 90% of the use cases for the frontier. You've made a case that the labs need these customers. You haven't made a case that the labs have these customers. | ||
| ▲ | ainch 16 minutes ago | parent [-] | |
No I think you're right that the amount of compute spent on office work is lower than coding - although I don't have any sense for the right share. The best source I could find was an OpenAI report [1] which mentions that ~66% of enterprise token generation is via Codex, which I would expect to skew entirely towards coding. But it's hard to say how the remainder is split, what proportion is 'frontier', or whether it's representative for Anthropic. On your questions - I've spoken to a number of execs and seniors behind closed doors but nothing public I can point to. Anecdotally, I've spoken to senior leaders at banks spending billions of tokens on one-off tasks like prepping execs for earnings calls or piloting end-to-end agent worflows for specific use cases (but mostly piecemeal/one-off). Many new financial analysts I've spoke to are also leaning on Fable to produce research docs and models - I hear that the models are improving rapidly for these tasks. This lot have been blindsided by the spend growth [2], the same as for coders in enterprise (e.g. Uber blowing annual budget in 4 months [3]), so I do think there's appetite and budget for a capable, cheaper open model - but Kimi does not obviously fill that role across the board, the way it might for coding. That said, I still largely agree with you - most of the office work stuff is still fairly piecemeal compared to the mass deployment of LLMs across software development. I think it's fair to say I could've focussed on coding more rather than taking AA's benchmark distribution as representative - perhaps a more balanced title would be "Kimi K3 is not cheap across the board"? I guess there's also some ambiguity about what 'cheap' means - as I said elsewhere in this thread, I think when someone people talk about the price of Chinese models, they imagine Deepseek competing with o1 for 1/20th of the price. Even thought it is better priced for coding, Kimi isn't Deepseek-level cheap. I do, however, think you could debate whether coding will remain at >50% total token usage going forwards - big enterprises are hunting for ways to get value out of LLMs, and the labs are investing a correspondingly large amount in generating demonstrations and RL environments to get the models up to par. That said, it's also possible that Chinese labs will shift focus to white collar applications now they've demonstrated a commanding lead on cost efficiency for coding, so, I mean who knows - it'll be interesting to get some detail when Anthropic IPOs. Sorry for the long reply! Appreciate it's quite meandering... [1] https://cdn.openai.com/pdf/5d1e1489-21c0-43e4-9d42-f87efdbf0... [2] https://www.reuters.com/business/finance/australias-cba-flag... [3] https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-c... | ||