| ▲ | bnchrch 4 days ago | |
I don't think the "don't really save you money" hot take holds water in every case. Coding, maybe. But for operationalized/repeatable tasks it definitely does. For example I have a workflow that I was running in April that effectively would cost $30k in token spend for each full run. However now, with GLM 5.3-flash, we've brought the cost down to $7k-9k with our evals showing we've had no loss in recall, precision etc.. | ||