So, basically it’s a little bit of optimisation everywhere to reduce cost and latency (partially found by Sol).
The only interesting part is that they trained the model for both task success and token spent.