>token use is higher than k3 and far higher than proprietary models
GLM sets effort to max by default historically.
Aa also benchmarked k3 at max