did you test kimi k3 or qwen 3.8 max on the same task ? or plan to test them ? I respect those genuine users tests other than those benchmarks that models are trained and overfitted to them