| ▲ | m00dy 6 hours ago | |
I would want to see three things before drawing strong conclusions: End-to-end tokens/sec and cost on realistic coding agent trajectories, including tool outputs and retries, not isolated decode benchmarks. Cache hit rates and prefill cost for branching, multi-turn sessions. Router-load distributions after post-training, where expert collapse or specialization problems often show up. | ||