Yeah I would hold that models don’t know how to simplify because most rl/benchmarks doesn’t penalize complexity