surely the LLM can do that. It is RL'd against some reward, if the known strategies are clearly suboptimal with easy improvement, it'll find it most likely