I can't say this enough, a well tune deterministic solver always beats an LLM. In most of my experiments, the best the LLM can do is match the solvers results.