This is a problem with soft rules and LLM in general, and it makes sense that it gets worse on complicated tasks with long context.
Glad to be able to put some numbers on it.