In my work on LLM as a judge, I prefer to use LLM decisions as features in a downstream classic ML model for the final decision. It works really well
https://softwaredoug.com/blog/2025/01/21/llm-judge-decision-...