| ▲ | jll29 2 hours ago | |
I don't know who downvoted the parent or why, but it's a fair question IMHO. The answer is there can be dramatic difference running a benchmark one time, because LLMs are not deterministic. A proper methodology would ask each question 20 times and calculate the mean correctness across experiments. The reason is that the temperature parameter introduces random behavior. | ||