Anecdotally, +1. I’d also say this benchmark matches my experiences and how much I trust the model output