| ▲ | bovermyer 4 hours ago | |
This stood out to me as a little concerning: > The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall. | ||
| ▲ | orangecat 3 hours ago | parent [-] | |
That seems to be for the "AA-Omniscience" test where you get +1 for a correct answer, -1 for a wrong answer, and 0 for "I don't know". If a model is more than 50% confident in its answer, it should go ahead and submit it even though it will sometimes be wrong. I'd be curious to see a version of the test where models are asked to give a probability that their answers are correct so we can see how calibrated they are. | ||