| ▲ | jascha_eng 6 hours ago | |
Imo omniscience correlates better to how useful the model is in practice than the intelligence index. But you have to use both together of course. | ||
| ▲ | smartbit 3 hours ago | parent [-] | |
Too late to edit, now including AA-omni Score [0] and 'by Domain' Software Engineering [1]. Also added comparison to the eyeballs-median of the 'top 10' models and then you see that indeed Mistral Large 4 Preview scores miserable in the AA-Omni indices.
[0] https://artificialanalysis.ai/evaluations/omniscience
[1] https://artificialanalysis.ai/evaluations/omniscience?detail... | ||