The last time grok made these statement, I tried using it for my workflows and it did not perform as good as opus or even sonnet.
My guess is that xai benchmaxxes a lot but fails in actual capacity to produce good models.