Remix.run Logo
DetroitThrow 6 hours ago

I've tried it on my "let's run every model in parallel and see which finds more edge cases" type of tasks, and Grok 4.5 was really behind Opus/ChatGPT but ahead of Gemini - despite having a strong showing on benchmarks.

That makes me really skeptical of it being GPT5.6-tier, much less Fable-tier, based on some of these benchmarks alone. But I'll test here shortly.

DetroitThrow 5 hours ago | parent [-]

It's still not as good as GPT5.6 or Opus5 but it's better than KimiK3. Good job xAI team.