Remix.run Logo
▲ oh_no an hour ago

Luna is 1 point being on AA's index at 1/4 the cost, yes it "doesn't score as well" but paying 4x for 1 point is crazy if you're going off benchmarks.

AA has Haiku 5.5 as cheaper than 4.1 Flash (both on Max, which isn't ideal but what can ya do) and a 4 point intelligence gap.

Why do people like to think open models are more competitive than they are?

▲pimeys 44 minutes ago | parent | next [-]

It is super bad on a bit more complex workflows and starts repeating same errors with the same tool until the cycle breaker hits.

6 is worse than 5.6 here.

But it is amazing on generating a report on content generated by better agentic models such as DeepSeek or GLM, which both do a mediocre/bad job on reports.

▲gregwebs an hour ago | parent | prev [-]

DeepSeek's own paper advises against using Max, showing that it normally doesn't perform that much better. I am not using it on Max, so that's not a useful benchmark for me. I have seen other benchmarks where Flash does significantly (30%) better than Luna.