| ▲ | oh_no an hour ago | |
Luna is 1 point being on AA's index at 1/4 the cost, yes it "doesn't score as well" but paying 4x for 1 point is crazy if you're going off benchmarks. AA has Haiku 5.5 as cheaper than 4.1 Flash (both on Max, which isn't ideal but what can ya do) and a 4 point intelligence gap. Why do people like to think open models are more competitive than they are? | ||
| ▲ | pimeys 44 minutes ago | parent | next [-] | |
It is super bad on a bit more complex workflows and starts repeating same errors with the same tool until the cycle breaker hits. 6 is worse than 5.6 here. But it is amazing on generating a report on content generated by better agentic models such as DeepSeek or GLM, which both do a mediocre/bad job on reports. | ||
| ▲ | gregwebs an hour ago | parent | prev [-] | |
DeepSeek's own paper advises against using Max, showing that it normally doesn't perform that much better. I am not using it on Max, so that's not a useful benchmark for me. I have seen other benchmarks where Flash does significantly (30%) better than Luna. | ||