| ▲ | dakolli a day ago | |
I didn't accept a single edit from this model over the entire week, just saying. I do not understand how it's being benchmarked on par with Sol and other larger models. | ||
| ▲ | jazzpush2 a day ago | parent | next [-] | |
It was certainly almost RL-fried to overfit the benchmarks, at the expense of actual usability. See Opus 5. | ||
| ▲ | respectattentio a day ago | parent | prev [-] | |
is it a benchmarkmaxxing model?! | ||