| ▲ | theplumber an hour ago | |
It would be great if all these benchmarks would also provide a slop metric. The main issue I see is not finding bugs but reporting “stupid” bugs and and if the recommendations are followed you end up overengineering the wrong things basically producing AI slop. GPT is prime candidate for this. Fable as well but less than GPT | ||