| ▲ | pu_pe an hour ago | |
The paper underlying this blog post is fundamentally flawed because of benchmark ceilings. If we define only simple tasks like asking what is the capital of France, all models will converge to 100%, obviously. But as bigger models get more capable we want them to replace more and more complex tasks, in as short time as possible. Then of course there is the economics of it. Do people prefer to spend $5000 upfront to get things done 5x slower, or would they rather pay $20 a month for that? | ||