| ▲ | Vetch 5 hours ago | |||||||
The most charitable explanation I can think of for this is something like regression to the mean. When a model is first released, there'll be a subset of users who, just by chance, sample the highest quality band of the distribution that answers their query. Some of them will rush over to social media and post about how amazing a model is. Over time, those users' mental model of responses will converge but they'll perceive the model's return to typical performance as a downgrade. This guess/explanation predicts that most users won't match what the initial social media hype claims, doesn't discount user experience as simple habituation nor does it assume companies are lying when they say there have been no changes to the model itself (quantization included). I also think there's an aspect where initial testing is more forgiving because the more persnickety polish bits can be ignored and tests are likely to have similar structure to things that can be trained for. Meanwhile, actual specific work items are a broader unusual distribution with more stringent acceptance criteria. Personally, I can detect a separation between Sol and Astra (but not as large as that between Opus and Fable). While they can solve most of the same problems, Astra takes less time, is less frustrating to talk to, is cleaner, notices more, spins wheels less and requires less corrections. | ||||||||
| ▲ | roywiggins an hour ago | parent | next [-] | |||||||
Another effect may be that with newer models, people try out the hard problems they got stuck with in older models, and when that model happens to succeed, they are very impressed. Perhaps the new model really is better... but perhaps merely being prompted to try again with the difficult problem just gave them a new dice roll and they came up lucky. And then regression to the mean kicks in. | ||||||||
| ▲ | kmeisthax 3 hours ago | parent | prev [-] | |||||||
To add onto this, if you use a shiny new model and it gives you a turd, you're not going to tweet about it ("hey guys, look what I made with Astra! Nothing!"), and even if you do nobody is going to interact with it so it does poorly in the algorithm, because it has to compete with all the people using the new model to make something that looks impressive. Then people get tired of the magic trick and the logic flips. | ||||||||
| ||||||||