| ▲ | chrsw a day ago | ||||||||||||||||||||||
I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results. What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage? | |||||||||||||||||||||||
| ▲ | sebzim4500 16 hours ago | parent | next [-] | ||||||||||||||||||||||
>What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage? As far as the API goes, it would be really obvious. I run a small service that uses LLMs extensively, and if a model suddenly dropped in performance it would be straightforward for us to prove it. We regularly run comparisons where we generate completions with alternative models to e.g. see if we could get away with using cheap models for easy cases, if the baseline outputs deteriorated it would be all over our metrics. | |||||||||||||||||||||||
| ▲ | Wowfunhappy 21 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
> What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage? ...I mean, if they were actually doing this despite saying that they don't—promising one product and delivering something else—I think that would be fraud, no? And, maybe it's one thing to secretly defraud normies like us (although class action lawsuits do exist), but I don't think major enterprises or the US military would take too kindly to it. | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | dist-epoch 19 hours ago | parent | prev [-] | ||||||||||||||||||||||
It's called hedonic adaptation. > What is actually stopping these model companies You can say this about any company in the world, selling anything. It's trivially measurable, and there are people running the same benchmark on the leading models every day and measuring if they degrade. Spoiler: they don't. But you can always say "the conspiracy goes higher", and that the companies know about these daily benchmarks and are routing them to "quality" envs. | |||||||||||||||||||||||