Remix.run Logo
torginus 8 hours ago

I think the results might be underwhelming - AI providers need to turn a profit and can't subsidize, and they're working off of the commodity hardware everyone does.

I wouldn't be surprised if they started offering potentiall bad quantizations with much reduced capability at lower prices (without telling the users, of course)

user43928 6 hours ago | parent [-]

I would be surprised, considering that OpenRouter requires disclosing the quantization and shows automatic benchmarks to compare between providers for the same model.

m00dy 5 hours ago | parent [-]

The latter is a joke

user43928 5 hours ago | parent [-]

I saw that it runs GPQA Diamond and TAU-Bench Airline and shows the results over a 32 day rolling average.

Other than that they track Tool call error rate and Structured output error rate.

I only discovered this today, and it seems like a good idea. What are the problems in practice?