Remix.run Logo
cainxinth 7 hours ago

It's the same story every time OpenAI or Anthropic releases a new model. They are generous with compute for the first few days, and use maximum fidelity with uncompressed weights. Everything runs at its best to make a good first impression. But eventually they pare things back and the models perform a little worse.

Vetch 5 hours ago | parent | next [-]

The most charitable explanation I can think of for this is something like regression to the mean. When a model is first released, there'll be a subset of users who, just by chance, sample the highest quality band of the distribution that answers their query. Some of them will rush over to social media and post about how amazing a model is. Over time, those users' mental model of responses will converge but they'll perceive the model's return to typical performance as a downgrade.

This guess/explanation predicts that most users won't match what the initial social media hype claims, doesn't discount user experience as simple habituation nor does it assume companies are lying when they say there have been no changes to the model itself (quantization included).

I also think there's an aspect where initial testing is more forgiving because the more persnickety polish bits can be ignored and tests are likely to have similar structure to things that can be trained for. Meanwhile, actual specific work items are a broader unusual distribution with more stringent acceptance criteria.

Personally, I can detect a separation between Sol and Astra (but not as large as that between Opus and Fable). While they can solve most of the same problems, Astra takes less time, is less frustrating to talk to, is cleaner, notices more, spins wheels less and requires less corrections.

roywiggins an hour ago | parent | next [-]

Another effect may be that with newer models, people try out the hard problems they got stuck with in older models, and when that model happens to succeed, they are very impressed.

Perhaps the new model really is better... but perhaps merely being prompted to try again with the difficult problem just gave them a new dice roll and they came up lucky. And then regression to the mean kicks in.

kmeisthax 3 hours ago | parent | prev [-]

To add onto this, if you use a shiny new model and it gives you a turd, you're not going to tweet about it ("hey guys, look what I made with Astra! Nothing!"), and even if you do nobody is going to interact with it so it does poorly in the algorithm, because it has to compete with all the people using the new model to make something that looks impressive. Then people get tired of the magic trick and the logic flips.

zaphirplane an hour ago | parent [-]

Really? there would be complaints, it’s expensive and doesn’t do as well

baby 6 hours ago | parent | prev | next [-]

You think they introduce stronger quantization after a few days?

boredatoms 6 hours ago | parent [-]

For sure they quickly move to q8, the output quality difference to bf16 is small compared to the speed/capacity gain

NineStarPoint 5 hours ago | parent | next [-]

Yeah q8 made so littler difference back when I was testing such things I'd be surprised if people could quickly notice that as a change. It's got to be either further quantized or some other type of optimization that kicks in when people notice the drop.

selectodude 5 hours ago | parent [-]

NVFP4 would buy them a huge increase in capacity but I think it would be noticeable.

Caracas288 4 hours ago | parent [-]

Why doesn't someone just try to measure this next time!?

embedding-shape an hour ago | parent [-]

Can't really measure without being sure you aren't being messed around with, when it's a remote platform. Stupidly easy to detect when people run such benchmarks/tests against you as well.

torginus 4 hours ago | parent | prev | next [-]

Some people here have remarked previously that while reduced precision doesn't show up in quick prompts, it does severely impact these models' ability to perform long running tasks - to the point that running these big models with severe quantization might be counterproductive as smaller but less quantized ones perform better.

nonethewiser 4 hours ago | parent | prev [-]

Could this explain Opus?

holoduke an hour ago | parent | prev | next [-]

I am sure every input send to openai is prechecked by a dumb model and then send to another one. They heavily tweak this to improve performance.

blurbleblurble 6 hours ago | parent | prev | next [-]

Or a lot worse

holler 6 hours ago | parent [-]

so, AGI is cancelled?

__MatrixMan__ 6 hours ago | parent [-]

AGI for the peasants is cancelled.

elwell 2 hours ago | parent [-]

Trogdor - the AGInator

dooglius 5 hours ago | parent | prev [-]

Do you have hard evidence of this assertion?

simlevesque 5 hours ago | parent | next [-]

We can't have hard evidence. It's a SaaS and they own the code and the machine it runs on.

So it may be a widespread hallucination. But there's no evidence of that either.

dooglius 2 hours ago | parent | next [-]

Run a benchmark with a large number of samples, rerun a few days later. Compare results, use statistics to see if there's a statistically significant difference.

fragmede 4 hours ago | parent | prev [-]

We could still have soft evidence though. Make a Todo app on Monday, and make a Todo app on Tuesday, and see what it makes in comparison.

marcus_cemes 4 hours ago | parent | next [-]

You would need a significant sample size to make any sort of conclusion from such a probabilistic process. Then there's the issue of how you would actually grade/compare.

ArvidSu 4 hours ago | parent | prev | next [-]

You only need to come up with a catchy "SomethingBench" name, post it on reddit/x and now you're an ai sage. Not to disparage the launch/after comparison though, I'd genuinely enjoy a data point like that

luckydata 4 hours ago | parent | prev [-]

someone already does that https://aistupidlevel.info/

chaimtweiss 3 hours ago | parent [-]

It's actually a extremely cool site, and fascinating to view the results off the AI bots i use.

bradly 5 hours ago | parent | prev [-]

There is a toot from an Open AI person a couple days ago saying they are "pulling all the levers" because of capacity issues. I have no idea what the heck the person is talking about, but I'm guessing there are consequence for those levers.

    > "Demand for Astra is really unprecedented. We're pulling all the levers possible to sustain the demand, but I've not seen anything like it until now and we went through very steep growth before. Priority will always be to keep excellent service for existing users, but we might have to pause new Pro subscriptions for a bit if this continues."
nonethewiser 4 hours ago | parent [-]

Depends on the nature of the levers