Remix.run Logo
▲ Razengan 2 hours ago

Theory (Conjecture? Hypothesis?): What we notice as "model nerfing" is the company diverting compute to training/running new unreleased models..

Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced

▲Centigonal 2 hours ago | parent | next [-]

wouldn't less compute result in slower inference, rather than worse performance?

▲latentsea 2 hours ago | parent | next [-]

They could potentially quantize the model and run it at lower quality taking less VRAM.

▲poizan42 25 minutes ago | parent | prev | next [-]

My guess is that they are dynamically changing the quality of the model to always keep the speed above some floor. So once it gets below that they switch to a worse quant or reduce reasoning level, or some combination of both.

▲btown 2 hours ago | parent | prev [-]

The more likely thing that would happen is that the provider begins silently interpreting (perhaps some) high effort-level requests as medium, etc., or having a classifier do this far more subtly. As such, the load on the cluster is less, and more resources can be devoted to training. Whether the frontier labs actually do this is purely conjecture at this point.

▲JohnBooty 31 minutes ago | parent | next [-]

I assume there's classification going on where a really basic "Hi how are you?" style request sent to a high-effort instance can be routed to a lower-level instance. This... is pretty much fine with me, assuming they do a good job of it.

I would also assume they use nebulous labels like "Medium Effort" or "High Effort" map to quantitative amounts of compute allocation... and that these amounts can be varied manually or automatically. Right?

I mean, there's a reason why they call it "High Effort" and not "Exactly 5 Minutes of GPU Time on Exactly 10 GPUs." They want to be able to move those sliders and tweak those knobs.

▲nightpool 2 hours ago | parent | prev [-]

why is that more likely?

▲zxilly an hour ago | parent [-]

Because they already did so. The model in Codex will get lower `juice` than API version.

▲jackmott42 an hour ago | parent | prev [-]

There is no nerfing, look at the data before coming up with a theory as to why the nerfing that isn't even happening is happening.

fuck