| ▲ | Centigonal 2 hours ago | ||||||||||||||||||||||
wouldn't less compute result in slower inference, rather than worse performance? | |||||||||||||||||||||||
| ▲ | latentsea 2 hours ago | parent | next [-] | ||||||||||||||||||||||
They could potentially quantize the model and run it at lower quality taking less VRAM. | |||||||||||||||||||||||
| ▲ | poizan42 25 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||
My guess is that they are dynamically changing the quality of the model to always keep the speed above some floor. So once it gets below that they switch to a worse quant or reduce reasoning level, or some combination of both. | |||||||||||||||||||||||
| ▲ | btown 2 hours ago | parent | prev [-] | ||||||||||||||||||||||
The more likely thing that would happen is that the provider begins silently interpreting (perhaps some) high effort-level requests as medium, etc., or having a classifier do this far more subtly. As such, the load on the cluster is less, and more resources can be devoted to training. Whether the frontier labs actually do this is purely conjecture at this point. | |||||||||||||||||||||||
| |||||||||||||||||||||||