Remix.run Logo
▲ Majromax an hour ago

> This feels like an obvious jab at Navier-Stokes, but where we get to shift the numbers around to hide where the compute actually is being spent ... compute is being spent. It's either being spent in amortization to make the search smarter ahead of time, during training, or its being spent after.

I think that argument is recursive? These posts aren't very complicated for either of us, but they're written on devices that are fabricated with billions of dollars of semiconductor equipment. At what point do we just acknowledge that we stand on the shoulders of giants?

To me, the distinguishing factor is that the expense not special-purpose but upfront. The model here is trained without foreknowledge of what problems it will solve. Solutions like nine loops are genuine expressions of a pre-existing model capability, even if that capability has not pre-existed for very long.

▲mccoyb an hour ago | parent [-]

Oh I agree with you! I just don't see "we're standing on the shoulders of giants" in most of these marketing blog posts?

If I'm wrong here, I'd love reference links. I think of these companies as trying to inspire the idea that Claude (or GPT) are these special alien entities, in a sense?