| ▲ | dualvariable 2 hours ago | |
This question would be better answered if people were careful about distinguishing between "models" and "transformer architecture". If you bake a given transformer architecture into silicon and then, a year later, changes in transformer architecture give a large inference performance boost, you may have to throw away all that now nearly-useless silicon that gets outperformed by humble GPUs. | ||