| ▲ | yieldcrv 5 hours ago | |
It’s not just about small models, that’s only one part of evolution Some groups are baking models into silicone, Deepmind has an example, it gets 18,000 tokens/sec on Llama 3.1, not sure about parameter size | ||
| ▲ | fooker 2 hours ago | parent | next [-] | |
> Some groups are baking models into silicone While some other groups are baking silicone into models :) | ||
| ▲ | intrasight 5 hours ago | parent | prev [-] | |
I think this is the future - at least it will be for on-device models. Apple, for instance, will "bake silicon" once a year for their current model, and use that chip in all their devices. | ||