| ▲ | bobbylarrybobby 5 hours ago | |
The models themselves have far from plateaued. Maybe someone finds a way to get a really capable model down to, say, 12GB of ram. Then we'd be in business. | ||
| ▲ | 2 hours ago | parent | next [-] | |
| [deleted] | ||
| ▲ | swiftcoder 3 hours ago | parent | prev [-] | |
Agreed. We've just seen DeepSeek post-train their ~300 billion parameter flash model to outperform their 1.6 trillion parameter pro model, in the space of a few months. There would seem to still be quite a few opportunities on the table to bring big model smarts down to the smaller models | ||