Remix.run Logo
bobbylarrybobby 5 hours ago

The models themselves have far from plateaued. Maybe someone finds a way to get a really capable model down to, say, 12GB of ram. Then we'd be in business.

2 hours ago | parent | next [-]
[deleted]
swiftcoder 3 hours ago | parent | prev [-]

Agreed. We've just seen DeepSeek post-train their ~300 billion parameter flash model to outperform their 1.6 trillion parameter pro model, in the space of a few months. There would seem to still be quite a few opportunities on the table to bring big model smarts down to the smaller models