| ▲ | Aurornis 2 hours ago | |||||||||||||||||||||||||
> I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon. Taalas needed a giant chip (6nm) for an 8B model. At best you could use a more advanced node to try to put a MoE model across several chips working together, but you can’t have GPT Sol size models on a single chip like that. | ||||||||||||||||||||||||||
| ▲ | greenknight 2 hours ago | parent [-] | |||||||||||||||||||||||||
Nope. But we are hitting some pretty impressive levels with 128B models. The other thing is, a lot of the time, model performance is improved with more 'thinking' time. The thinking time is just more tokens... but instead of say 1000 tokens or 10,000 tokens worth of thinking its 1,000,000... how does that improve model performance? Could a 128B model hit levels of GPT Sol? | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||