| ▲ | misiti3780 19 hours ago | |||||||
is anyone doing this ? | ||||||||
| ▲ | LarsDu88 19 hours ago | parent | next [-] | |||||||
The closest example I've seen is ChatJimmy: https://chatjimmy.ai/ a prototype from Taalas running Llama 8B Scaling this up to 2.8 Trillion (350X increase), will certainly be challenging. If I was younger and had the right background, I'd love to dive into attempting somethign like this | ||||||||
| ||||||||
| ▲ | alach11 12 hours ago | parent | prev | next [-] | |||||||
News just broke today that Google is planning on doing this: https://news.ycombinator.com/item?id=48986351 | ||||||||
| ▲ | gopalv 19 hours ago | parent | prev | next [-] | |||||||
Taalas HC1 is the closest thing to this. Last I saw they posted Deepseek R1 numbers in Feb of this year. The challenge is rolling out a new one every 7-8 weeks as the weights change & cheap enough for a hyper scaler to afford to buy one and save enough on power over the next 8 weeks as a payoff. | ||||||||
| ▲ | carterschonwald 19 hours ago | parent | prev | next [-] | |||||||
not sure about that, but im actively working on designing ultra sparse models that i want to have perform competitively with stuff 100-10_000 times larger. ehich does yield similar throughput. time will tell id it works out | ||||||||
| ▲ | 535188B17C93743 19 hours ago | parent | prev | next [-] | |||||||
Yeah, I've heard of Taalas doing it. Not sure of others but I'm sure lots of companies are considering it, especially as we start to hit points of depreciating returns in training. | ||||||||
| ▲ | eckr 19 hours ago | parent | prev | next [-] | |||||||
There was a startup that did this for Llama 3, I forgot their name. Etched is also doing some similar things I believe. | ||||||||
| ▲ | robgough 19 hours ago | parent | prev | next [-] | |||||||
obligatory link to https://chatjimmy.ai | ||||||||
| ▲ | hnfong 19 hours ago | parent | prev [-] | |||||||
It sounds vaguely similar to what Cerebras.ai is doing | ||||||||