Remix.run Logo
▲ aeve890 3 hours ago

>Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware

That's low hanging for you?

▲guyomes 3 hours ago | parent | next [-]

If we throw in hardware dedicated to a specific LLM, it seems to be a rather low hanging fruit. Especially considering that this is already happening for vision models [1].

[1]: "FPGA-based CNN Acceleration using Pattern-Aware Pruning" https://inria.hal.science/hal-04689673/document

▲mdp2021 34 minutes ago | parent [-]

> hardware dedicated to a specific LLM

That wording screams "Taalas". Which, importantly, is not the only player trying to abate the distance between data and arithmetics...

▲msdz 3 hours ago | parent | prev | next [-]

Maybe they meant in the sense of “untapped potential”, because so far a lot of the focus has been on increasing model capabilities, not necessarily performance/power budget.

▲TeMPOraL 2 hours ago | parent [-]

Yes. Point is, it's untapped only because everyone is running in the race (even if out of curiosity), and there's just not enough people with means to tap into these side threads. For the past few years, there's been many interesting papers that circulated the industry, got recognized as worthwhile pursuits, and then dropped because running behind the Big Vendors had massively better ROI.

▲mdp2021 an hour ago | parent | prev | next [-]

> That's low hanging for you

An important part of the industry is studying that: it is built-up effort. Sooner or later, the fruits will be harvested. The targeted preparation has been there for years now.

▲blurbleblurble 3 hours ago | parent | prev | next [-]

It's likely quite close. There are so many papers proving concepts that would bring this, they just haven't been combined in production.

▲TeMPOraL 2 hours ago | parent | prev | next [-]

Yes. It's well within realm of possibility, but so far wasn't pursued because the Big Vendors went all-in into capability growth (rightfully testing "the bitter lesson" to its limits) and got themselves stuck in an arms race, while everyone else is barely keeping up and/or starstruck with fascination, exploring what these models can do.

This got everyone racing forward and right now there is not enough human attention left in the world to productionize this, or any of the other "side threads". When the race slows down, people will catch up, branch out, and loop back.

▲blurbleblurble 2 hours ago | parent [-]

Just like renewable energy and so many other things. Hyperconcentration of capital is really tragic. I hope things turn around.

▲TeMPOraL 2 hours ago | parent [-]

They will. That's the fallacy of the "S-curve" everyone likes to commit these days actually giving a positive outlook.

Assuming it won't get to full RSI, the current approach will burn out - most likely economically. The race slows down, people branch out, look back, start picking up the "untapped potential"/low-hanging fruits, and you have new S-curves launching in place of the one that just tapered off (hence a fallacy - a stack of S-curves adds up to continuing exponential growth).

In other words: it comes and goes. Hyperconcentrated capital will eventually deconcentrate.

▲ekabod 3 hours ago | parent | prev [-]

That's a high hanging fruit, not low.