Remix.run Logo
▲ TeMPOraL 3 hours ago

Jev is one.

Diffusion transformers are not "easy" but underfunded.

Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.

E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.

Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.

Or imagine automated sliding doors that don't suck.

--

[0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.

▲mdp2021 an hour ago | parent | next [-]

There are "low hanging fruits" - easier to achieve goals -, and there are super-fruits, milestone-fruits.

Among the most important ones:

-- the long-known Problem of Transparency, applied to the apparent emergent intelligence in NNs. Why does it happen - in detail?

-- then, a Theory of Apparent Intelligence through NNs. Transforming the results achieved into a Science. Which allows to do what we are doing - but in a lean and targeted way.

-- then, a General Theory of Intelligence, that includes the above to go beyond current architectures and get those features of Intelligence we expect and still not have.

The long-term direction we got into must lead to this.

(You note a ponderant detail of the above when you note the importance of explaining the emergence of a World Model from a Language Model.)

▲patcon 19 minutes ago | parent | next [-]

If anyone is interested, following Dr Michael Levin's Thoughtforms.life podcast is the cutting edge of where all this previously fuzzy stuff is becoming more concrete. So long as you can tolerate distinguished scientists flailing about as they discuss consciousness and life and developmental biology (and other less-obviously living things, like algorithms) as involving "free lunches" and "ingressing patterns from the platonic realm" :)

▲TeMPOraL an hour ago | parent | prev [-]

Those are the absolutely fascinating parts, and I sincerely hope AI won't get out of control before we're able to tackle some of these.

▲blurbleblurble 3 hours ago | parent | prev | next [-]

Diffusion models combined with these new looping techniques are gonna change the whole conversation about efficiency. Imagine control net but in one or more conceptual latent spaces.

But also harnesses and more generally new insights on "the control flow problem" could end up squeezing a ton of performance out of small models.

▲seizethecheese 2 hours ago | parent | prev | next [-]

I commend you for actually answering, independent of what I think of the answers.

▲dominotw 41 minutes ago | parent [-]

not a very good answer though

▲flipping_beacon 3 hours ago | parent | prev | next [-]

Definitely agree with edge computation, although inference extensively researched and funded if SOTA LLMs hit a dead end tomorrow,there is still a lot to explore and research in inference and edge computation

▲Amekedl 2 hours ago | parent | prev | next [-]

yeah your reply, nobody can predict the future.

Enough stuff can happen, software use itself might change, and that could really cause anything. "What will we do with all the gpus" might become a question if for a magnitude of tech and reasons leaked-opus-9 runs on a macbook m6 or 7

▲aeve890 3 hours ago | parent | prev | next [-]

>Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware

That's low hanging for you?

▲guyomes 3 hours ago | parent | next [-]

If we throw in hardware dedicated to a specific LLM, it seems to be a rather low hanging fruit. Especially considering that this is already happening for vision models [1].

[1]: "FPGA-based CNN Acceleration using Pattern-Aware Pruning" https://inria.hal.science/hal-04689673/document

▲mdp2021 36 minutes ago | parent [-]

> hardware dedicated to a specific LLM

That wording screams "Taalas". Which, importantly, is not the only player trying to abate the distance between data and arithmetics...

▲msdz 3 hours ago | parent | prev | next [-]

Maybe they meant in the sense of “untapped potential”, because so far a lot of the focus has been on increasing model capabilities, not necessarily performance/power budget.

▲TeMPOraL 2 hours ago | parent [-]

Yes. Point is, it's untapped only because everyone is running in the race (even if out of curiosity), and there's just not enough people with means to tap into these side threads. For the past few years, there's been many interesting papers that circulated the industry, got recognized as worthwhile pursuits, and then dropped because running behind the Big Vendors had massively better ROI.

▲mdp2021 an hour ago | parent | prev | next [-]

> That's low hanging for you

An important part of the industry is studying that: it is built-up effort. Sooner or later, the fruits will be harvested. The targeted preparation has been there for years now.

▲blurbleblurble 3 hours ago | parent | prev | next [-]

It's likely quite close. There are so many papers proving concepts that would bring this, they just haven't been combined in production.

▲TeMPOraL 2 hours ago | parent | prev | next [-]

Yes. It's well within realm of possibility, but so far wasn't pursued because the Big Vendors went all-in into capability growth (rightfully testing "the bitter lesson" to its limits) and got themselves stuck in an arms race, while everyone else is barely keeping up and/or starstruck with fascination, exploring what these models can do.

This got everyone racing forward and right now there is not enough human attention left in the world to productionize this, or any of the other "side threads". When the race slows down, people will catch up, branch out, and loop back.

▲blurbleblurble 2 hours ago | parent [-]

Just like renewable energy and so many other things. Hyperconcentration of capital is really tragic. I hope things turn around.

▲TeMPOraL 2 hours ago | parent [-]

They will. That's the fallacy of the "S-curve" everyone likes to commit these days actually giving a positive outlook.

Assuming it won't get to full RSI, the current approach will burn out - most likely economically. The race slows down, people branch out, look back, start picking up the "untapped potential"/low-hanging fruits, and you have new S-curves launching in place of the one that just tapered off (hence a fallacy - a stack of S-curves adds up to continuing exponential growth).

In other words: it comes and goes. Hyperconcentrated capital will eventually deconcentrate.

▲ekabod 3 hours ago | parent | prev [-]

That's a high hanging fruit, not low.

▲alightsoul an hour ago | parent | prev [-]

[dead]