| ▲ | namibj 7 hours ago | |
Oh, is the principle of sparse universal transformers finally in SoTA LLMs? I guess we did manage to eventually seriously crash into the wall "more compute than normal (non-looped/unique-weights) transformers can efficiently consume with the limited training data we have", plus massive focus on highly hands-off agentic tool use reasoning... https://arxiv.org/abs/2310.07096 Edit: read much of the article, it's brute force predecessor was explicitly called out as an almost-ancient example: > The looped transformer is nothing new, and the basic idea already appeared in the Universal Transformers paper from 2018 | ||