Remix.run Logo
▲ Dust: Pretraining Transformers Without Backpropagation(qlabs.sh)
76 points by E-Reverance 3 hours ago | 6 comments
▲polyomino 2 hours ago | parent | next [-]

Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains? It would be cool to apply this to different stages and see if that affects the learning trajectory

▲api 2 hours ago | parent | prev [-]

It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?

▲vatsachak 2 hours ago | parent [-]

Not necessarily, backprop is highly parallelizable since it is just a bunch of matrix mults.

Something like Dust skips the backward pass on backprop. But other techniques like Neural Predictive Coding can be completely asynchronous, each "weight" can fire independent of those far away from it. Innocenti, et. al have shown that NPC gradients converge to backprop within a certain "regime".

The win with asynchronous techniques like NPC is that you do not need the extreme co-ordination that backprop requires and hence should be computationally much easier given the right device.

Although at this point the industry has so much money in the forward-backward pass system that I doubt a backprop successor would win unless someone makes NPC hardware feasible and can prove scaling up to billions of params

▲tyromaniac 2 minutes ago | parent | next [-]

Also has the "advantage" of being slightly more biologically plausible as the optimization happens locally rather than globally.

That idea was taken further by N'dri et al in PCL, in which "activation energy" was minimized as well, and inhibitory neurons added https://www.nature.com/articles/s41467-025-64234-z.pdf

While trying to find the link for that I stumbled upon

https://arxiv.org/pdf/2605.12732

Which also looks pretty interesting

▲janalsncm 19 minutes ago | parent | prev | next [-]

Knowing nothing about this, I wonder if it could be useful in situations where we can’t reliably sync with all the workers. Something like folding@home, where all the workers are just shaking weights and if one of them finds a winner it uploads to the central server?

▲ACCount39 23 minutes ago | parent | prev [-]

The reason why I don't see the promise for ML-only applications is that the coordination backprop requires comes very cheap to us.

"Much easier given the right device" - the "right" there just isn't shaped like the devices we actually build. And the price of "not having backprop" is usually expending more FLOPs, getting worse sample efficiency, etc.

The biggest "device" that doesn't do backprop is the brain, and that's because the brain doesn't have the connectivity or the coordination to pull it off. Both of those are "expensive" for something like it to implement. Cheap for us though. We aren't stuck with neurons that only get locally available information and have to implement learning rules based on that. So, skill issue?