|
| ▲ | tyromaniac 33 minutes ago | parent | next [-] |
| Also has the "advantage" of being slightly more biologically plausible as the optimization happens locally rather than globally. That idea was taken further by N'dri et al in PCL, in which "activation energy" was minimized as well, and inhibitory neurons added
https://www.nature.com/articles/s41467-025-64234-z.pdf While trying to find the link for that I stumbled upon https://arxiv.org/pdf/2605.12732 Which also looks pretty interesting |
|
| ▲ | strbean 10 minutes ago | parent | prev | next [-] |
| Would these alternatives to backprop make it more feasible to have constant live-training going on in a model? Giving it something akin to neuro-plasticity? |
|
| ▲ | janalsncm an hour ago | parent | prev | next [-] |
| Knowing nothing about this, I wonder if it could be useful in situations where we can’t reliably sync with all the workers. Something like folding@home, where all the workers are just shaking weights and if one of them finds a winner it uploads to the central server? |
| |
| ▲ | vatsachak 22 minutes ago | parent [-] | | No real advantage over Neural Nets here; backprop matmuls can be calculated layer by layer so you can chunk backprop across different machines. The real advantage comes from energy savings, you require no global co-ordination | | |
| ▲ | janalsncm a few seconds ago | parent [-] | | Imagine we did that, split up a model layers as A->B->C. C will need to wait for B to compute a forward pass, which is waiting for A to compute its forward pass. To compute the forward pass, B needs all of the outputs from A, which is an upload and a download (maybe these can be done concurrently). Then A waits for B to compute its backwards pass, which is waiting for C to do the same thing. Again you are sending around potentially gigabytes of data. This is in contrast to mining bitcoins for example which doesn’t require any coordination from miners because their work is completely independent, and the answer is very small compared to the work needed to get it. |
|
|
|
| ▲ | ACCount39 an hour ago | parent | prev [-] |
| The reason why I don't see the promise for ML-only applications is that the coordination backprop requires comes very cheap to us. "Much easier given the right device" - the "right" there just isn't shaped like the devices we actually build. And the price of "not having backprop" is usually expending more FLOPs, getting worse sample efficiency, etc. The biggest "device" that doesn't do backprop is the brain, and that's because the brain doesn't have the connectivity or the coordination to pull it off. Both of those are "expensive" for something like it to implement. Cheap for us though. We aren't stuck with neurons that only get locally available information and have to implement learning rules based on that. So, skill issue? |
| |
| ▲ | vatsachak 28 minutes ago | parent [-] | | I mean co-ordination requires energy though. The brain wattage looks at GPUs and says "skill issue". But you're right that we look at natural energy production techniques and say "skill issue" |
|