| ▲ | akssri 4 hours ago | |
The intuition here is okay - but the math is hand-wavy with imprecise terms like "blow-up" etc. The statements however, if taken to mean optimality, are also incorrect. Reverse-mode AD (backprop) is generally quite efficient for scalar outputs (more generally, when n_inputs >> n_outputs), but it's not strictly optimal even for this particular scalar-output case. Consider for eg. a MLP, with 4-layers with dims (1, N, 1, N, 1) - reverse-mode here does ~3N multiplies, but the optimal is ~2N. The optimal ordering for gradient accumulation is in fact NP-hard on general DAGs, but such 'cross-mode' AD is apparently quite hard to implement and not often seen given the marginal gains. Griewank-Walther's excellent book is a excellent reference for this and much more, https://epubs.siam.org/doi/book/10.1137/1.9780898717761 They also had a library called ADOL-C that had mixed-mode. | ||