| ▲ | killerstorm 4 hours ago | |
Making model to consists of many small modules is inefficient on GPU, especially as routing adds data dependencies, etc, and especially with pytorch (compared to a custom kernel). The difference might be smaller on a CPU which has limited parallelism. But it's basically equivalent to a very deep model which might be problematic for training. | ||