| ▲ | himata4113 2 hours ago | |||||||
MoE experts were likely trained independently / in a sparse format. Training anything beyond 2t on typical systems would be infuriantingly slow, you could do 4t on nvidias room-scale solution, but for a reasonable training speed / batch size it caps around 3t. | ||||||||
| ▲ | sosodev 2 hours ago | parent [-] | |||||||
Do you have any resources to share regarding independent expert training? I was under the impression that it's not feasible. | ||||||||
| ||||||||