What is 'expert parallelism' in the context of Mixture-of-Experts (MoE) models on multi-GPU systems?
-
A
Distributing different expert sub-networks across GPUs so each GPU hosts a subset of experts
-
B
Assigning expert human reviewers to validate GPU outputs
-
C
Running the same expert network on all GPUs simultaneously for redundancy
-
D
Using specialized GPUs with higher memory bandwidth for attention layers only