Which technique in computer vision training uses two networks — a teacher and a student — where the student learns to mimic the teacher's output distribution rather than hard labels?
-
A
Data augmentation
-
B
Knowledge distillation
-
C
Batch normalization
-
D
Dropout regularization