What is the vanishing gradient problem in deep neural networks?
-
A
Gradients become too large during backpropagation
-
B
Gradients shrink exponentially as they propagate backward through layers
-
C
The network forgets earlier training data over time
-
D
Weights converge to zero during initialization