What problem does gradient clipping solve during training of deep neural networks, particularly RNNs?
-
A
Vanishing gradients caused by deep networks
-
B
Exploding gradients that cause parameter updates to become excessively large
-
C
Overfitting due to large weight magnitudes
-
D
Slow convergence caused by sparse gradients