In seq2seq models, what problem does the attention mechanism primarily address?
-
A
Slow training speed on long sequences
-
B
The fixed-length bottleneck of the encoder's final hidden state
-
C
Out-of-vocabulary token handling
-
D
Gradient vanishing in the decoder LSTM