What is the role of positional encoding in the original Transformer model?
-
A
To normalize token embeddings before attention
-
B
To inject sequence order information since attention is permutation-invariant
-
C
To reduce the dimensionality of token embeddings
-
D
To initialize the softmax output layer