What is the primary purpose of the Transformer's multi-head attention mechanism?
-
A
To reduce model size by pruning redundant heads
-
B
To allow the model to jointly attend to information from different representation subspaces
-
C
To speed up sequential token generation
-
D
To replace positional encodings with learned embeddings