What is 'attention sink' phenomenon observed in large language models with long contexts?
-
A
Attention weights collapsing to zero for all tokens beyond the context window
-
B
Disproportionately high attention assigned to early tokens (especially the first) regardless of relevance
-
C
Attention heads specializing exclusively in syntactic rather than semantic patterns
-
D
Gradient flow being blocked at attention layers during backpropagation