What preprocessing challenge arises specifically when handling Chinese or Japanese text that doesn't exist with English?
-
A
Capitalization normalization
-
B
Word boundary segmentation since words are not space-delimited
-
C
Removing punctuation marks
-
D
Handling contractions