Named Entity Recognition Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Named Entity Recognition flashcards as text
What is the role of a 'token classification head' when fine-tuning BERT for NER?
Answer: It applies a linear layer to each token's contextual representation to predict NER tags
A linear (dense) layer is placed on top of each token's BERT output embedding to project it into the NER label space, then trained with cross-entropy loss.
Which evaluation strategy counts an entity as correct only if its entire span and type exactly match the gold annotation?
Answer: Exact (strict) match F1
Strict (exact) match F1 requires both span boundaries and entity type to be correct; partial overlap does not count as a true positive.
What is 'data augmentation' commonly used for in low-resource NER?
Answer: Generating additional training examples by replacing entities with synonyms or back-translation
Entity-level substitution (swapping entity mentions with semantically similar ones) and back-translation are popular augmentation strategies to expand small NER training sets.
In multi-task learning for NER, what is a common auxiliary task trained jointly with entity recognition?
Answer: Part-of-speech tagging or chunking
POS tagging and chunking share syntactic structure with NER, so training them jointly often improves entity recognition through shared representations.
Which component of a pipeline-based information extraction system comes directly AFTER NER?
Answer: Relation extraction
Relation extraction identifies semantic relationships between entity pairs that NER has already identified, making it the natural downstream task.
What is 'transfer learning' in the context of domain-specific NER (e.g., clinical NER)?
Answer: Pre-training on general text then fine-tuning on domain-specific annotated data
Transfer learning pre-trains a language model on large corpora (e.g., PubMed) and then fine-tunes it on a small clinical NER dataset, achieving strong results with limited labeled data.
What problem does 'label inconsistency' cause in NER training data, and how is it typically addressed?
Answer: It introduces conflicting supervision signals; addressed through annotation guidelines and inter-annotator agreement checks
When the same entity is sometimes labeled and sometimes not (inconsistent annotation), it confuses the model; clear guidelines and high inter-annotator agreement (Cohen's kappa) mitigate this.