โ† All NLP Flashcard Decks

Named Entity Recognition Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Named Entity Recognition flashcards as text
  1. Which tagging scheme uses B- prefix for the beginning of an entity and I- prefix for continuation?

    Answer: IOB2 (BIO)

    IOB2 (also called BIO) uses B- for the first token of an entity and I- for subsequent tokens, with O for non-entity tokens.

  2. What is the primary advantage of the BILOU tagging scheme over BIO?

    Answer: It explicitly marks the last token and unit entities, aiding disambiguation

    BILOU adds L- (Last) and U- (Unit/singleton) tags, helping models distinguish entity boundaries more precisely than BIO.

  3. In a CRF-based NER model, what does the CRF layer model that a simple softmax classifier does not?

    Answer: Label dependencies between adjacent tokens

    CRF (Conditional Random Field) models the joint probability of the entire label sequence, capturing dependencies between consecutive labels.

  4. Which evaluation metric is most commonly reported for NER systems?

    Answer: Entity-level F1 score

    NER is evaluated using entity-level F1, where an entity is counted correct only if both its span and type are exactly correct.

  5. What challenge does 'nested NER' address?

    Answer: Entities where one entity span is contained within another

    Nested NER handles cases like 'Bank of America' (ORG) containing 'America' (LOC), where standard flat NER misses inner entities.

  6. Which dataset is a well-known English NER benchmark that includes PER, ORG, LOC, and MISC entity types from news wire?

    Answer: CoNLL-2003

    CoNLL-2003 is the canonical English NER benchmark derived from Reuters news, with four entity categories: PER, ORG, LOC, and MISC.

  7. What problem does 'mention detection' solve as a subtask before full NER?

    Answer: Identifying entity span boundaries without classifying type

    Mention detection identifies where entity spans are in text, leaving type classification to a subsequent step, which can improve pipeline modularity.