โ† All NLP Flashcard Decks

Text Classification Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Text Classification flashcards as text
  1. Which machine learning algorithm is considered a generative classifier commonly used for text classification due to its simplicity and strong baseline performance?

    Answer: Naive Bayes

    Naive Bayes is a generative classifier that models the joint probability of features and labels, and despite its 'naive' independence assumption, it performs surprisingly well on text data.

  2. In text classification, what is the primary purpose of TF-IDF (Term Frequency-Inverse Document Frequency)?

    Answer: To weight words by importance relative to a corpus

    TF-IDF weights a word higher when it appears frequently in a document but rarely across the corpus, highlighting distinctive terms for classification.

  3. Which evaluation metric is most appropriate for a highly imbalanced text classification dataset where the positive class is rare?

    Answer: F1-Score

    F1-Score balances precision and recall, making it more informative than accuracy when the class distribution is skewed and the minority class matters most.

  4. What is multi-label text classification?

    Answer: Assigning multiple non-exclusive labels to a single text instance

    Multi-label classification allows each document to belong to several categories simultaneously, such as a news article tagged as both 'politics' and 'economy'.

  5. In the Bag-of-Words (BoW) model for text classification, which information is explicitly discarded?

    Answer: Word order and syntax

    The BoW model treats a document as an unordered set of word counts, discarding grammatical structure, word order, and context.

  6. Which deep learning architecture introduced the concept of self-attention that dramatically improved text classification benchmarks?

    Answer: Transformer

    The Transformer architecture uses self-attention mechanisms to capture long-range dependencies in text, enabling models like BERT to achieve state-of-the-art classification results.

  7. What does 'zero-shot text classification' mean in the context of modern NLP?

    Answer: Classifying text into categories that were not seen during model training

    Zero-shot classification leverages large pre-trained language models to assign labels to new, unseen categories by understanding the semantic meaning of the label names.