← All AML Flashcard Decks

Deep Learning & Neural Networks Flashcards

7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Deep Learning & Neural Networks flashcards as text
  1. What is the 'lottery ticket hypothesis' in the context of neural network pruning?

    Answer: Dense networks contain sparse subnetworks ('winning tickets') that, when trained in isolation from initialization, match full network accuracy

    Frankle & Carlin (2019) found that large networks contain small subnetworks with lucky initializations that can be trained alone to match full network performance.

  2. In neural ordinary differential equations (Neural ODEs), what replaces the discrete sequence of layers in a standard ResNet?

    Answer: A continuous-depth dynamics function defined by an ODE solved with a numerical integrator

    Neural ODEs parameterize hidden-state dynamics with a neural network and solve the resulting ODE, treating depth as a continuous variable.

  3. What is 'catastrophic forgetting' in continual learning, and which method most directly addresses it by constraining weight updates?

    Answer: Erasing previously learned task knowledge when training on new tasks; addressed by Elastic Weight Consolidation (EWC)

    EWC adds a regularization term that penalizes changes to weights deemed important for previous tasks, slowing forgetting while allowing new learning.

  4. In graph neural networks (GNNs), what is the 'over-smoothing' problem that limits the number of layers?

    Answer: Node representations converging to indistinguishable vectors as more neighborhood aggregation layers are added

    With many aggregation steps, all node features converge toward the same stationary distribution, losing local discriminative information.

  5. Which technique is used to prevent sequence-to-sequence models from attending to future tokens during auto-regressive training?

    Answer: Causal (look-ahead) masking that sets future attention scores to negative infinity

    A causal mask sets upper-triangular attention logits to −∞, so after softmax those positions receive zero weight and no future information leaks.

  6. What is the key innovation of 'flash attention' compared to standard scaled dot-product attention?

    Answer: It reorders computation to avoid materializing the full N×N attention matrix in HBM, reducing memory I/O bottlenecks

    Flash Attention tiles computation in SRAM, never writing the full attention matrix to slow GPU HBM memory, achieving significant speed and memory improvements.

  7. In the context of neural network quantization, what is 'quantization-aware training' (QAT) and how does it differ from post-training quantization (PTQ)?

    Answer: QAT simulates quantization effects during training so the model adapts its weights, while PTQ applies quantization to a pretrained model without retraining

    QAT inserts fake quantization operators during forward passes so the model learns weight distributions robust to precision reduction, yielding higher accuracy than PTQ.