← All AML Flashcard Decks

Deep Learning & Neural Networks Flashcards

7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Deep Learning & Neural Networks flashcards as text
  1. Which technique addresses the vanishing gradient problem by allowing gradients to flow directly through skip connections?

    Answer: Residual connections (ResNets)

    Residual connections let gradients bypass layers via identity shortcuts, preventing them from vanishing in very deep networks.

  2. In a transformer's multi-head attention, why are queries, keys, and values projected into multiple subspaces?

    Answer: To allow the model to attend to information from different representation subspaces simultaneously

    Multiple heads let the model jointly attend to information from different subspaces, capturing diverse relationship types in parallel.

  3. What is the primary purpose of the temperature parameter in a softmax output during inference?

    Answer: To control the sharpness or diversity of the probability distribution

    A lower temperature sharpens the distribution toward the argmax, while a higher temperature flattens it, increasing output diversity.

  4. Which regularization method randomly zeroes entire feature maps (channels) during training in CNNs?

    Answer: Spatial dropout

    Spatial dropout drops entire 2D feature maps, enforcing stronger feature independence than element-wise dropout in convolutional layers.

  5. In a GAN, what condition describes the theoretical equilibrium where the generator perfectly replicates the data distribution?

    Answer: The discriminator outputs 0.5 for all inputs

    At Nash equilibrium the discriminator cannot distinguish real from generated samples, outputting 0.5 (random chance) for every input.

  6. What does 'depthwise separable convolution' decompose a standard convolution into?

    Answer: A depthwise convolution followed by a pointwise (1×1) convolution

    Depthwise separable convolutions apply a single filter per channel then combine channels with 1×1 convolutions, dramatically reducing parameters.

  7. Which loss function is most appropriate for training a neural network on a multi-label classification task where each sample can belong to multiple classes?

    Answer: Binary cross-entropy applied independently per label

    Binary cross-entropy treats each label independently with a sigmoid activation, allowing multiple classes to be simultaneously positive.