โ† All DSE Flashcard Decks

DSE - Data Science Deep Learning and Neural Networks Flashcards

6 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 DSE - Data Science Deep Learning and Neural Networks flashcards as text
  1. What does GAN stand for, and what are its two main components?

    Answer: Generative Adversarial Network; generator and discriminator

    A GAN consists of a generator that creates synthetic data and a discriminator that tries to distinguish real from fake data; they are trained adversarially until the generator produces convincing outputs.

  2. In transformer models, what is the role of the 'attention mechanism'?

    Answer: Computing a weighted sum of all positions in a sequence so each token can directly attend to relevant tokens regardless of distance

    The attention mechanism computes relevance scores between all pairs of tokens and takes a weighted combination of their representations, allowing the model to capture long-range dependencies without recurrence.

  3. Which activation function is used in the output layer of a binary classification neural network to produce a probability between 0 and 1?

    Answer: Sigmoid

    The sigmoid function maps any real-valued input to the range (0, 1), making it directly interpretable as a probability for binary classification.

  4. Why is proper weight initialization important in training deep neural networks?

    Answer: Poor initialization can cause gradients to vanish or explode, preventing effective training

    Initializing weights too small causes vanishing gradients while initializing too large causes exploding gradients; methods like Xavier/He initialization set weights to maintain stable gradient flow.

  5. Which data augmentation technique helps neural networks generalize better by injecting small random perturbations into training inputs?

    Answer: Noise injection / data augmentation

    Adding noise (e.g., Gaussian noise, random flips, crops) to training samples artificially expands the dataset and forces the network to learn features that are robust to small variations.

  6. Which of the following is an example of a hyperparameter (not a learned parameter) in a neural network?

    Answer: The learning rate used by the optimizer

    Hyperparameters like learning rate, batch size, and number of layers are set before training and not updated by gradient descent, unlike weights and biases which are learned.