โ† All AI Flashcard Decks

AI Engineer: Neural Networks and Deep Learning Flashcards

6 cards from real AI practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 AI Engineer: Neural Networks and Deep Learning flashcards as text
  1. What is the vanishing gradient problem in deep neural networks?

    Answer: Gradients become extremely small, making early layers learn very slowly

    Vanishing gradients occur when backpropagated gradients shrink exponentially through layers, causing weights in early layers to update negligibly.

  2. Which activation function is most commonly used in hidden layers of modern deep neural networks to mitigate the vanishing gradient problem?

    Answer: ReLU (Rectified Linear Unit)

    ReLU outputs max(0, x), maintaining gradient magnitude for positive inputs and largely avoiding vanishing gradients compared to sigmoid or tanh.

  3. What is the purpose of batch normalization in a neural network?

    Answer: To normalize activations within a layer to stabilize and accelerate training

    Batch normalization normalizes layer inputs to zero mean and unit variance, reducing internal covariate shift and enabling higher learning rates.

  4. In a convolutional neural network (CNN), what does 'stride' control?

    Answer: How many pixels the filter moves at each step

    Stride determines how many pixels the convolutional filter shifts at each step, controlling the spatial dimensions of the output feature map.

  5. What is dropout regularization in neural networks?

    Answer: Randomly setting a fraction of neurons to zero during training to prevent overfitting

    Dropout randomly deactivates neurons during each training step, forcing the network to learn redundant representations and reducing overfitting.

  6. Which optimizer adapts the learning rate for each parameter individually using estimates of first and second moments of gradients?

    Answer: Adam

    Adam combines adaptive learning rates with momentum, maintaining per-parameter moving averages of gradients and their squares.