Deep Learning and Neural Networks Flashcards
6 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Deep Learning and Neural Networks flashcards as text
Which deep learning architecture is best suited for sequential text classification tasks due to its ability to process inputs in order?
Answer: LSTM
LSTMs maintain a cell state across timesteps, making them well-suited for sequential data where context from earlier inputs affects later predictions.
What is the 'vanishing gradient' problem in deep neural networks?
Answer: Gradients shrinking exponentially as they propagate back through many layers
In deep networks, repeated multiplication of small gradient values during backpropagation causes them to shrink toward zero, preventing early layers from learning.
Which technique involves training a neural network on augmented versions of training data (flips, crops, rotations) to improve generalization?
Answer: Data augmentation
Data augmentation artificially expands the training set by applying label-preserving transformations, reducing overfitting and improving model robustness.
In deep learning, what does a residual (skip) connection accomplish in architectures like ResNet?
Answer: Allows gradients to flow directly through skip paths, easing optimization of very deep networks
Skip connections bypass one or more layers, adding the input directly to the output, which preserves gradient flow and enables training of networks hundreds of layers deep.
What is the purpose of the softmax function in the output layer of a neural network for classification?
Answer: Convert raw logits into a probability distribution that sums to 1
Softmax exponentiates each logit and divides by the sum of all exponentiated logits, producing a valid probability distribution over class labels.
Which hyperparameter controls how much the model's weights are updated in response to the estimated error on each mini-batch?
Answer: Learning rate
The learning rate scales the gradient before it is subtracted from the weights; too high causes divergence, too low causes slow convergence.