Deep Learning and Neural Networks Flashcards
6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Deep Learning and Neural Networks flashcards as text
What is the 'attention mechanism' in deep learning primarily used for?
Answer: Allowing the model to focus on relevant parts of the input when producing each output
Attention computes a weighted sum of input representations, letting the model dynamically focus on the most relevant tokens or features for each prediction.
In deep learning, what is a hyperparameter?
Answer: A configuration value set before training that controls the learning process
Hyperparameters such as learning rate, batch size, and number of layers are set before training and are not updated by the optimization algorithm.
What problem does the ReLU activation function help solve compared to sigmoid?
Answer: Vanishing gradients in deep networks by maintaining non-zero gradients for positive inputs
ReLU outputs the input directly for positive values, providing a constant gradient of 1 and mitigating the vanishing gradient problem that plagues sigmoid.
What distinguishes a deep neural network from a shallow one?
Answer: Deep networks contain multiple hidden layers enabling hierarchical feature learning
Depth allows hierarchical composition of features — early layers detect edges, middle layers detect shapes, and later layers detect high-level concepts.
Which technique is used to prevent overfitting by penalizing large weights in a neural network's loss function?
Answer: L2 regularization (weight decay)
L2 regularization adds a penalty proportional to the sum of squared weights to the loss, discouraging large weight values and reducing overfitting.
In a neural network, what is the purpose of the softmax function in the output layer for classification tasks?
Answer: To convert raw logits into a probability distribution that sums to 1
Softmax exponentiates each logit and divides by the sum of all exponentiated logits, producing normalized probabilities that represent class likelihoods.