Pearson IT Specialist: Artificial Intelligence (INF-307) — Questions and Answers
Question 1: What does the Transformer's multi-head attention allow compared to single-head attention?
- Faster tokenization
- Processing longer sequences without additional memory
- Attending to different parts of the sequence from multiple representational subspaces simultaneously (Correct answer)
- Using larger batch sizes
Correct answer: Attending to different parts of the sequence from multiple representational subspaces simultaneously
Multi-head attention runs several attention functions in parallel across different learned linear projections, allowing the model to jointly attend to information from different positions and representation subspaces.
Question 2: What is 'data privacy' in the context of AI model training?
- Ensuring that personal data used in training is handled, stored, and used in accordance with privacy laws (Correct answer)
- Removing duplicate records from training data
- Encrypting model weights
- Using only public datasets
Correct answer: Ensuring that personal data used in training is handled, stored, and used in accordance with privacy laws
Data privacy in AI means protecting individuals' personal information used during model training from unauthorized access or misuse.
Question 3: What is semantic segmentation in computer vision?
- Detecting text in images
- Estimating depth from a single image
- Assigning a class label to every pixel in the image without distinguishing between instances (Correct answer)
- Labeling each object instance with a unique ID
Correct answer: Assigning a class label to every pixel in the image without distinguishing between instances
Semantic segmentation classifies every pixel into a category (road, sky, car) but treats all instances of the same class as one region, unlike instance segmentation.
Question 4: What is 'AI alignment' research primarily concerned with?
- Synchronizing distributed training across GPUs
- Aligning model architecture layers properly
- Matching AI predictions to database schemas
- Ensuring AI systems behave in accordance with human values and intentions (Correct answer)
Correct answer: Ensuring AI systems behave in accordance with human values and intentions
AI alignment studies how to ensure advanced AI systems pursue goals that are beneficial and consistent with human values.
Question 5: What is zero-shot classification in the context of large language models?
- Classifying inputs into categories the model has never explicitly been trained on, using only natural language descriptions (Correct answer)
- Running inference without any GPU acceleration
- Removing all zero-value parameters from a model
- Training a model on zero labeled examples then evaluating on training data
Correct answer: Classifying inputs into categories the model has never explicitly been trained on, using only natural language descriptions
Zero-shot classification leverages a model's pre-trained knowledge to assign labels described in natural language without any task-specific training examples.
Question 6: In a generative model like a Variational Autoencoder (VAE) for images, what does the latent space represent?
- A compressed, continuous lower-dimensional representation that captures the essential factors of variation in the data (Correct answer)
- The weights of the convolutional filters
- The loss history during training
- The pixel values of the reconstructed image
Correct answer: A compressed, continuous lower-dimensional representation that captures the essential factors of variation in the data
The VAE's latent space encodes images as probability distributions over a compact representation, from which new samples can be decoded to generate novel images.
Question 7: What is the primary advantage of using Long Short-Term Memory (LSTM) over a simple RNN?
- LSTMs require less memory
- LSTMs do not need backpropagation
- LSTMs are faster to train
- LSTMs can capture long-range dependencies by controlling information flow with gates (Correct answer)
Correct answer: LSTMs can capture long-range dependencies by controlling information flow with gates
LSTM gates (input, forget, output) regulate what information is stored, forgotten, or passed on, mitigating the vanishing gradient problem in long sequences.
Question 8: Which principle of responsible AI states that AI systems should cause minimal harm and consider the well-being of all stakeholders?
- Fairness
- Transparency
- Non-maleficence (Correct answer)
- Accountability
Correct answer: Non-maleficence
Non-maleficence ('do no harm') requires AI systems to avoid causing physical, psychological, financial, or social harm to individuals or society.
Question 9: Which query language is used to retrieve and manipulate data stored in RDF (Resource Description Framework) knowledge graphs?
- SQL
- Cypher
- XPath
- SPARQL (Correct answer)
Correct answer: SPARQL
SPARQL (SPARQL Protocol and RDF Query Language) is the W3C standard query language designed specifically for querying and updating RDF-based knowledge graphs and linked data.
Question 10: What does the 'reward signal' represent in reinforcement learning?
- The probability of transitioning between states
- Feedback from the environment indicating the immediate value or desirability of the agent's last action (Correct answer)
- The total number of states in the environment
- The learning rate of the neural network
Correct answer: Feedback from the environment indicating the immediate value or desirability of the agent's last action
The reward signal provides scalar feedback after each action, guiding the agent toward behaviors that yield higher cumulative reward over time.
Question 11: Which optimizer adapts the learning rate for each parameter based on historical gradient information?
- Momentum
- Adam (Correct answer)
- SGD
- Batch gradient descent
Correct answer: Adam
Adam (Adaptive Moment Estimation) maintains per-parameter adaptive learning rates using estimates of first and second moments of gradients.
Question 12: What is 'value alignment' in AI safety research?
- Calibrating AI confidence scores to match actual accuracy
- The challenge of designing AI systems whose goals and behaviors align with human values and intentions (Correct answer)
- Standardizing AI model file formats across organizations
- Ensuring AI models have consistent numerical outputs
Correct answer: The challenge of designing AI systems whose goals and behaviors align with human values and intentions
Value alignment addresses the problem of ensuring that as AI systems become more capable, their objectives remain consistent with human values rather than pursuing misaligned proxy goals.
Question 13: What does 'fine-tuning' a pre-trained language model involve?
- Pruning the model to reduce its size
- Continuing training on a task-specific labeled dataset to adapt the model to a new task (Correct answer)
- Re-training the model entirely from scratch on new data
- Converting the model to a different programming language
Correct answer: Continuing training on a task-specific labeled dataset to adapt the model to a new task
Fine-tuning continues gradient-based training of a pre-trained model on a smaller task-specific dataset, adapting general representations to the target task.
Question 14: What is tokenization in natural language processing?
- Measuring the sentiment of a sentence
- Converting text to uppercase
- Splitting text into meaningful units such as words or subwords (Correct answer)
- Removing stop words from a document
Correct answer: Splitting text into meaningful units such as words or subwords
Tokenization breaks raw text into tokens (words, subwords, or characters) that serve as the basic units for NLP model input.
Question 15: What is 'model transparency' in responsible AI?
- Ensuring the model has no parameters
- Making model weights publicly downloadable
- Publishing all training code without exception
- Openly documenting how a model works, what data it was trained on, and its limitations (Correct answer)
Correct answer: Openly documenting how a model works, what data it was trained on, and its limitations
Model transparency involves disclosing key information about the model's design, training data, and performance to enable accountability.
Question 16: Which metric is most appropriate when false negatives are more costly than false positives?
- Recall (Correct answer)
- Precision
- Accuracy
- Specificity
Correct answer: Recall
Recall (sensitivity) measures the proportion of actual positives correctly identified, minimizing false negatives.
Question 17: Which of the following is an example of unsupervised learning?
- Spam email classification
- House price prediction
- Customer clustering by purchase behavior (Correct answer)
- Handwritten digit recognition with labels
Correct answer: Customer clustering by purchase behavior
Clustering groups unlabeled data by similarity, requiring no predefined output labels.
Question 18: What is 'reward shaping' in reinforcement learning?
- Using imitation learning to replace the reward function
- Adding supplementary reward signals to guide the agent toward desired behaviors when the natural reward is sparse or delayed (Correct answer)
- Normalizing reward values to fall between -1 and 1
- Changing the discount factor during training
Correct answer: Adding supplementary reward signals to guide the agent toward desired behaviors when the natural reward is sparse or delayed
Reward shaping augments the environment's sparse reward with additional dense rewards that reflect progress toward the goal, speeding up learning without changing the optimal policy.
Question 19: What is 'hallucination' in large language models?
- Repeating the same token endlessly
- Generating confident but factually incorrect or fabricated information (Correct answer)
- The model generating images during text tasks
- A training instability that causes random outputs
Correct answer: Generating confident but factually incorrect or fabricated information
LLM hallucination occurs when the model produces plausible-sounding but false or invented information with unwarranted confidence.
Question 20: What is coreference resolution in NLP?
- Determining the sentiment of each sentence
- Translating pronouns across languages
- Identifying all expressions in a text that refer to the same real-world entity (Correct answer)
- Detecting duplicate sentences in a document
Correct answer: Identifying all expressions in a text that refer to the same real-world entity
Coreference resolution links pronouns and noun phrases that refer to the same entity, enabling coherent understanding of who or what is being discussed.
Question 21: What does the NIST AI Risk Management Framework (AI RMF) provide to US organizations?
- A voluntary framework for identifying, assessing, and managing risks in AI systems (Correct answer)
- Federal regulations for AI certification
- A standard AI training curriculum for employees
- A mandatory AI licensing program
Correct answer: A voluntary framework for identifying, assessing, and managing risks in AI systems
The NIST AI RMF provides a voluntary, flexible framework to help organizations manage AI risks across the AI lifecycle.
Question 22: In a decision tree, what does 'pruning' accomplish?
- Boosts minority class samples
- Normalizes feature values
- Removes branches to reduce overfitting (Correct answer)
- Adds more branches to improve accuracy
Correct answer: Removes branches to reduce overfitting
Pruning removes branches that provide little power, simplifying the tree to improve generalization.
Question 23: In a seq2seq model for machine translation, what is the role of the encoder?
- To score candidate translations
- To compress the source sentence into a context representation (Correct answer)
- To generate the translated output token by token
- To perform beam search over possible outputs
Correct answer: To compress the source sentence into a context representation
The encoder processes the source sequence and produces a fixed-size context vector (or sequence of hidden states) that summarizes its meaning for the decoder.
Question 24: What is a 'language model' in NLP?
- A model that assigns probabilities to sequences of words or predicts the next word in a sequence (Correct answer)
- A model that translates between languages
- A grammar checker for written documents
- A classification model for identifying topics
Correct answer: A model that assigns probabilities to sequences of words or predicts the next word in a sequence
A language model learns the probability distribution over sequences of words, enabling tasks like text generation, completion, and scoring sentence fluency.
Question 25: What does precision measure in a classification model?
- The total accuracy across all classes
- The speed of inference
- The fraction of positive predictions that were correct (Correct answer)
- The fraction of actual positives that were found
Correct answer: The fraction of positive predictions that were correct
Precision is true positives divided by all positive predictions, measuring prediction correctness.
Question 26: What is batch normalization designed to do?
- Apply dropout to each batch
- Convert labels into one-hot encodings
- Increase the batch size during training
- Normalize layer inputs to speed up and stabilize training (Correct answer)
Correct answer: Normalize layer inputs to speed up and stabilize training
Batch normalization normalizes activations within a mini-batch, reducing internal covariate shift and allowing higher learning rates.
Question 27: Which algorithm is a non-parametric method that classifies new points based on the majority class of their nearest neighbors?
- Naive Bayes
- Linear Regression
- K-Nearest Neighbors (Correct answer)
- Support Vector Machine
Correct answer: K-Nearest Neighbors
K-Nearest Neighbors classifies a point by looking at the k closest training examples and using a majority vote.
Question 28: Case-based reasoning (CBR) in AI solves new problems by:
- Training a neural network on historical data
- Randomly sampling candidate solutions and selecting the best one
- Deriving solutions from first-principles logical axioms
- Retrieving and adapting solutions from similar past cases stored in memory (Correct answer)
Correct answer: Retrieving and adapting solutions from similar past cases stored in memory
CBR follows a retrieve-reuse-revise-retain cycle: it finds past cases similar to the current problem, adapts their solutions, evaluates the result, and stores the new case for future use.
Question 29: Which of the following is a primary goal of Artificial Intelligence research?
- Coupling
- Reasoning (Correct answer)
- Data
- Mastering
Correct answer: Reasoning
Reasoning is a primary goal of Artificial Intelligence research because it involves enabling machines to draw inferences, make decisions, and solve problems logically, similar to human cognitive processes. AI systems strive to mimic or surpass human reasoning abilities to interpret information, understand contexts, and generate appropriate responses or actions. This capability is fundamental to developing intelligent agents that can operate autonomously and effectively.
Question 30: What does data augmentation in computer vision typically involve?
- Compressing training images to reduce disk usage
- Applying random transformations like flips, rotations, and crops to training images to improve generalization (Correct answer)
- Converting images to grayscale before training
- Collecting additional labeled images from the internet
Correct answer: Applying random transformations like flips, rotations, and crops to training images to improve generalization
Data augmentation artificially expands training data by applying label-preserving transformations, helping models become robust to variations in orientation, scale, and lighting.
Question 31: What is transfer learning in deep learning?
- Transferring data between cloud storage providers
- Training a model on multiple datasets simultaneously
- Using a pre-trained model's learned weights as a starting point for a new task (Correct answer)
- Moving training between different hardware
Correct answer: Using a pre-trained model's learned weights as a starting point for a new task
Transfer learning reuses weights from a model trained on a large dataset (e.g., ImageNet) as initialization for a related task, reducing training time and data requirements.
Question 32: Which word embedding model learns vector representations by predicting surrounding words (skip-gram) or predicting a word from context (CBOW)?
- Word2Vec (Correct answer)
- GloVe
- BERT
- FastText
Correct answer: Word2Vec
Word2Vec offers two architectures: skip-gram predicts context words from a target, and CBOW predicts a target word from its context window.
Question 33: What is an 'adversarial example' in the context of AI security?
- A training sample with an incorrect label
- An example used to evaluate model robustness
- An input deliberately modified with small perturbations to cause an AI model to make wrong predictions (Correct answer)
- A competitor's AI product
Correct answer: An input deliberately modified with small perturbations to cause an AI model to make wrong predictions
Adversarial examples are inputs crafted by adding carefully chosen, often imperceptible perturbations that reliably fool AI classifiers into producing incorrect outputs.
Question 34: What does 'fairness' in AI typically require?
- Models that are free from any statistical patterns
- Equitable treatment and outcomes for individuals across demographic groups (Correct answer)
- Identical outputs for all inputs
- Equal model accuracy across all users regardless of group
Correct answer: Equitable treatment and outcomes for individuals across demographic groups
AI fairness aims to ensure that model decisions do not systematically disadvantage individuals based on protected attributes.
Question 35: Which architecture introduced the concept of encoder-decoder with skip connections, widely used for image segmentation?
- VGG16
- ResNet-50
- U-Net (Correct answer)
- AlexNet
Correct answer: U-Net
U-Net uses a contracting encoder path and an expansive decoder path with skip connections between corresponding layers, enabling precise pixel-level segmentation.
Question 36: Fuzzy logic is used in AI primarily to:
- Represent and reason about degrees of truth in situations involving uncertainty or vagueness (Correct answer)
- Handle binary true/false decisions more efficiently
- Speed up gradient computation in neural networks
- Encrypt knowledge base entries for security
Correct answer: Represent and reason about degrees of truth in situations involving uncertainty or vagueness
Fuzzy logic extends classical binary logic by allowing truth values between 0 and 1, making it suitable for reasoning about imprecise or vague concepts like 'tall', 'warm', or 'fast'.
Question 37: Which algorithm is best described as finding the hyperplane that maximizes the margin between two classes?
- Decision tree
- Logistic regression
- Support Vector Machine (Correct answer)
- K-means
Correct answer: Support Vector Machine
Support Vector Machines find the optimal separating hyperplane by maximizing the margin between the closest data points of each class.
Question 38: What is the primary difference between extractive and abstractive text summarization?
- Extractive selects existing sentences from the source; abstractive generates new sentences (Correct answer)
- Extractive is always shorter than abstractive summaries
- Extractive uses neural networks; abstractive uses rule-based systems
- Extractive summarization works only on emails; abstractive on articles
Correct answer: Extractive selects existing sentences from the source; abstractive generates new sentences
Extractive summarization picks and ranks existing sentences from the document, while abstractive summarization generates novel sentences that may not appear verbatim in the source.
Question 39: Which technique allows a language model to answer questions about a document it wasn't trained on by providing that document as context?
- Retrieval-augmented generation (Correct answer)
- Model distillation
- Knowledge distillation
- Prompt tuning
Correct answer: Retrieval-augmented generation
Retrieval-augmented generation (RAG) retrieves relevant documents at inference time and includes them in the model's context, enabling factual answers beyond the model's training knowledge.
Question 40: What does the subword tokenization algorithm BPE (Byte Pair Encoding) do?
- Iteratively merges the most frequent character pairs to build a vocabulary of subword units (Correct answer)
- Splits text only on whitespace
- Removes all punctuation from text
- Converts all tokens to a fixed embedding size
Correct answer: Iteratively merges the most frequent character pairs to build a vocabulary of subword units
BPE starts with individual characters and repeatedly merges the most frequent adjacent pair until a target vocabulary size is reached, balancing between word and character tokenization.
Question 41: What is transfer learning in the context of neural networks?
- Transferring gradients between layers
- Training from scratch on new data
- Reusing a pre-trained model's weights as a starting point for a new task (Correct answer)
- Compressing a model into fewer parameters
Correct answer: Reusing a pre-trained model's weights as a starting point for a new task
Transfer learning leverages weights learned on a large dataset (e.g., ImageNet) and fine-tunes them for a different task.
Question 42: What is 'accountability' in responsible AI governance?
- Ensuring that developers, deployers, and organizations can be held responsible for the outcomes and harms caused by AI systems (Correct answer)
- Publishing benchmark results for all AI models
- Logging every model inference for auditing
- Tracking model training costs
Correct answer: Ensuring that developers, deployers, and organizations can be held responsible for the outcomes and harms caused by AI systems
Accountability means that identifiable parties are responsible for AI system outcomes, can be questioned about their decisions, and bear consequences when systems cause harm.
Question 43: What is the 'right to explanation' under GDPR in the context of automated AI decisions?
- The right to opt out of all AI recommendations
- The right to delete your data from AI training sets
- The right of individuals to receive meaningful information about the logic behind automated decisions that significantly affect them (Correct answer)
- The right to request a copy of your personal data
Correct answer: The right of individuals to receive meaningful information about the logic behind automated decisions that significantly affect them
GDPR Article 22 grants individuals the right not to be subject to solely automated decisions with significant effects, and to obtain meaningful explanations of the decision logic.
Question 44: Which NLP task determines whether the relationship between two sentences is entailment, contradiction, or neutral?
- Coreference resolution
- Natural language inference (Correct answer)
- Question answering
- Semantic role labeling
Correct answer: Natural language inference
Natural language inference (NLI) classifies the logical relationship between a premise and a hypothesis sentence as entailment, contradiction, or neutral.
Question 45: Which US agency is primarily responsible for regulating AI in consumer financial products?
- FTC (Federal Trade Commission)
- CFPB (Consumer Financial Protection Bureau) (Correct answer)
- FDA
- NIST
Correct answer: CFPB (Consumer Financial Protection Bureau)
The CFPB oversees AI-driven credit scoring and lending decisions to ensure they comply with fair lending laws.
Pearson IT Specialist: Artificial Intelligence (INF-307)
The Pearson IT Specialist Artificial Intelligence exam (INF-307) validates foundational knowledge and practical skills in AI, covering problem definition, data engineering, AI algorithms and models, application deployment, and monitoring AI systems in production.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds