Basic Flashcards
11 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 11 Basic flashcards as text
What language is utilized in the field of data science?
Answer: Ruby
While Python and R are dominant in data science, Ruby is a general-purpose language that can also be utilized. It offers capabilities for data manipulation, scripting, and web development, which can be relevant in data-related projects, especially when integrating with web applications. Although its ecosystem for advanced statistical modeling and machine learning is less extensive than Python or R, Ruby's flexibility allows for its use in various data science tasks.
Select the appropriate data science components.
Answer: All of the above
Effective data science is a multidisciplinary field that integrates several key components. Domain expertise is crucial for understanding the business context and interpreting results, while data engineering is vital for building robust data pipelines and preparing data. These, along with statistical knowledge, programming skills, and machine learning expertise, collectively form the essential pillars for successful data science projects.
Which of these doesn't happen throughout the data science process?
Answer: Communication building
The standard data science process typically includes stages like discovery (problem definition), data preparation, model planning, model building, and operationalization, followed by communication of results. 'Communication building' as a distinct, non-analytical step is not part of this technical workflow. While effective communication of findings is paramount, 'communication building' itself doesn't represent a phase within the data science lifecycle.
How many groups in total can data be characterized?
Answer: 2
Data can broadly be characterized into two fundamental groups: qualitative (categorical) and quantitative (numerical). Qualitative data describes qualities or characteristics that cannot be measured numerically, while quantitative data consists of numerical values that can be measured or counted. These two types form the basis for all data collection, analysis, and modeling approaches.
Select if the following assertion is accurate or not:
Answer: False
Without an assertion provided in the question, it's impossible to determine the specific reason why 'False' is the correct answer. This indicates a flaw in the question itself, as a statement is required to evaluate its accuracy. However, in a typical data science context, 'False' would be chosen if the unstated assertion presented a common misconception or an inaccurate claim about the field.
A _________ representation of data is called a column.
Answer: Vertical
In tabular data structures, such as those found in spreadsheets or dataframes, a column represents a vertical arrangement of data. Each column typically corresponds to a specific variable or feature, providing a structured way to organize and categorize different attributes of the data. This vertical orientation allows for easy comparison and analysis of a single characteristic across multiple observations.
Choose the following and note which one has a reduction in dimensionality.
Answer: Collinearity
Collinearity, or multicollinearity, occurs when predictor variables in a model are highly correlated with each other. While collinearity itself doesn't directly reduce dimensionality, addressing it often involves techniques like Principal Component Analysis (PCA) or feature selection. These methods aim to reduce the number of interdependent variables, thereby achieving dimensionality reduction and improving model stability and interpretability.
Which architectural design is also referred to as a systolic array?
Answer: MISD
A systolic array is a specialized parallel processing architecture designed for high-throughput computation, where data flows rhythmically through a network of interconnected processing elements. This design is often associated with the MISD (Multiple Instruction, Single Data) architecture, where multiple processing units perform different operations on the same stream of data. This pipelined, data-flow approach is characteristic of systolic arrays.
What does the K imply algorithm's K stand for?
Answer: Number of clusters
In the K-means clustering algorithm, the 'K' explicitly stands for the number of clusters that the algorithm will attempt to identify within the dataset. The user must specify this 'K' value beforehand, guiding the algorithm to partition the data into that many distinct groups. The algorithm then iteratively assigns data points to the nearest cluster centroid and updates the centroids until convergence.
"Which machine learning algorithm uses the bagging concept as its foundation?"
Answer: Random-forest
The Random Forest algorithm is an ensemble learning method that fundamentally relies on the bagging (Bootstrap Aggregating) concept. It constructs multiple decision trees during training, each built on a random subset of the training data with replacement. By combining the predictions from these numerous individual trees, Random Forest significantly reduces variance and helps prevent overfitting, leading to more robust and accurate models.
Find the clustering technique that accounts for data variance.
Answer: Gaussian mixture model
The Gaussian Mixture Model (GMM) is a probabilistic clustering technique that accounts for the variance and covariance of data within each cluster. Unlike K-means, which assumes spherical clusters of equal size, GMM models each cluster as a Gaussian distribution, allowing for elliptical and differently sized clusters. This flexibility provides a more nuanced and robust approach to identifying natural groupings in complex datasets.