← All CPA Flashcard Decks

Data Analytics Flashcards

7 cards from real CPA practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Data Analytics flashcards as text
  1. What does ETL stand for in data analytics pipelines?

    Answer: Extract, Transform, Load

    ETL stands for Extract, Transform, Load — the three-phase process of pulling data from sources, reshaping it, and loading it into a target system.

  2. A data analyst wants to identify which rows in a pandas DataFrame contain null values. Which method should they use?

    Answer: df.isnull()

    df.isnull() returns a boolean DataFrame of the same shape where True indicates a missing (null) value.

  3. In statistics, what does a p-value less than 0.05 typically indicate?

    Answer: The result is statistically significant at the 5% level

    A p-value < 0.05 means there is less than a 5% probability of observing the result (or more extreme) if the null hypothesis were true, so we reject it.

  4. Which SQL keyword is used to remove duplicate rows from a query result?

    Answer: DISTINCT

    The DISTINCT keyword in a SELECT statement eliminates duplicate rows, returning only unique combinations of the selected columns.

  5. What type of data is represented by the Python list ['red', 'blue', 'green', 'red']?

    Answer: Nominal categorical data

    Colors are nominal categorical data — they represent distinct categories with no inherent order or numerical meaning.

  6. In a box plot, what do the whiskers typically represent?

    Answer: The range within 1.5 × IQR from the quartiles

    By Tukey's convention, box plot whiskers extend to the farthest data point within 1.5 × IQR beyond Q1 and Q3; points beyond are plotted as outliers.

  7. Which technique splits a dataset into training and testing subsets to evaluate a predictive model's performance?

    Answer: Train-test split

    A train-test split reserves a portion of data (commonly 20-30%) as an unseen test set to measure how well a trained model generalizes.