Lean Six Sigma Black Belt Measure Phase: Data Analysis Questions and Answers — Questions and Answers
Question 1: A Black Belt is analyzing data from a manufacturing process to establish a performance baseline. The dataset is large and appears to have multiple sources of variation. To better understand the underlying patterns, the Black Belt decides to separate the data based on shift, machine, and operator. What is this data analysis technique called?
- Regression Analysis
- Stratification (Correct answer)
- Gage R&R
- Hypothesis Testing
Correct answer: Stratification
Stratification is the technique of dividing data into distinct sub-groups or layers (strata) to identify patterns or sources of variation that might otherwise be obscured when looking at the data as a whole. By separating the data by shift, machine, and operator, the Black Belt can analyze the performance of each individual stratum to pinpoint specific areas contributing to variation.
Question 2: In the Measure phase, a project team conducted a Gage R&R study to assess the reliability of their measurement system. The study revealed a high percentage of variation attributable to reproducibility. What is the MOST likely cause of this issue?
- The measurement instrument requires calibration.
- There is significant variation within the measurements taken by a single appraiser.
- The operational definitions for taking the measurement are unclear or inconsistently applied by different appraisers. (Correct answer)
- The process being measured is naturally unstable.
Correct answer: The operational definitions for taking the measurement are unclear or inconsistently applied by different appraisers.
Reproducibility refers to the variation in the average of the measurements made by different appraisers using the same measuring instrument when measuring the identical characteristic on the same part. High reproducibility indicates that the appraisers are not consistent with each other, which is often due to inadequate training or unclear operational definitions and procedures.
Question 3: A Six Sigma team is evaluating a process's ability to meet customer specifications. They calculate a Cpk of 1.5 and a Cp of 2.0. What can be concluded from these process capability indices?
- The process is not capable of meeting specifications.
- The process is perfectly centered between the specification limits.
- The process has the potential to be highly capable, but it is currently off-center.
- The process variation is too wide compared to the specification limits. (Correct answer)
Correct answer: The process variation is too wide compared to the specification limits.
Cp measures the potential capability by comparing the process spread to the specification spread, while Cpk accounts for the process centering. A Cp of 2.0 indicates the process spread is half the width of the specifications, suggesting high potential capability. However, since Cpk (1.5) is less than Cp (2.0), it signifies that the process mean is not centered between the specification limits, which reduces the actual capability.
Question 4: Which of the following graphical tools is most appropriate for visualizing the distribution, central tendency, and spread of a continuous dataset, and for identifying potential outliers?
- Pareto Chart
- Scatter Plot
- Control Chart
- Box Plot (Correct answer)
Correct answer: Box Plot
A Box Plot (or Box and Whisker Plot) is specifically designed to display the five-number summary of a set of data: minimum, first quartile, median, third quartile, and maximum. This makes it an excellent tool for quickly understanding the distribution's central tendency (median), spread (interquartile range), and identifying potential outliers which appear as individual points beyond the 'whiskers'.
Question 5: A Black Belt is preparing to analyze process data that has been collected. Many statistical tools assume the data follows a normal distribution. Which statistical test should be used to formally verify this assumption?
- t-test
- ANOVA
- Anderson-Darling test (Correct answer)
- Chi-Square test
Correct answer: Anderson-Darling test
The Anderson-Darling test is a statistical test used to determine if a given sample of data is drawn from a specific distribution, most commonly the normal distribution. In the Measure phase, it's a critical step before performing capability analysis or other statistical tests that require normally distributed data. A p-value from the test helps to objectively decide if the data fits the normal distribution.
Question 6: During the data analysis portion of the Measure phase, a team creates a scatter plot that shows a strong positive correlation between ambient temperature in a factory and the number of product defects. What is the next logical step for the team to take?
- Immediately implement a new cooling system for the factory.
- Conclude that temperature is the definitive root cause of the defects.
- Formulate a hypothesis to test the causal relationship between temperature and defects in the Analyze phase. (Correct answer)
- Discard the data as it only shows correlation, not causation.
Correct answer: Formulate a hypothesis to test the causal relationship between temperature and defects in the Analyze phase.
While a strong correlation suggests a relationship, it does not prove causation. The Measure phase is about collecting data and identifying potential relationships. The appropriate next step is to move into the Analyze phase, where this observed relationship can be formally stated as a hypothesis and tested through more rigorous statistical methods or designed experiments to determine if a true cause-and-effect relationship exists.
A Black Belt is analyzing data from a manufacturing process to establish a performance baseline.
The dataset is large and appears to have multiple sources of variation.
To better understand the underlying patterns, the Black Belt decides to separate the data based on shift, machine, and operator.
What is this data analysis technique called?