ACP Data Engineering & Workflow Automation Flashcards
6 cards from real ACP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 ACP Data Engineering & Workflow Automation flashcards as text
Which Python library is specifically designed for defining, scheduling, and monitoring data pipeline workflows as Directed Acyclic Graphs (DAGs)?
Answer: All of the above
Luigi, Apache Airflow, and Prefect are all Python-native workflow orchestration frameworks that model pipelines as DAGs with scheduling and monitoring capabilities.
In pandas, which method is used to apply a custom function to every row or column of a DataFrame?
Answer: df.apply()
`df.apply()` applies a function along an axis (rows with `axis=1`, columns with `axis=0`), enabling custom transformations across the entire DataFrame.
Which pandas method efficiently removes duplicate rows from a DataFrame, keeping only the first occurrence by default?
Answer: df.drop_duplicates()
`df.drop_duplicates()` returns a DataFrame with duplicate rows removed, with `keep='first'` as default and options for `keep='last'` or `keep=False` to drop all duplicates.
When building a data pipeline, what does the term 'idempotency' mean in the context of pipeline task execution?
Answer: Running a task multiple times produces the same result as running it once
An idempotent pipeline task can be safely re-executed without side effects — running it once or ten times yields the same final state, which is critical for reliable data engineering.
Which pandas method stacks a DataFrame from wide format (one column per variable) to long format (one row per observation)?
Answer: df.melt()
`df.melt()` unpivots a DataFrame from wide to long format by converting specified columns into rows, creating `variable` and `value` columns.
In pandas, what does `groupby()` followed by `agg()` allow you to do?
Answer: Apply multiple aggregation functions to groups simultaneously
`df.groupby('col').agg({'col1': 'sum', 'col2': 'mean'})` groups rows and applies different aggregation functions to different columns in one vectorized operation.