Big Data & Cloud Analytics Flashcards
7 cards from real DAC practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Big Data & Cloud Analytics flashcards as text
Which characteristic distinguishes a 'lakehouse' architecture?
Answer: Combines data lake storage with warehouse-style management and ACID transactions
A lakehouse merges low-cost lake storage with warehouse features like ACID transactions and governance.
What is the primary purpose of Apache Airflow in a data platform?
Answer: Orchestrate and schedule data pipeline workflows
Apache Airflow authors, schedules, and monitors workflow pipelines as directed acyclic graphs (DAGs).
In cloud cost management for analytics, what does 'query cost' in serverless warehouses like BigQuery usually depend on?
Answer: The amount of data scanned by the query
Serverless warehouses typically bill based on the volume of data a query scans.
Which technique reduces data scanned by only reading columns needed for a query?
Answer: Columnar storage / column pruning
Columnar storage lets engines read only the columns referenced, avoiding unneeded data.
What does 'schema-on-read' mean in a big data context?
Answer: Structure is applied when data is queried, not when stored
Schema-on-read defers structure until query time, allowing flexible ingestion of raw data.
Which AWS service lets you run SQL queries directly against data stored in S3 without loading it into a database?
Answer: Amazon Athena
Amazon Athena is a serverless query service that runs SQL directly on data in S3.
In distributed computing, what is 'data skew' and why is it a problem?
Answer: Uneven distribution of data causing some nodes to overload
Data skew is uneven partitioning that overloads some nodes, creating stragglers and slowing jobs.