Big Data & Cloud Analytics Flashcards
7 cards from real DAC practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Big Data & Cloud Analytics flashcards as text
In the Hadoop ecosystem, which component is primarily responsible for resource management and job scheduling?
Answer: YARN
YARN (Yet Another Resource Negotiator) manages cluster resources and schedules jobs in Hadoop.
Which AWS service is a fully managed data warehouse optimized for analytics on petabyte-scale datasets?
Answer: Amazon Redshift
Amazon Redshift is AWS's columnar, petabyte-scale managed data warehouse.
The 'three Vs' of big data traditionally refer to volume, velocity, and which third characteristic?
Answer: Variety
The classic three Vs are volume, velocity, and variety.
Which processing model does Apache Spark use to keep intermediate data in memory for faster iterative computation?
Answer: Resilient Distributed Datasets (RDDs)
Spark's RDD abstraction keeps data in memory, enabling fast iterative processing.
In Google Cloud, which serverless service is designed for running SQL analytics over massive datasets?
Answer: BigQuery
BigQuery is Google Cloud's serverless, highly scalable data warehouse for SQL analytics.
What is the primary purpose of a data lake compared to a traditional data warehouse?
Answer: Store raw data in native formats at scale
A data lake stores raw, unstructured, and structured data in its native format at scale.
Which term describes automatically adjusting cloud compute resources up or down based on workload demand?
Answer: Elastic scaling
Elastic scaling (auto-scaling) adds or removes resources automatically as demand changes.