ADE Cloud & Big Data Technologies Flashcards
6 cards from real ADE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 ADE Cloud & Big Data Technologies flashcards as text
Which AWS service is commonly used as a managed Hadoop and Spark cluster for big data processing?
Answer: Amazon EMR
Amazon EMR (Elastic MapReduce) is the AWS managed service for running Apache Hadoop and Spark clusters for big data workloads.
In Apache Spark, what is the primary distributed data structure used for in-memory processing?
Answer: RDD
The Resilient Distributed Dataset (RDD) is Spark's fundamental distributed data abstraction enabling fault-tolerant in-memory parallel computation.
Which cloud storage format is optimized for analytical queries due to its columnar storage layout?
Answer: Parquet
Parquet is a columnar storage format designed for efficient analytical queries, reducing I/O by reading only relevant columns.
What is the purpose of a data lake in a modern cloud architecture?
Answer: To serve as a central repository for raw data in any format
A data lake stores raw data in any format (structured, semi-structured, unstructured) at scale, enabling flexible downstream processing.
Which GCP service provides a fully managed, serverless big data query engine for analyzing data stored in Google Cloud Storage?
Answer: BigQuery
BigQuery is Google's fully managed, serverless data warehouse that supports SQL analytics over petabyte-scale datasets stored in GCS.
In Azure, which service is used to ingest, process, and analyze streaming data in real time?
Answer: Azure Stream Analytics
Azure Stream Analytics is a real-time analytics service designed to process high-throughput streaming data from sources like IoT devices and event hubs.