← All ADE Flashcard Decks

ADE Cloud & Big Data Technologies Flashcards

6 cards from real ADE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 ADE Cloud & Big Data Technologies flashcards as text
  1. Which AWS service is commonly used as a managed Hadoop and Spark cluster for big data processing?

    Answer: Amazon EMR

    Amazon EMR (Elastic MapReduce) is the AWS managed service for running Apache Hadoop and Spark clusters for big data workloads.

  2. In Apache Spark, what is the primary distributed data structure used for in-memory processing?

    Answer: RDD

    The Resilient Distributed Dataset (RDD) is Spark's fundamental distributed data abstraction enabling fault-tolerant in-memory parallel computation.

  3. Which cloud storage format is optimized for analytical queries due to its columnar storage layout?

    Answer: Parquet

    Parquet is a columnar storage format designed for efficient analytical queries, reducing I/O by reading only relevant columns.

  4. What is the purpose of a data lake in a modern cloud architecture?

    Answer: To serve as a central repository for raw data in any format

    A data lake stores raw data in any format (structured, semi-structured, unstructured) at scale, enabling flexible downstream processing.

  5. Which GCP service provides a fully managed, serverless big data query engine for analyzing data stored in Google Cloud Storage?

    Answer: BigQuery

    BigQuery is Google's fully managed, serverless data warehouse that supports SQL analytics over petabyte-scale datasets stored in GCS.

  6. In Azure, which service is used to ingest, process, and analyze streaming data in real time?

    Answer: Azure Stream Analytics

    Azure Stream Analytics is a real-time analytics service designed to process high-throughput streaming data from sources like IoT devices and event hubs.