โ† All ADE Flashcard Decks

ADE Cloud & Big Data Technologies Flashcards

6 cards from real ADE practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 ADE Cloud & Big Data Technologies flashcards as text
  1. What does the term 'data lakehouse' refer to?

    Answer: An architecture combining the flexibility of data lakes with the management features of data warehouses

    A data lakehouse merges the low-cost storage of data lakes with ACID transactions, schema enforcement, and BI support from data warehouses.

  2. Which Apache technology is used for distributed message streaming and serves as a backbone for real-time data pipelines?

    Answer: Apache Kafka

    Apache Kafka is a distributed event streaming platform used to build real-time data pipelines and streaming applications at scale.

  3. In cloud-based big data architectures, what is 'schema-on-read'?

    Answer: Defining the schema when data is read rather than when it is stored

    Schema-on-read means data is stored in raw form and the schema is applied at query time, providing flexibility for diverse data types.

  4. Which Azure service is used to orchestrate large-scale data movement and transformation pipelines?

    Answer: Azure Data Factory

    Azure Data Factory is a cloud-based ETL and data integration service for creating data-driven workflows to orchestrate data movement and transformation.

  5. What is the primary advantage of using columnar storage over row-based storage for analytical workloads?

    Answer: Improved read performance for aggregate queries on specific columns

    Columnar storage allows analytical queries to read only the columns needed, drastically reducing I/O and improving aggregate query performance.

  6. Which AWS service provides a fully managed ETL service that automatically generates ETL code?

    Answer: AWS Glue

    AWS Glue is a serverless ETL service that discovers data schemas via crawlers and auto-generates PySpark or Scala ETL scripts.