← All MS-DS Master of Data science Flashcard Decks

FREE MS-DS Master of Data science Big Data Technologies Questions and Answers Flashcards

6 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 FREE MS-DS Master of Data science Big Data Technologies Questions and Answers flashcards as text
  1. What is the primary purpose of Apache ZooKeeper in a big data ecosystem?

    Answer: Distributed coordination and configuration management

    Apache ZooKeeper provides centralized services for maintaining configuration information, naming, distributed synchronization, and group services across distributed applications.

  2. In Apache Kafka, what is a partition's role in achieving scalability for data streaming?

    Answer: It allows a topic's data to be distributed across multiple brokers for parallel consumption

    Kafka partitions divide a topic's data across multiple brokers, enabling parallel reads and writes that scale horizontally with the number of consumers.

  3. Which data processing paradigm does Apache Flink natively support that distinguishes it from batch-first frameworks?

    Answer: True event-time stream processing with exactly-once semantics

    Apache Flink is built as a stream-first processing engine that natively handles event-time semantics and provides exactly-once state consistency guarantees.

  4. What is the function of the YARN ResourceManager in a Hadoop cluster?

    Answer: It allocates cluster resources to competing applications

    The YARN ResourceManager is the master daemon that arbitrates resources among all applications running in the cluster by managing NodeManagers and scheduling containers.

  5. Which technique does Apache Spark use to recover lost data partitions without requiring full data replication?

    Answer: Lineage-based recomputation from the DAG of transformations

    Spark tracks the lineage of transformations used to build each RDD, allowing it to recompute only the lost partitions from the original data source rather than storing redundant copies.

  6. In a data lake architecture, what is the primary advantage of using a schema-on-read approach over schema-on-write?

    Answer: Data can be ingested in its raw format and structured at query time for different use cases

    Schema-on-read allows raw data to be stored without predefined structure, enabling different consumers to apply their own schemas when reading, which maximizes flexibility for diverse analytical needs.