โ† All DSE Flashcard Decks

Big Data Technologies Flashcards

7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Big Data Technologies flashcards as text
  1. What is the role of Apache ZooKeeper in a big data cluster?

    Answer: Providing distributed coordination and configuration management services

    ZooKeeper provides distributed coordination services such as leader election, configuration management, and distributed locking for cluster components.

  2. In Apache Flink, what distinguishes it from Apache Spark Streaming?

    Answer: Flink is a true streaming engine where batch is a special case, whereas Spark Streaming uses micro-batches

    Flink treats streaming as the first-class primitive and handles batch as a finite stream, while Spark Streaming originally processed streams as a series of small batches.

  3. What does 'replication factor' mean in HDFS?

    Answer: The number of copies of each data block stored across DataNodes

    The replication factor (default 3) defines how many copies of each HDFS block are written to different DataNodes for fault tolerance.

  4. Which pattern best describes the Lambda Architecture for big data systems?

    Answer: Combining batch and speed layers to serve both historical and real-time queries

    Lambda Architecture uses a batch layer for comprehensive historical processing and a speed layer for real-time updates, with a serving layer merging both views.

  5. What is Apache Sqoop primarily used for?

    Answer: Bulk transfer of data between relational databases and Hadoop

    Sqoop is designed to efficiently transfer bulk data between relational databases (via JDBC) and Hadoop storage systems like HDFS or HBase.

  6. In the context of big data, what is 'schema-on-read'?

    Answer: Applying the schema interpretation only when data is queried or read

    Schema-on-read defers schema enforcement to query time, allowing raw data to be stored in any format and interpreted flexibly when accessed.

  7. What is the purpose of a Kafka offset?

    Answer: A sequential identifier that tracks a consumer's read position within a partition

    An offset is a unique sequential number assigned to each message in a partition, allowing consumers to track and resume their read position.