โ† All MS-DS Master of Data science Flashcard Decks

Big Data Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Big Data flashcards as text
  1. Which Apache Flink concept tracks the progress of event time in a stream and triggers time-based operations?

    Answer: Watermarks

    Watermarks are monotonically increasing timestamps injected into the stream to signal that all events up to a certain event time have arrived.

  2. In the context of big data, what does 'data skew' refer to?

    Answer: Uneven distribution of data across partitions causing hotspots

    Data skew occurs when some partitions receive disproportionately more data, causing certain tasks to run much longer than others.

  3. Which storage abstraction in Apache Spark allows it to persist intermediate results and avoid recomputation?

    Answer: Caching/persistence

    Spark's cache() and persist() methods store RDD or DataFrame data in memory (or disk) so reused datasets are not recomputed from scratch.

  4. A Kafka consumer group with 4 consumers subscribing to a topic with 3 partitions will result in:

    Answer: 3 consumers active and 1 idle

    Kafka assigns at most one consumer per partition within a group, so one consumer will be idle when consumers outnumber partitions.

  5. Which technique resolves the problem of small files in HDFS degrading NameNode performance?

    Answer: Sequence files or HAR archives to consolidate small files

    Small files consume NameNode memory proportionally; combining them into Sequence Files or HAR archives reduces the number of metadata entries.

  6. What distinguishes a 'hot path' from a 'cold path' in big data architectures?

    Answer: Hot path processes data in near real-time; cold path processes historical data in batch

    The hot path handles streaming data for immediate insights while the cold path runs batch jobs over stored historical data.

  7. Which NoSQL data model is most appropriate for storing social network relationships with complex graph traversals?

    Answer: Graph database

    Graph databases like Neo4j natively represent entities as nodes and relationships as edges, enabling efficient multi-hop traversals.