Big Data Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Big Data flashcards as text
Which Apache Flink concept tracks the progress of event time in a stream and triggers time-based operations?
Answer: Watermarks
Watermarks are monotonically increasing timestamps injected into the stream to signal that all events up to a certain event time have arrived.
In the context of big data, what does 'data skew' refer to?
Answer: Uneven distribution of data across partitions causing hotspots
Data skew occurs when some partitions receive disproportionately more data, causing certain tasks to run much longer than others.
Which storage abstraction in Apache Spark allows it to persist intermediate results and avoid recomputation?
Answer: Caching/persistence
Spark's cache() and persist() methods store RDD or DataFrame data in memory (or disk) so reused datasets are not recomputed from scratch.
A Kafka consumer group with 4 consumers subscribing to a topic with 3 partitions will result in:
Answer: 3 consumers active and 1 idle
Kafka assigns at most one consumer per partition within a group, so one consumer will be idle when consumers outnumber partitions.
Which technique resolves the problem of small files in HDFS degrading NameNode performance?
Answer: Sequence files or HAR archives to consolidate small files
Small files consume NameNode memory proportionally; combining them into Sequence Files or HAR archives reduces the number of metadata entries.
What distinguishes a 'hot path' from a 'cold path' in big data architectures?
Answer: Hot path processes data in near real-time; cold path processes historical data in batch
The hot path handles streaming data for immediate insights while the cold path runs batch jobs over stored historical data.
Which NoSQL data model is most appropriate for storing social network relationships with complex graph traversals?
Answer: Graph database
Graph databases like Neo4j natively represent entities as nodes and relationships as edges, enabling efficient multi-hop traversals.