← All Data Engineering Flashcard Decks

Real-Time Streaming Architectures Flashcards

6 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Real-Time Streaming Architectures flashcards as text
  1. A financial services company wants to monitor credit card transactions for fraudulent activity. The system must group all transactions by a specific user that occur closely together, but the time between these bursts of activity is unpredictable. A new group should start only after a significant period of user inactivity. Which windowing strategy is most suitable for this scenario?

    Answer: Session Windows

    Session windows are designed specifically for this use case. They group events based on periods of activity, which are terminated by a predefined gap of inactivity (a timeout). This allows the system to dynamically create windows for each user's transaction burst without having fixed start or end times.

  2. In a real-time streaming architecture that processes data using event time, what is the primary function of a watermark?

    Answer: To signal to the processing engine a point in time beyond which it is unlikely to receive older events, allowing it to close windows and trigger computations.

    A watermark is a mechanism that tracks the progress of event time in a stream. It provides a heuristic for the system to know when it has likely received all the events for a particular time window, enabling it to finalize calculations for that window even if some data arrives out of order or late. This balances the need for accuracy with the need to produce timely results and manage state.

  3. A data engineering team is building a real-time dashboard to display the number of unique visitors on their website over the last 15 minutes. Which type of stream processing is fundamentally required to implement this feature?

    Answer: Stateful processing

    Calculating the number of unique visitors requires the system to remember which visitor IDs it has already seen within the current time window. This memory of past events is known as 'state'. Therefore, this is a stateful processing operation. A stateless operation processes each event independently without knowledge of previous events.

  4. Which of the following is a defining characteristic of the Kappa Architecture compared to the Lambda Architecture?

    Answer: It simplifies the architecture by using a single stream-processing engine to handle both real-time queries and historical data reprocessing from an append-only log.

    The Kappa Architecture was proposed as a simplification of the Lambda Architecture. Its core principle is to eliminate the batch layer and use a single, unified stream processing pipeline for all tasks. Historical analysis is achieved by replaying the data from a canonical, immutable log (like Apache Kafka) through the same streaming engine.

  5. In a streaming data pipeline, a processing stage is consistently unable to keep up with the rate of data it receives from an upstream stage. This leads to growing memory usage, increased latency, and potential system instability. What is the term for the mechanism designed to mitigate this specific problem?

    Answer: Backpressure

    Backpressure is a flow control mechanism where a slower downstream consumer can signal to a faster upstream producer to reduce the rate of data transfer. This prevents the consumer from being overwhelmed, which could lead to buffer overflows, data loss, or system failure.

  6. An IoT application collects sensor readings from thousands of devices. Each reading is tagged with the precise timestamp when the measurement was taken. Due to variable network conditions, these readings arrive at the central processing system out of order and with significant delays. For accurate time-series analysis (e.g., calculating hourly averages), which time characteristic must the system rely on?

    Answer: Event Time

    Event Time refers to the timestamp embedded within the data record itself, indicating when the event actually occurred in the real world. Using Event Time is crucial for correctness when data can arrive late or out of order, as it ensures that analysis reflects the true sequence of events, not the sequence in which they were processed. Processing Time is the time on the machine executing the job and would lead to inaccurate results in this scenario.