โ† All Data Engineering Flashcard Decks

Cloud Data Storage Solutions Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Cloud Data Storage Solutions flashcards as text
  1. Which AWS storage class is most cost-effective for archival data accessed less than once a year with retrieval times of hours?

    Answer: S3 Glacier Deep Archive

    S3 Glacier Deep Archive offers the lowest storage cost for rarely accessed long-term archives.

  2. What is the primary benefit of object storage over block storage for a data lake?

    Answer: Massive scalability with rich metadata

    Object storage scales to virtually unlimited capacity and attaches metadata to each object, ideal for data lakes.

  3. In Google Cloud Storage, which feature automatically transitions objects between storage classes based on age?

    Answer: Lifecycle Management

    Lifecycle Management applies rules to move or delete objects as they age, optimizing cost.

  4. Which storage option is best suited for a high-throughput streaming workload requiring low-latency appends?

    Answer: A managed message/log store like Kafka or Kinesis

    Log-based streaming stores are designed for high-throughput, low-latency append operations.

  5. What does 'eventual consistency' mean in a distributed cloud object store?

    Answer: Updates propagate over time and reads may briefly return stale data

    Eventual consistency means replicas converge to the latest value after a propagation delay.

  6. Which factor most directly increases the cost of querying data stored in S3 via Amazon Athena?

    Answer: Amount of data scanned per query

    Athena charges based on the volume of data scanned, so partitioning and columnar formats reduce cost.

  7. What is the main advantage of using a columnar format like Parquet in cloud storage?

    Answer: Reduced scan size and better compression for analytics

    Parquet stores data by column, enabling column pruning and high compression for analytical queries.