Cloud Data Storage Solutions Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Cloud Data Storage Solutions flashcards as text
Which Azure service provides hierarchical namespace support optimized for big-data analytics?
Answer: Azure Data Lake Storage Gen2
ADLS Gen2 adds a hierarchical namespace on top of Blob Storage for analytics workloads.
What is the purpose of partitioning data in cloud object storage for a query engine?
Answer: To prune irrelevant data and reduce scan volume
Partitioning lets query engines skip directories that don't match filter predicates, cutting scan cost.
Which consistency model does Amazon S3 now provide for read-after-write on new objects?
Answer: Strong read-after-write consistency
S3 provides strong read-after-write consistency for PUTs of new objects.
What is a key reason to enable object versioning in a cloud bucket?
Answer: To protect against accidental overwrites and deletions
Versioning retains previous object copies, allowing recovery from unintended changes.
Which open table format adds ACID transactions and time travel on top of a data lake?
Answer: Apache Iceberg
Apache Iceberg (like Delta Lake and Hudi) brings ACID transactions and snapshots to lake storage.
What does 'data egress cost' refer to in cloud storage pricing?
Answer: Cost to transfer data out of the cloud provider's network
Egress charges apply when data leaves the provider's network, often to the internet or another region.
Which approach minimizes small-file problems in a cloud data lake?
Answer: Compacting many small files into larger files
Compaction merges many small files into fewer large files, improving read throughput and reducing overhead.