Data Serialization Formats Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Data Serialization Formats flashcards as text
Which file format is the default underlying storage format for Delta Lake tables?
Answer: Parquet
Delta Lake uses Parquet as its underlying storage format, adding a transaction log (_delta_log) on top to provide ACID transactions and schema enforcement.
What does 'run-length encoding' (RLE) do in the context of Parquet column compression?
Answer: Replaces consecutive repeated values by storing the value and its repetition count
Run-length encoding replaces consecutive repeated values with a single value and a count (e.g., five 'US' values become 'US×5'), which is very effective for sorted or low-cardinality columns.
Which formats are best suited for storing deeply nested and semi-structured data in a data lake?
Answer: JSON and Parquet (which supports nested schemas)
JSON natively supports nested structures, and Parquet also supports complex nested schemas including arrays, maps, and structs, making both well-suited for semi-structured data in data lakes.
What is the primary purpose of a Schema Registry (such as Confluent Schema Registry) in a streaming data pipeline?
Answer: To centrally manage and validate schemas enabling compatible serialization across producers and consumers
A Schema Registry provides a central repository for managing Avro, JSON Schema, or Protobuf schemas, ensuring producers and consumers use compatible schemas and enabling controlled schema evolution.
When would you choose ORC over Parquet as your data storage format?
Answer: When the primary processing engine is Apache Hive
ORC was optimized specifically for Apache Hive and offers superior performance in Hive-based workloads, while Parquet is generally preferred for Spark and multi-engine environments.
What is 'column pruning' in the context of reading columnar file formats?
Answer: The ability of a query engine to read only the columns referenced by a query, skipping others
Column pruning is the optimization where the query engine reads only the specific columns referenced in a query from the file, skipping all other columns and significantly reducing disk I/O.
Which serialization format has the highest storage overhead per record due to repeating field names with every row?
Answer: JSON
JSON repeats field names as text strings with every record (e.g., {"name":"Alice","age":30} for every row), resulting in significant overhead compared to binary formats that store schemas separately.