Auto Loader and Data Ingestion Flashcards
7 cards from real Databricks Certified Data Engineer Associate practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Auto Loader and Data Ingestion flashcards as text
What is the primary advantage of using Auto Loader over `spark.read` for cloud storage ingestion?
Answer: It automatically identifies and processes only new files incrementally
Auto Loader incrementally identifies and processes only new files as they arrive in cloud storage, avoiding reprocessing of previously seen files.
Which Spark source format string is used to configure Auto Loader in a readStream call?
Answer: cloudFiles
Auto Loader is accessed via `spark.readStream.format('cloudFiles')`, which is the Databricks-specific streaming source for cloud file ingestion.
What is the purpose of the `cloudFiles.schemaLocation` option in Auto Loader?
Answer: Stores the inferred schema so it persists and evolves across stream restarts
`cloudFiles.schemaLocation` tells Auto Loader where to persist the inferred schema on cloud storage, allowing it to reuse and evolve the schema across multiple stream restarts.
What is the key difference between Auto Loader's 'directory listing' mode and 'file notification' mode?
Answer: File notification uses cloud services (e.g., AWS SNS/SQS) for real-time updates; directory listing polls the directory periodically
File notification mode sets up cloud event infrastructure (like AWS SNS + SQS) so new file arrivals trigger processing immediately, while directory listing scans the directory on each trigger.
When Auto Loader detects new columns in incoming files with schema evolution enabled, which option controls this behavior?
Answer: cloudFiles.schemaEvolutionMode
`cloudFiles.schemaEvolutionMode` controls how Auto Loader handles schema changes, with options such as 'addNewColumns', 'rescue', and 'failOnNewColumns'.
Which of the following is a REQUIRED option when writing an Auto Loader stream output?
Answer: checkpointLocation
`checkpointLocation` is required for all Structured Streaming writes, including Auto Loader, to enable fault tolerance and exactly-once processing guarantees.
What happens to files already present in a directory when Auto Loader is started for the very first time?
Answer: They are processed along with all subsequent new files that arrive
On its first run, Auto Loader processes all existing files in the source directory and then continues to pick up new files as they arrive.