AWS Certified Data Engineer – Associate (DEA-C01) — Questions and Answers
Question 1: Which Azure service provides hierarchical namespace support optimized for big-data analytics?
- Azure Files
- Azure Queue Storage
- Azure Blob Storage (flat)
- Azure Data Lake Storage Gen2 (Correct answer)
Correct answer: Azure Data Lake Storage Gen2
ADLS Gen2 adds a hierarchical namespace on top of Blob Storage for analytics workloads.
Question 2: What does setting a task SLA (Service Level Agreement) in Airflow accomplish?
- It alerts when a task exceeds an expected completion time (Correct answer)
- It speeds up the task
- It encrypts the output
- It increases retries automatically
Correct answer: It alerts when a task exceeds an expected completion time
An SLA triggers a notification when a task runs longer than its defined expected duration.
Question 3: A data engineering team is building a data lake on a major cloud platform. The primary requirement is to store massive volumes of raw, semi-structured (JSON logs) and unstructured (images, videos) data in its native format. Which cloud storage solution is most appropriate for the foundational layer of this data lake?
- Object Storage (e.g., Amazon S3, Google Cloud Storage, Azure Blob Storage) (Correct answer)
- Block Storage (e.g., Amazon EBS, Google Persistent Disk)
- An in-memory cache (e.g., Redis, Memcached)
- A managed relational database (e.g., Cloud SQL, RDS)
Correct answer: Object Storage (e.g., Amazon S3, Google Cloud Storage, Azure Blob Storage)
Object storage is designed for storing vast amounts of unstructured and semi-structured data. It offers a flat namespace, high durability and availability, virtually limitless scalability, and a low cost per GB, making it the ideal foundation for a data lake where raw data is landed before processing.
Question 4: What is the main reason to colocate compute with data in a cluster?
- Force a global sort
- Increase replication factor
- Reduce network transfer by processing data locally (Correct answer)
- Avoid using partitions
Correct answer: Reduce network transfer by processing data locally
Data locality minimizes network I/O by running tasks where the data already resides.
Question 5: When designing a dimensional model, what is the primary advantage of using a Star Schema compared to a Snowflake Schema?
- Improved data integrity through normalization.
- Reduced data redundancy and lower storage costs.
- Simplified data maintenance and updates.
- Faster query performance due to fewer joins. (Correct answer)
Correct answer: Faster query performance due to fewer joins.
The primary advantage of a Star Schema is its simplicity and faster query performance. Because dimension tables are denormalized and connect directly to the central fact table, queries require fewer joins, which typically results in faster data retrieval. [1, 3, 6] Snowflake schemas, while more storage-efficient, require more complex joins, which can slow down query performance. [5]
Question 6: Which technique mitigates skew by adding random prefixes to hot keys?
- Bucketing by replica
- Salting (Correct answer)
- Broadcasting
- Watermarking
Correct answer: Salting
Salting spreads a hot key across multiple reducers by appending random values.
Question 7: Data skew in a distributed join causes which symptom?
- A few tasks run far longer than the rest (Correct answer)
- Disk usage drops to zero
- All tasks finish at the exact same time
- The query returns wrong results
Correct answer: A few tasks run far longer than the rest
Skew concentrates rows on a few keys, overloading the tasks that process them.
Question 8: What is the main benefit of data compression in storage systems?
- Increased password strength
- Better screen brightness
- Reduced storage costs and faster I/O for large datasets (Correct answer)
- Faster typing speed
Correct answer: Reduced storage costs and faster I/O for large datasets
Compression shrinks data size, lowering storage costs and often speeding up read/write operations.
Question 9: A data engineering team is designing a data warehouse for a large financial institution. They need to prioritize storage efficiency and data integrity due to complex, multi-level hierarchies in their customer and account dimensions. Query performance is a secondary concern. Which schema design would be most appropriate for this scenario?
- Flat Denormalized Model
- Data Vault
- Snowflake Schema (Correct answer)
- Star Schema
Correct answer: Snowflake Schema
A Snowflake Schema is the most appropriate choice because it normalizes dimension tables into multiple related tables. This reduces data redundancy and improves data integrity, which is crucial for complex hierarchies. [1, 5] While this leads to more complex queries with more joins, it aligns with the stated priorities of storage efficiency and data integrity over query speed. [3, 6]
Question 10: A retail company wants to analyze sales performance. Their data model includes a central 'Sales' table with measures like 'quantity_sold' and 'total_amount'. This table is linked to other tables such as 'Product', 'Store', and 'Date' which contain descriptive attributes. What is the primary role of the 'Sales' table in this dimensional model?
- To normalize hierarchical data to save storage space.
- To provide descriptive context for the business process.
- To maintain historical changes of dimensional attributes.
- To store quantitative measures of business events. (Correct answer)
Correct answer: To store quantitative measures of business events.
The 'Sales' table is a fact table. The primary role of a fact table in a dimensional model is to store the quantitative, numeric measures of business events or transactions. [4, 8] In this scenario, 'quantity_sold' and 'total_amount' are the facts, while the linked tables ('Product', 'Store', 'Date') are dimension tables that provide context.
Question 11: Too many small files ("small files problem") hurts query performance mainly by:
- Removing partition keys
- Forcing row-based storage
- Adding per-file open/metadata overhead (Correct answer)
- Corrupting the data
Correct answer: Adding per-file open/metadata overhead
Each tiny file incurs listing and open costs that dominate when files are numerous.
Question 12: Which US law protects the privacy of patient health information?
- GLBA
- CCPA
- HIPAA (Correct answer)
- GDPR
Correct answer: HIPAA
HIPAA governs protected health information in the US.
Question 13: A retroactive correction must change history without creating a new version. Which type behavior applies?
- Type 2 insert
- Type 3 column shift
- Type 1 overwrite applied to a specific historical row (Correct answer)
- Type 0 freeze
Correct answer: Type 1 overwrite applied to a specific historical row
Overwriting in place corrects an existing version without spawning a new historical record.
Question 14: What is a common reason to use a sensor in an orchestration tool?
- To wait for an external condition like a file or partition to appear (Correct answer)
- To allocate GPU resources
- To compress output files
- To encrypt connection strings
Correct answer: To wait for an external condition like a file or partition to appear
Sensors pause a workflow until a specified external event or condition is met.
Question 15: Which encryption option lets you manage your own keys for cloud-stored data while the provider performs encryption?
- Client-side hashing
- Customer-managed keys via KMS (SSE-KMS) (Correct answer)
- No encryption
- Provider-managed keys (SSE-S3)
Correct answer: Customer-managed keys via KMS (SSE-KMS)
Customer-managed keys through a key management service give control over key rotation and access policies.
Question 16: In a hybrid Type 6 SCD, which behaviors are combined?
- Type 2 and a fact table
- Type 1, 2, and 3 (1+2+3=6) in one dimension (Correct answer)
- Two separate dimensions merged
- Type 0 and Type 1 only
Correct answer: Type 1, 2, and 3 (1+2+3=6) in one dimension
Type 6 blends overwrite, new-row history, and a current-value column for flexible reporting.
Question 17: Which statement best describes the difference between scheduling and orchestration?
- Scheduling triggers jobs by time; orchestration manages dependencies and coordination (Correct answer)
- Orchestration only runs single tasks
- Scheduling handles error recovery exclusively
- They are identical terms
Correct answer: Scheduling triggers jobs by time; orchestration manages dependencies and coordination
Scheduling decides when jobs run, while orchestration coordinates dependencies, retries, and data flow across tasks.
Question 18: In a distributed query engine like Apache Spark, you are joining a very large fact table (billions of rows) with a small dimension table (a few hundred rows). Which join strategy is the most efficient and should be used by the optimizer in this scenario?
- Cartesian Product Join
- Shuffle Hash Join
- Sort Merge Join
- Broadcast Hash Join (Correct answer)
Correct answer: Broadcast Hash Join
A Broadcast Hash Join is ideal when one table is significantly smaller than the other. The small table is duplicated (broadcast) to every worker node. The join can then be performed locally on each node without a costly network shuffle of the large table's data, making it highly efficient for this use case.
Question 19: A data engineer needs to provide a dataset to a third-party analytics vendor. The dataset contains user email addresses and phone numbers, which must be protected. The requirement is to replace these sensitive fields with irreversible, unique identifiers to prevent re-identification of individuals, while still allowing the vendor to count unique users. Which data protection technique should be used?
- Data Virtualization
- Anonymization using a cryptographic hash function (Correct answer)
- Encryption at Rest
- Data Masking with Shuffling
Correct answer: Anonymization using a cryptographic hash function
Anonymization using a one-way cryptographic hash function (like SHA-256) is the correct technique. Hashing transforms the PII into a fixed-length string that is computationally infeasible to reverse, thus protecting the original data. [7, 19, 21] Since the same input always produces the same output, it allows the vendor to count distinct individuals without exposing the actual PII. [11] Encryption is reversible, and data masking may not guarantee irreversibility or uniqueness suitable for counting.
Question 20: Which consistency model does Amazon S3 now provide for read-after-write on new objects?
- Causal consistency only
- No consistency guarantee
- Eventual consistency
- Strong read-after-write consistency (Correct answer)
Correct answer: Strong read-after-write consistency
S3 provides strong read-after-write consistency for PUTs of new objects.
Question 21: Which of the following is a primary advantage of using a traditional ETL process, particularly in industries with strict data privacy regulations like healthcare or finance?
- Lower initial setup costs due to not needing a separate transformation server.
- The ability to handle unstructured data more effectively.
- Faster data ingestion speeds compared to ELT.
- The capability to cleanse, mask, or anonymize sensitive data before it is loaded into the target data warehouse. (Correct answer)
Correct answer: The capability to cleanse, mask, or anonymize sensitive data before it is loaded into the target data warehouse.
A significant advantage of the ETL process is its ability to enhance data security and compliance. Since the transformation step occurs *before* the load step, organizations can implement rules to remove, mask, or anonymize Personally Identifiable Information (PII) or other sensitive data in a controlled staging environment. This ensures that such sensitive information never resides in its raw form within the target data warehouse, which can be a critical requirement for regulations like HIPAA or GDPR.
Question 22: Why is overwriting a natural key value risky in a dimension table?
- It speeds up queries too much
- It always violates normalization
- It removes the need for surrogate keys
- It can break historical fact-to-dimension relationships (Correct answer)
Correct answer: It can break historical fact-to-dimension relationships
Surrogate keys insulate facts from natural-key changes, which otherwise could corrupt historical joins.
Question 23: A large e-commerce platform runs its application on hundreds of virtual machines. The operations team needs to centralize all application logs in real-time for monitoring and security analysis. Which of the following represents the most common and effective architectural pattern for this task?
- Writing a custom script on each server to periodically use SSH/SCP to copy log files to a central server.
- Installing a lightweight agent (e.g., Fluentd, Filebeat) on each server to tail log files and stream events to a central log aggregator. (Correct answer)
- Directly writing logs from the application on each server to a central relational database.
- Configuring a cron job on each server to batch-compress and upload log files to cloud storage every hour.
Correct answer: Installing a lightweight agent (e.g., Fluentd, Filebeat) on each server to tail log files and stream events to a central log aggregator.
The standard and most robust pattern for centralized logging is to use a dedicated log shipping agent. Lightweight agents like Fluentd, Filebeat, or the OpenTelemetry Collector are designed specifically to run on each source machine, tail log files efficiently, and forward the log events in a streaming fashion to a central system (like Elasticsearch, OpenSearch, or a message queue). This approach is scalable, resilient to network issues, and provides real-time data. Hourly batch uploads are not real-time, direct database writes would cause a performance bottleneck, and custom SSH scripts are brittle and hard to manage at scale.
Question 24: A full-load ingestion is most appropriate when:
- The source table is small or lacks a reliable change-tracking column (Correct answer)
- The source provides a transaction log
- The table has billions of rows updated constantly
- Sub-second latency is required
Correct answer: The source table is small or lacks a reliable change-tracking column
Full loads suit small tables or sources without a dependable change indicator, since reloading everything is simplest.
Question 25: In a MapReduce job, what is the primary purpose of the shuffle phase?
- To split the input file into blocks
- To group and transfer mapper output to reducers by key (Correct answer)
- To compress the final output
- To launch the driver program
Correct answer: To group and transfer mapper output to reducers by key
The shuffle phase sorts mapper output and routes records with the same key to the same reducer.
Question 26: Which testing approach validates the statistical properties and distributions of data rather than individual row correctness?
- Unit testing
- Integration testing
- Schema validation testing
- Statistical/distribution testing (Correct answer)
Correct answer: Statistical/distribution testing
Statistical testing checks that data distributions, mean values, standard deviations, and outlier rates remain within expected bounds.
Question 27: In orchestration, what is a 'critical path'?
- The cheapest set of tasks
- The encrypted data route
- The path with the fewest tasks
- The longest dependency chain that determines minimum total runtime (Correct answer)
Correct answer: The longest dependency chain that determines minimum total runtime
The critical path is the longest sequence of dependent tasks, setting the floor on total execution time.
Question 28: What is 'column pruning' in the context of reading columnar file formats?
- The ability of a query engine to read only the columns referenced by a query, skipping others (Correct answer)
- Removing duplicate or redundant columns from a dataset schema
- Deleting outdated columns from a schema during migration
- Compressing columns that contain sparse or mostly null data
Correct answer: The ability of a query engine to read only the columns referenced by a query, skipping others
Column pruning is the optimization where the query engine reads only the specific columns referenced in a query from the file, skipping all other columns and significantly reducing disk I/O.
Question 29: In Airflow, what does the schedule_interval '@daily' do?
- Runs continuously
- Runs the DAG once at midnight each day (Correct answer)
- Runs every hour
- Runs only when triggered manually
Correct answer: Runs the DAG once at midnight each day
'@daily' schedules the DAG to run once per day at midnight.
Question 30: Which process describes Extract, Transform, Load (ETL)?
- Encrypting data before transmission
- Pulling data from sources, reshaping it, then storing it in a target system (Correct answer)
- Loading raw data first, then querying it directly
- Compressing files for backup
Correct answer: Pulling data from sources, reshaping it, then storing it in a target system
ETL extracts data from sources, transforms it into a usable shape, and loads it into a destination.
Question 31: In dbt (data build tool), what type of test checks that a column contains no NULL values?
- accepted_values
- relationships
- unique
- not_null (Correct answer)
Correct answer: not_null
The `not_null` test in dbt asserts that a specified column has no NULL values in the dataset.
Question 32: Which scenario is best handled by event-driven orchestration rather than time-based scheduling?
- Triggering a pipeline the moment a new file lands in cloud storage (Correct answer)
- A weekly cleanup job
- A monthly financial report
- A fixed nightly batch load
Correct answer: Triggering a pipeline the moment a new file lands in cloud storage
Event-driven orchestration reacts to events like file arrival rather than waiting for a clock.
Question 33: Which storage choice best supports globally distributed, low-latency reads of a content catalog?
- Single-region cold archive
- A single relational primary
- Object storage fronted by a CDN (Correct answer)
- Tape storage
Correct answer: Object storage fronted by a CDN
A CDN caches objects near users worldwide, delivering low-latency reads from object storage origins.
Question 34: In a DAG, what does 'fan-out' followed by 'fan-in' typically represent?
- Splitting work into parallel tasks, then aggregating their results (Correct answer)
- A scheduling delay
- Two unrelated DAGs
- Encrypting then decrypting data
Correct answer: Splitting work into parallel tasks, then aggregating their results
Fan-out runs tasks in parallel and fan-in collects their outputs into a downstream task.
Question 35: Which encryption approach protects data while it is stored on disk in a data warehouse?
- Encryption at rest (Correct answer)
- Encryption in transit
- Token bucketing
- TLS handshake
Correct answer: Encryption at rest
Encryption at rest secures data persisted on storage media.
Question 36: Which normal form requires that every non-key attribute is fully functionally dependent on the entire primary key, not just part of it?
- First Normal Form (1NF)
- Third Normal Form (3NF)
- Boyce-Codd Normal Form (BCNF)
- Second Normal Form (2NF) (Correct answer)
Correct answer: Second Normal Form (2NF)
2NF eliminates partial dependencies, requiring every non-key attribute to depend on the whole composite primary key.
Question 37: What is the main goal of data anonymization?
- Encrypt data in transit
- Speed up queries
- Compress large datasets
- Remove identifying information so individuals cannot be re-identified (Correct answer)
Correct answer: Remove identifying information so individuals cannot be re-identified
Anonymization strips identifiers to prevent re-identification.
Question 38: What is the main advantage of using a columnar format like Parquet in cloud storage?
- Reduced scan size and better compression for analytics (Correct answer)
- Simpler human readability
- Native support for binary blobs only
- Faster row-by-row inserts
Correct answer: Reduced scan size and better compression for analytics
Parquet stores data by column, enabling column pruning and high compression for analytical queries.
Question 39: What does denormalizing a dimension table generally improve?
- Data redundancy elimination
- Transaction integrity
- Query simplicity and read performance by reducing joins (Correct answer)
- Write performance and storage efficiency
Correct answer: Query simplicity and read performance by reducing joins
Denormalized (flat) dimensions reduce the number of joins, simplifying and speeding up analytical queries.
Question 40: What is a 'task instance' in Airflow?
- A worker node
- A specific run of a task for a particular execution date (Correct answer)
- A database connection
- A reusable operator template
Correct answer: A specific run of a task for a particular execution date
A task instance is one execution of a task tied to a specific DAG run and date.
Question 41: What is the purpose of partitioning data in cloud object storage for a query engine?
- To prune irrelevant data and reduce scan volume (Correct answer)
- To increase replication factor
- To enforce schema validation
- To encrypt sensitive fields
Correct answer: To prune irrelevant data and reduce scan volume
Partitioning lets query engines skip directories that don't match filter predicates, cutting scan cost.
Question 42: In Apache Airflow, what does a DAG represent?
- A database connection pool
- A directed acyclic graph defining task dependencies (Correct answer)
- A distributed file storage layer
- A message queue partition
Correct answer: A directed acyclic graph defining task dependencies
A DAG (Directed Acyclic Graph) defines tasks and the order in which they run, with no cycles.
Question 43: A watermark in stream processing primarily helps the system to:
- Authenticate the producer
- Decide when a time window can be considered complete despite late data (Correct answer)
- Compress the stream
- Encrypt the event payload
Correct answer: Decide when a time window can be considered complete despite late data
Watermarks track event-time progress, letting the engine close windows and handle late-arriving data sensibly.
Question 44: A financial services company's compliance department is auditing a critical regulatory report and discovers a discrepancy in a key metric. To investigate, they need to trace the metric back through all the transformations and data pipelines to its original source systems. What data governance capability is essential for this investigation?
- Data Cataloging
- Master Data Management (MDM)
- Data Encryption
- Data Lineage (Correct answer)
Correct answer: Data Lineage
Data lineage provides a complete audit trail of data's journey, showing its origin, every transformation it undergoes, and its final destination. [1, 4, 6] This visibility is crucial for root cause analysis of errors, impact analysis of changes, and meeting regulatory compliance requirements by proving the provenance of data in reports. [5, 15]
Question 45: What is the primary purpose of column-level masking in a database?
- Replicate columns across nodes
- Compress columns to save space
- Index columns faster
- Hide sensitive field values from unauthorized users (Correct answer)
Correct answer: Hide sensitive field values from unauthorized users
Masking obscures sensitive values so unauthorized users see redacted data.
Question 46: Which Airflow concept controls how many task instances can run concurrently for a single DAG?
- concurrency (max_active_tasks) (Correct answer)
- retries
- schedule_interval
- max_active_runs
Correct answer: concurrency (max_active_tasks)
The concurrency (max_active_tasks) setting limits how many task instances run at once within a DAG.
Question 47: In dimensional modeling, a Type 2 Slowly Changing Dimension handles updates by what method?
- Deleting the old record entirely
- Storing changes only in logs
- Adding a new row to preserve historical versions (Correct answer)
- Overwriting the existing value in place
Correct answer: Adding a new row to preserve historical versions
SCD Type 2 inserts a new row with validity dates, keeping the full history of the attribute.
Question 48: In Kafka, what does log compaction retain?
- Messages newer than the retention window only
- The latest value for each message key (Correct answer)
- All messages forever regardless of key
- Only the first message per partition
Correct answer: The latest value for each message key
Log compaction keeps at least the most recent value for every key, discarding older duplicates.
Question 49: What does RBAC stand for in access control?
- Realtime Batch Access Coordination
- Rule-Based Audit Compliance
- Resource Backup And Caching
- Role-Based Access Control (Correct answer)
Correct answer: Role-Based Access Control
RBAC assigns permissions based on user roles.
Question 50: How is a high-end_date often represented for the current Type 2 row?
- The load timestamp
- 0000-00-00
- A far-future sentinel date like 9999-12-31 (Correct answer)
- NULL only
Correct answer: A far-future sentinel date like 9999-12-31
A far-future sentinel keeps BETWEEN range queries simple while marking the row as open.
Question 51: In Prefect or Dagster, what is a key advantage over a pure cron schedule?
- Built-in retries, observability, and dependency management (Correct answer)
- No code is needed
- It runs without any compute
- It only supports one task per flow
Correct answer: Built-in retries, observability, and dependency management
Modern orchestrators add observability, retries, and dependency handling that cron lacks.
Question 52: When a streaming consumer cannot keep up with the producer, which technique signals the producer to slow down?
- Backpressure (Correct answer)
- Compaction
- Denormalization
- Sharding
Correct answer: Backpressure
Backpressure is the mechanism that throttles upstream producers when downstream consumers fall behind.
Question 53: What is the purpose of a 'data quality scorecard' in enterprise data engineering?
- To provide a consolidated view of data quality metrics across datasets, pipelines, and business domains (Correct answer)
- To rank data engineers by their productivity
- To calculate the cost savings from data quality improvements
- To assign ownership of data assets to specific teams
Correct answer: To provide a consolidated view of data quality metrics across datasets, pipelines, and business domains
A data quality scorecard aggregates quality metrics (completeness, accuracy, freshness) across datasets to give stakeholders a unified quality view.
Question 54: In Kafka, increasing the replication factor primarily improves:
- Fault tolerance and durability (Correct answer)
- Consumer offset commit speed
- Message ordering guarantees
- Producer compression ratio
Correct answer: Fault tolerance and durability
More replicas mean the data survives more broker failures, raising durability.
Question 55: Which metric measures the percentage of required fields that are populated with non-null values?
- Completeness rate (Correct answer)
- Validity rate
- Consistency rate
- Accuracy rate
Correct answer: Completeness rate
Completeness rate measures what percentage of expected data fields actually contain values, identifying missing data problems.
Question 56: In a mature data governance framework, which of the following roles is primarily responsible for the day-to-day management of a specific data domain, including defining data quality rules and ensuring compliance with policies for that domain?
- Chief Data Officer (CDO)
- Data Engineer
- Data Custodian
- Data Steward (Correct answer)
Correct answer: Data Steward
A Data Steward is a subject matter expert for a specific data domain (e.g., Customer Data, Product Data). They are responsible for the hands-on governance of that data, including defining its meaning, setting quality standards, and managing its lifecycle. [18, 24] A Data Custodian is more focused on the technical implementation, storage, and security of the data, while a CDO focuses on overall data strategy. [2, 17, 20]
Question 57: What does 'data profiling' accomplish in a data engineering workflow?
- Optimizes query execution plans for faster performance
- Partitions large tables to reduce scan costs
- Analyzes dataset characteristics like cardinality, null rates, and distributions to understand data quality (Correct answer)
- Encrypts sensitive data fields for compliance
Correct answer: Analyzes dataset characteristics like cardinality, null rates, and distributions to understand data quality
Data profiling examines datasets to produce statistical summaries (null counts, distinct values, min/max) that reveal data quality issues before transformation.
Question 58: What does a data retention policy define?
- Which encryption algorithm to use
- The number of replicas per table
- How fast queries must run
- How long data is kept before deletion or archival (Correct answer)
Correct answer: How long data is kept before deletion or archival
Retention policies specify how long data is stored before disposal.
Question 59: In a consistent hashing ring, what happens when one node is removed?
- The ring rebuilds from scratch
- All keys are reshuffled
- Replication stops
- Only its keys remap to the next node (Correct answer)
Correct answer: Only its keys remap to the next node
Consistent hashing limits remapping to the departing node's keys, minimizing disruption.
Question 60: Which of the following statements most accurately describes the primary difference between a data lake and a data warehouse in terms of data structure?
- A data warehouse stores data in a normalized form (3NF), while a data lake uses a denormalized star schema.
- A data lake uses a "schema-on-write" approach, while a data warehouse uses "schema-on-read".
- A data lake stores only unstructured data, while a data warehouse stores only structured data.
- A data lake stores raw data in its native format, applying structure during analysis ("schema-on-read"). (Correct answer)
Correct answer: A data lake stores raw data in its native format, applying structure during analysis ("schema-on-read").
The fundamental difference lies in when the schema is applied. A data warehouse requires a predefined schema before data is loaded (schema-on-write). In contrast, a data lake stores data in its raw, native format and the schema is applied when the data is read or queried for a specific analysis (schema-on-read), providing greater flexibility.
Question 61: Which serialization format uses binary encoding and requires a schema to serialize and deserialize data?
- XML
- CSV
- JSON
- Apache Avro (Correct answer)
Correct answer: Apache Avro
Apache Avro uses binary encoding and requires a schema (defined in JSON) to serialize and deserialize data, making it compact and efficient.
Question 62: Which concept describes ensuring data residency requirements are met for a given jurisdiction?
- Caching data in memory
- Deduplicating records
- Encrypting data twice
- Storing data within required geographic boundaries (Correct answer)
Correct answer: Storing data within required geographic boundaries
Data residency requires keeping data within specified geographic regions.
Question 63: What does 'storage tiering' aim to optimize in a cloud environment?
- User authentication
- Schema design
- The balance between access cost, retrieval latency, and storage price (Correct answer)
- Network bandwidth only
Correct answer: The balance between access cost, retrieval latency, and storage price
Tiering matches data to the storage class whose cost and latency fit its access pattern.
Question 64: What is the primary risk of a poorly designed long-running monolithic task in a DAG?
- It is too easy to test
- It cannot be scheduled
- It uses no resources
- Failures require re-running everything with no granular recovery (Correct answer)
Correct answer: Failures require re-running everything with no granular recovery
Monolithic tasks lack checkpoints, so any failure forces a full re-run.
Question 65: What is a data mart?
- A tool for writing SQL
- A subset of a data warehouse focused on a specific business area (Correct answer)
- A backup of the entire data warehouse
- A real-time streaming buffer
Correct answer: A subset of a data warehouse focused on a specific business area
A data mart is a focused, department-specific subset of a data warehouse serving a particular business function.
AWS Certified Data Engineer – Associate (DEA-C01)
The AWS Certified Data Engineer – Associate validates expertise in designing, building, and maintaining data pipelines and architectures on AWS. It covers data ingestion, transformation, storage management, operations, and security and governance.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds