Introduction To Data Engineering Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Introduction To Data Engineering flashcards as text
What is the difference between OLTP and OLAP systems?
Answer: OLTP handles transactional operations; OLAP handles analytical queries on aggregated data
OLTP is optimized for transactions while OLAP is optimized for complex analytical queries.
What is a 'data mart'?
Answer: A subset of a data warehouse focused on a specific business area
A data mart is a focused subset of a warehouse serving a particular department or function.
Which practice improves data quality by catching errors before data is used?
Answer: Data validation
Data validation checks data against rules to catch errors and ensure quality before use.
What is horizontal scaling (scaling out) in distributed data systems?
Answer: Adding more machines to share the workload
Horizontal scaling adds more nodes to distribute load, unlike vertical scaling which upgrades one machine.
What is the role of a message queue like Apache Kafka in data engineering?
Answer: To buffer and stream data reliably between producers and consumers
Kafka acts as a durable, high-throughput streaming platform decoupling data producers from consumers.
What does 'denormalization' do to a database design?
Answer: Intentionally adds redundancy to improve read/query performance
Denormalization adds redundant data to reduce joins and speed up read-heavy queries.
Which metric describes the time delay between data creation and its availability for use?
Answer: Data latency
Data latency measures the delay between when data is generated and when it becomes usable.