Big Data and Cloud Analytics Flashcards
6 cards from real Data and Analytics practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Big Data and Cloud Analytics flashcards as text
What is Google BigQuery's primary architectural advantage?
Answer: Serverless, scalable analytics separating storage from compute
BigQuery's serverless architecture separates storage from compute, allowing independent scaling and eliminating infrastructure management.
What is Apache Kafka primarily used for?
Answer: Real-time data streaming and event-driven pipelines
Apache Kafka is a distributed event streaming platform used for high-throughput, real-time data pipelines and stream processing.
What is the purpose of a CDN (Content Delivery Network) in analytics applications?
Answer: To reduce latency by caching content closer to users geographically
A CDN distributes cached content across geographically dispersed servers to reduce latency and improve load times for end users.
What does 'schema-on-read' mean in the context of data lakes?
Answer: Schema is applied when data is queried, not when it's stored
Schema-on-read means raw data is stored without a fixed structure, and the schema is applied dynamically when the data is queried.
What is Snowflake's multi-cluster, shared data architecture?
Answer: Separate storage and compute layers allowing multiple independent query engines
Snowflake separates storage from compute, allowing multiple virtual warehouses to independently access the same data without contention.
What is the purpose of data partitioning in big data systems?
Answer: To divide large datasets into smaller chunks for faster parallel processing
Partitioning divides data into smaller, manageable chunks that can be processed in parallel across multiple nodes, improving query performance.