Big Data and Cloud Analytics Flashcards
6 cards from real Data and Analytics practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Big Data and Cloud Analytics flashcards as text
What are the 3 Vs traditionally used to define Big Data?
Answer: Velocity, Volume, Variety
The original three defining characteristics of Big Data are Volume (size), Velocity (speed of generation), and Variety (different data types).
What is Apache Hadoop primarily used for?
Answer: Distributed storage and batch processing of large datasets
Apache Hadoop is a distributed computing framework that uses HDFS for storage and MapReduce for batch processing of large datasets across clusters.
What is the main advantage of Apache Spark over Hadoop MapReduce?
Answer: Spark processes data in-memory, making it significantly faster
Apache Spark's in-memory processing can be up to 100x faster than Hadoop MapReduce for iterative algorithms and interactive queries.
Which AWS service is a fully managed data warehouse solution?
Answer: Amazon Redshift
Amazon Redshift is AWS's fully managed, petabyte-scale cloud data warehouse optimized for analytics workloads.
What is a data lake?
Answer: A centralized repository storing raw data in native format at any scale
A data lake stores structured, semi-structured, and unstructured data in its raw format, enabling flexible analysis without predefined schemas.
What does the term 'ETL' stand for in data engineering?
Answer: Extract, Transform, Load
ETL (Extract, Transform, Load) is the process of extracting data from sources, transforming it to the desired format, and loading it into a target system.