Data Analysis & Reporting Flashcards
7 cards from real GCP practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Data Analysis & Reporting flashcards as text
Which BigQuery feature allows you to query data directly from Google Sheets without importing it?
Answer: External tables
BigQuery external tables let you query data stored in Google Sheets (and other sources like GCS) without loading it into BigQuery storage.
In Looker Studio, what is the purpose of a 'calculated field'?
Answer: To create new metrics derived from existing data source fields
Calculated fields in Looker Studio let you define custom metrics or dimensions using formulas applied to existing fields in your data source.
You need to analyze streaming sensor data with sub-second latency on GCP. Which service is most appropriate?
Answer: Cloud Dataflow with streaming pipeline
Cloud Dataflow's streaming pipelines process unbounded data in real time, making it ideal for low-latency sensor data analysis.
What does the BigQuery PARTITION BY clause in a DDL statement accomplish?
Answer: It divides a table into segments based on a column to reduce query costs
Partitioning divides a BigQuery table into segments (e.g., by date), so queries that filter on the partition column scan only relevant partitions, reducing cost.
Which Looker Studio chart type is best suited for showing the composition of a whole across multiple categories?
Answer: Pie chart
A pie chart is designed to show part-to-whole relationships, displaying how each category contributes to the total.
In BigQuery, what is a 'slot' in the context of query execution?
Answer: A unit of computational capacity used to execute SQL queries
BigQuery slots are virtual CPUs used to execute queries; projects have a default pool and can purchase dedicated slots for guaranteed capacity.
Which GCP service provides a managed environment for running Apache Spark and Hadoop jobs for data analysis?
Answer: Dataproc
Dataproc is GCP's managed Hadoop and Spark service, allowing you to run big data processing jobs without managing cluster infrastructure.