Data Analytics and Data Management Flashcards
7 cards from real AHIC practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Data Analytics and Data Management flashcards as text
Which data visualization type is most appropriate for displaying the distribution of length-of-stay values across a patient population?
Answer: Histogram
A histogram groups continuous values into bins and displays frequency, making it ideal for visualizing the distribution shape of a metric like length of stay.
In a relational database, a foreign key constraint ensures which type of data integrity?
Answer: Referential integrity
Referential integrity enforced by foreign keys ensures that a value in one table's column must match an existing primary key value in the referenced table.
A receiver operating characteristic (ROC) curve plots which two metrics against each other for a binary classifier at varying thresholds?
Answer: Sensitivity and 1-specificity
The ROC curve plots true positive rate (sensitivity) on the Y-axis against false positive rate (1-specificity) on the X-axis across decision thresholds.
Which data masking technique replaces real PHI values with fictitious but structurally valid values while preserving referential relationships across tables?
Answer: Synthetic data substitution
Synthetic data substitution replaces real PHI with realistic fake values (e.g., real-sounding names, valid date formats) while maintaining cross-table consistency for testing.
Which normalization form eliminates transitive dependencies, ensuring that non-key attributes depend only on the primary key and nothing else?
Answer: Third normal form (3NF)
Third normal form (3NF) requires that all non-key attributes be directly dependent on the primary key and not on other non-key attributes (no transitive dependency).
A health informatics team implements a data lake to store raw clinical, claims, and genomic data. What distinguishes a data lake from a traditional data warehouse?
Answer: Data lakes use schema-on-read and store data in its raw native format
Data lakes store raw data in any format (structured, semi-structured, unstructured) and apply schema at query time (schema-on-read), unlike warehouses that enforce schema-on-write.
Which measure evaluates the proportion of actual positive cases correctly identified by a predictive model, also known as the true positive rate?
Answer: Sensitivity (Recall)
Sensitivity (recall) = true positives / (true positives + false negatives), measuring how completely a model identifies all actual positive cases.