De-identification and Data Anonymization Flashcards
7 cards from real HIPAA practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 De-identification and Data Anonymization flashcards as text
A covered entity uses a coded system where each patient's name is replaced with a random number, and the code mapping is stored in a separate secure file. How does HIPAA classify this data?
Answer: PHI, because a code that can be translated back to an identity means the data is not truly de-identified
If a code can be used to identify an individual — even through a separate mapping — the coded data is still PHI under HIPAA because re-identification is possible.
Which of the following is an example of a 'direct identifier' that must be removed under both Safe Harbor de-identification and limited data set preparation?
Answer: Social Security numbers
Social Security numbers are direct identifiers that must be removed under Safe Harbor and also must be removed from limited data sets before they can be shared.
A biobank wants to share genomic data with researchers after removing all 18 Safe Harbor identifiers. Why might this still pose re-identification risks?
Answer: Genomic sequences are inherently identifying because they are unique to each individual and can be matched to public genetic databases
Genomic data is unique to each individual, and even without traditional identifiers, sequences can be matched against public genealogical or research databases to re-identify donors.
Under HIPAA's Safe Harbor method, how must URLs and IP addresses be handled?
Answer: Both URLs and IP addresses are among the 18 identifiers and must be completely removed
Both web URLs and IP addresses appear on HIPAA's list of 18 Safe Harbor identifiers and must be removed from datasets to achieve de-identification.
What is the primary advantage of using Expert Determination over Safe Harbor for de-identification?
Answer: Expert Determination allows retention of identifiers when statistical analysis shows re-identification risk is very small, offering more analytical flexibility
Expert Determination allows a statistician to justify retaining certain quasi-identifiers when proven low-risk, making it more flexible than mechanically removing all 18 Safe Harbor elements.
Which of the following correctly describes 'synthetic data' as an alternative to de-identification for HIPAA compliance?
Answer: Synthetic data is generated algorithmically to mimic the statistical properties of real PHI without being derived from any real individual's records
Synthetic data is computationally generated to mirror real data's statistical characteristics but contains no information derived from actual patients, so it is not PHI under HIPAA.
A covered entity de-identifies data using Safe Harbor and sends it to a marketing company. The marketing company combines the dataset with commercial data to re-identify individuals. Which statement is most accurate under HIPAA?
Answer: If the covered entity properly applied Safe Harbor and had no actual knowledge of re-identification, they fulfilled HIPAA obligations; the marketing company's actions may violate other laws but not HIPAA
HIPAA's de-identification standard shifts regulatory risk away from the covered entity when properly applied; however, the marketing company, as a non-covered entity, is not subject to HIPAA but may face FTC or state privacy law liability.