HIPAA De-identification and Data Anonymization 5 — Questions and Answers
Question 1: A covered entity uses a coded system where each patient's name is replaced with a random number, and the code mapping is stored in a separate secure file. How does HIPAA classify this data?
- De-identified, because no names appear in the dataset
- PHI, because a code that can be translated back to an identity means the data is not truly de-identified (Correct answer)
- A limited data set requiring only a data use agreement
- Encrypted PHI that may be shared freely as long as the key is protected
Correct answer: PHI, because a code that can be translated back to an identity means the data is not truly de-identified
If a code can be used to identify an individual — even through a separate mapping — the coded data is still PHI under HIPAA because re-identification is possible.
Question 2: Which of the following is an example of a 'direct identifier' that must be removed under both Safe Harbor de-identification and limited data set preparation?
- Admission year
- Three-digit ZIP code prefix
- Diagnosis codes
- Social Security numbers (Correct answer)
Correct answer: Social Security numbers
Social Security numbers are direct identifiers that must be removed under Safe Harbor and also must be removed from limited data sets before they can be shared.
Question 3: A biobank wants to share genomic data with researchers after removing all 18 Safe Harbor identifiers. Why might this still pose re-identification risks?
- Genomic data is always PHI regardless of any de-identification method
- Genomic sequences are inherently identifying because they are unique to each individual and can be matched to public genetic databases (Correct answer)
- The Safe Harbor list was written before genomics existed and does not address DNA
- Genomic data can only be shared under Expert Determination, never Safe Harbor
Correct answer: Genomic sequences are inherently identifying because they are unique to each individual and can be matched to public genetic databases
Genomic data is unique to each individual, and even without traditional identifiers, sequences can be matched against public genealogical or research databases to re-identify donors.
Question 4: Under HIPAA's Safe Harbor method, how must URLs and IP addresses be handled?
- IP addresses must be removed; URLs may be retained if they do not contain names
- Both URLs and IP addresses are among the 18 identifiers and must be completely removed (Correct answer)
- Neither URLs nor IP addresses are covered by Safe Harbor since they are IT artifacts
- URLs must be shortened but may remain; IP addresses must be removed
Correct answer: Both URLs and IP addresses are among the 18 identifiers and must be completely removed
Both web URLs and IP addresses appear on HIPAA's list of 18 Safe Harbor identifiers and must be removed from datasets to achieve de-identification.
Question 5: What is the primary advantage of using Expert Determination over Safe Harbor for de-identification?
- Expert Determination is cheaper and faster than removing all 18 identifiers
- Expert Determination allows retention of identifiers when statistical analysis shows re-identification risk is very small, offering more analytical flexibility (Correct answer)
- Expert Determination provides a legal safe harbor that Safe Harbor does not
- Expert Determination is required by OCR for datasets larger than 10,000 records
Correct answer: Expert Determination allows retention of identifiers when statistical analysis shows re-identification risk is very small, offering more analytical flexibility
Expert Determination allows a statistician to justify retaining certain quasi-identifiers when proven low-risk, making it more flexible than mechanically removing all 18 Safe Harbor elements.
Question 6: Which of the following correctly describes 'synthetic data' as an alternative to de-identification for HIPAA compliance?
- Synthetic data is generated algorithmically to mimic the statistical properties of real PHI without being derived from any real individual's records (Correct answer)
- Synthetic data is real patient data with all 18 Safe Harbor identifiers replaced by random values
- Synthetic data is PHI shared under a data use agreement for research synthesis
- Synthetic data is de-identified data that has been verified by two independent statisticians
Correct answer: Synthetic data is generated algorithmically to mimic the statistical properties of real PHI without being derived from any real individual's records
Synthetic data is computationally generated to mirror real data's statistical characteristics but contains no information derived from actual patients, so it is not PHI under HIPAA.
Question 7: A covered entity de-identifies data using Safe Harbor and sends it to a marketing company. The marketing company combines the dataset with commercial data to re-identify individuals. Which statement is most accurate under HIPAA?
- The covered entity is liable because they should have anticipated re-identification
- The marketing company violated HIPAA because they received what was effectively PHI
- If the covered entity properly applied Safe Harbor and had no actual knowledge of re-identification, they fulfilled HIPAA obligations; the marketing company's actions may violate other laws but not HIPAA (Correct answer)
- Both parties are equally liable under HIPAA's re-identification prohibition
Correct answer: If the covered entity properly applied Safe Harbor and had no actual knowledge of re-identification, they fulfilled HIPAA obligations; the marketing company's actions may violate other laws but not HIPAA
HIPAA's de-identification standard shifts regulatory risk away from the covered entity when properly applied; however, the marketing company, as a non-covered entity, is not subject to HIPAA but may face FTC or state privacy law liability.
A covered entity uses a coded system where each patient's name is replaced with a random number, and the code mapping is stored in a separate secure file.
How does HIPAA classify this data?