HIPAA De-identification and Data Anonymization — Questions and Answers
Question 1: Under HIPAA, what is the primary significance of de-identifying health information?
- De-identified information can be sold to any party for any purpose without restriction
- De-identified information is no longer PHI and is not subject to HIPAA's privacy protections (Correct answer)
- De-identified information requires less encryption but still needs a BAA for sharing
- De-identification protects organizations from all HIPAA penalties
Correct answer: De-identified information is no longer PHI and is not subject to HIPAA's privacy protections
Once properly de-identified under HIPAA standards, health information is no longer considered PHI and falls outside the scope of HIPAA's privacy and security requirements.
Under 45 CFR §164.514(a), health information that does not identify an individual and for which there is no reasonable basis to believe it could be used to identify an individual is not PHI. Such de-identified information is not subject to HIPAA's Privacy Rule requirements for protection, disclosure authorization, patient rights, or Notice of Privacy Practices. This creates significant value for healthcare analytics, research, and public health — organizations can freely use and share de-identified datasets. However, the de-identification must meet HIPAA's specific standards; partial de-identification is insufficient.
Question 2: HIPAA's Safe Harbor de-identification method requires removal of how many specific identifiers?
- 12 identifiers
- 15 identifiers
- 18 identifiers (Correct answer)
- 24 identifiers
Correct answer: 18 identifiers
HIPAA's Safe Harbor method requires removal of 18 specific categories of identifiers to de-identify health information.
45 CFR §164.514(b)(2) specifies 18 categories of identifiers that must be removed for Safe Harbor de-identification: (1) names, (2) geographic subdivisions smaller than state, (3) dates directly related to an individual except year (including birth, death, admission, discharge dates), (4) ages over 89, (5) telephone numbers, (6) fax numbers, (7) email addresses, (8) SSNs, (9) medical record numbers, (10) health plan beneficiary numbers, (11) account numbers, (12) certificate/license numbers, (13) vehicle identifiers/serial numbers/license plates, (14) device identifiers/serial numbers, (15) web URLs, (16) IP addresses, (17) biometric identifiers (fingerprints, voice), (18) full-face photographs and comparable images. Additionally, the organization must have no actual knowledge that the remaining information could re-identify the individual.
Question 3: Under HIPAA's Expert Determination method of de-identification, who must make the determination?
- A licensed physician
- A person with appropriate knowledge and experience applying statistical and scientific principles to health data (Correct answer)
- The covered entity's Privacy Officer
- A government-certified HIPAA auditor
Correct answer: A person with appropriate knowledge and experience applying statistical and scientific principles to health data
Expert Determination requires a qualified statistician or data scientist with appropriate experience in applying scientific principles to determine that re-identification risk is very small.
45 CFR §164.514(b)(1) requires that a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods for rendering information not individually identifiable: (1) apply such principles and methods and determines that the risk is very small that the information could be used to identify an individual; and (2) documents the methods and results of the analysis. A statistician, epidemiologist, or data scientist with relevant experience qualifies. The expert must document their methodology — this documentation is the organization's evidence of compliance if challenged.
Question 4: Which of the following items must be removed under HIPAA's Safe Harbor de-identification method?
- The patient's general diagnosis code (ICD-10)
- ZIP codes with populations less than 20,000 (Correct answer)
- The state where the patient received treatment
- The year of the patient's birth
Correct answer: ZIP codes with populations less than 20,000
Geographic data below the state level must be removed, including ZIP codes that could identify small populations. The first three digits of ZIP codes may be retained if the geographic area contains more than 20,000 people.
Under Safe Harbor, geographic data smaller than state must generally be removed. However, 45 CFR §164.514(b)(2)(i)(B) provides a limited exception: the first three digits of a ZIP code may be retained if the geographic unit formed by combining all ZIP codes with those same first three digits contains more than 20,000 people (according to current publicly available data from the U.S. Bureau of the Census). All ZIP codes with first three digits in geographic units containing 20,000 or fewer people must be replaced with '000'. This prevents identification of individuals in rural areas where small ZIP codes may contain very few residents.
Question 5: Under HIPAA Safe Harbor de-identification, how must ages over 89 be handled?
- They must be replaced with the number 89
- They may be aggregated into a single category of '90 or older' (Correct answer)
- All patient ages must be removed, not just ages over 89
- Ages over 89 may be retained if the patient has consented
Correct answer: They may be aggregated into a single category of '90 or older'
HIPAA allows ages over 89 to be aggregated into a category of '90 or older' rather than being completely removed, preserving some demographic information.
45 CFR §164.514(b)(2)(i)(C) specifically addresses ages: while ages over 89 must not be expressed as individual ages (since very old individuals can sometimes be uniquely identified by their age combined with other data), they may be aggregated into a single category of '90 or older.' This allows some preservation of demographic information for elderly populations while preventing re-identification. All ages 90 and above become simply '90 or older.' Ages 89 and under may be retained as specific ages under Safe Harbor.
Question 6: What does the term 're-identification' mean in the context of HIPAA de-identification?
- Creating a new patient record for a de-identified dataset
- The process of linking de-identified information back to the specific individual it pertains to (Correct answer)
- Updating patient records with new information
- Converting paper records to electronic format
Correct answer: The process of linking de-identified information back to the specific individual it pertains to
Re-identification is the process of combining de-identified data with other available information to link the data back to specific individuals, defeating the purpose of de-identification.
Re-identification occurs when ostensibly de-identified health information is successfully linked back to the individual it describes. Research has demonstrated that de-identified datasets can often be re-identified by combining them with other publicly available data — for example, date of birth + ZIP code + gender has been shown to uniquely identify 87% of Americans (Sweeney, 2000). Medical data can also be combined with social media, voter registration, or purchase data. HIPAA's Safe Harbor and Expert Determination methods are designed to prevent re-identification, but no de-identification is perfect. Organizations receiving de-identified data should contractually prohibit re-identification attempts.
Question 7: Under HIPAA, can a covered entity re-identify de-identified information that it originally de-identified?
- No, once information is de-identified it can never be re-identified under HIPAA
- Yes, covered entities may re-identify data they originally de-identified if they have a code key, subject to privacy rule restrictions (Correct answer)
- Only patients may request re-identification of their de-identified records
- Re-identification is only permitted by research institutions with IRB approval
Correct answer: Yes, covered entities may re-identify data they originally de-identified if they have a code key, subject to privacy rule restrictions
HIPAA allows covered entities to maintain a code system to re-identify de-identified data, provided the code cannot be used to identify individuals independently and access is restricted.
45 CFR §164.514(c) permits covered entities to assign a code or other means of record identification to allow de-identified information to be re-identified, provided that: (1) the code cannot be derived from or related to information about the individual; (2) the covered entity does not use or disclose the code for any other purpose; and (3) the covered entity does not disclose the mechanism for re-identification. This allows research collaboration where de-identified data is shared externally but the covered entity can re-link results to patients for follow-up care. The mechanism must be kept separate from the de-identified data.
Question 8: A research team wants to use patient data that retains some identifiers but has most removed. Under HIPAA, what is this arrangement called?
- A partial de-identification waiver
- A limited data set (Correct answer)
- A qualified research dataset
- A clinical data waiver
Correct answer: A limited data set
A HIPAA Limited Data Set is a partial de-identification approach retaining certain identifiers (geographic data, dates) for research/public health purposes, governed by a Data Use Agreement.
45 CFR §164.514(e) establishes the Limited Data Set as a third de-identification pathway (in addition to Safe Harbor and Expert Determination). A Limited Data Set has 16 direct identifiers removed (names, addresses, phone numbers, fax, email, SSN, medical record numbers, health plan numbers, account numbers, license/certificate numbers, vehicle identifiers, device identifiers, URLs, IP addresses, biometrics, full-face photos) but may retain geographic data at the city, state, ZIP level and dates (including birth and death dates). Limited Data Sets can only be shared for research, public health, or healthcare operations, and must be governed by a Data Use Agreement (DUA) rather than a BAA.
Question 9: Under HIPAA, what document must be executed before a covered entity can share a Limited Data Set?
- A Business Associate Agreement (BAA)
- A Data Use Agreement (DUA) (Correct answer)
- A Research Authorization Form
- A HIPAA Waiver of Authorization
Correct answer: A Data Use Agreement (DUA)
Limited Data Sets can only be shared under a Data Use Agreement that restricts the recipient's use of the data and prohibits re-identification.
45 CFR §164.514(e)(4) requires a Data Use Agreement between the covered entity and the limited data set recipient. The DUA must: establish the permitted uses and disclosures of the limited data set; prohibit the recipient from using or disclosing the information for purposes other than those permitted; prohibit re-identification or contacting individuals; require the recipient to use appropriate safeguards; and require the recipient to report any unauthorized use or disclosure. A DUA is not the same as a BAA — limited data set recipients are not business associates and the DUA is specific to this research/public health pathway.
Question 10: Which of the following is NOT one of the 18 Safe Harbor identifiers that must be removed under HIPAA de-identification?
- Email addresses
- Biometric identifiers including fingerprints
- ICD-10 diagnosis codes (Correct answer)
- Device identifiers and serial numbers
Correct answer: ICD-10 diagnosis codes
ICD-10 diagnosis codes are clinical codes that do not themselves identify individuals and are not among the 18 Safe Harbor identifiers that must be removed.
The 18 Safe Harbor identifiers (45 CFR §164.514(b)(2)) focus on information that directly identifies individuals: names, geographic details below state, dates, ages over 89, phone/fax numbers, email, SSNs, medical record numbers, health plan numbers, account numbers, license/certificate numbers, vehicle identifiers, device identifiers, URLs, IP addresses, biometric identifiers, full-face photos, and any other unique identifying numbers or codes. ICD-10 codes (diagnosis codes) describe medical conditions and by themselves cannot identify individuals — they may be retained in de-identified datasets, which is what makes de-identified clinical data valuable for research.
Question 11: A data analytics company receives de-identified health data from a hospital. Under HIPAA, what restrictions apply to this recipient?
- No HIPAA restrictions apply once data is properly de-identified
- If the data was provided as a limited data set, the recipient is bound by a DUA; if fully de-identified via Safe Harbor or Expert Determination, no HIPAA restrictions apply (Correct answer)
- All recipients of health data are subject to HIPAA regardless of de-identification status
- Recipients must sign a BAA before receiving any health-related data
Correct answer: If the data was provided as a limited data set, the recipient is bound by a DUA; if fully de-identified via Safe Harbor or Expert Determination, no HIPAA restrictions apply
Properly de-identified data (Safe Harbor or Expert Determination) is not PHI and HIPAA does not restrict its use; Limited Data Sets require a DUA but do not make recipients business associates.
The key distinction: Fully de-identified data (meeting Safe Harbor or Expert Determination) is no longer PHI under HIPAA. Recipients of such data are not subject to HIPAA — they can use, share, and analyze it freely (though other laws like state privacy acts, FTC regulations, or contractual restrictions may apply). Limited Data Sets retain some quasi-identifiers and remain subject to HIPAA; recipients must sign a DUA but are not considered business associates. This framework enables healthcare analytics, population health research, and AI/ML model development while protecting individual privacy.
Question 12: Under HIPAA's Safe Harbor method, what must happen with web URLs associated with patients?
- URLs may be retained if they don't include patient names
- All web URLs that could be used to identify an individual must be removed (Correct answer)
- Only URLs from patient-facing portals need to be removed
- Web URLs are not addressed in HIPAA's de-identification provisions
Correct answer: All web URLs that could be used to identify an individual must be removed
Web Uniform Resource Locators (URLs) are included in the 18 Safe Harbor identifiers and must be removed because they can link back to identifying information about individuals.
Web URLs are identifier #15 in the Safe Harbor list (45 CFR §164.514(b)(2)(i)(O)). Patient-specific URLs — such as unique URLs generated for patient portal accounts, appointment scheduling links, or personalized health education content — can be used to identify individuals if they contain user IDs, patient numbers, or other identifying parameters. Even URLs that don't appear to contain identifying information can lead to identifying data if clicked. Removing URLs prevents a de-identified dataset from containing links that could be followed to re-identify the subject.
Question 13: Why are biometric identifiers listed as Safe Harbor identifiers requiring removal under HIPAA de-identification?
- Because biometrics are always inaccurate and unreliable
- Because biometric identifiers like fingerprints and retinal scans are unique to individuals and can definitively identify them (Correct answer)
- Only fingerprints are included; other biometrics are not covered
- Biometrics are only identifiers in criminal law, not healthcare contexts
Correct answer: Because biometric identifiers like fingerprints and retinal scans are unique to individuals and can definitively identify them
Biometric identifiers are the ultimate unique identifiers — fingerprints, retinal scans, and voice prints are biologically unique to each individual and can definitively link data to a specific person.
Biometric identifiers (#17 in the Safe Harbor list) are uniquely identifying biological measurements — fingerprints, retinal scans, iris patterns, DNA sequences, voice prints, facial geometry (covered separately as identifier #18 for full-face photos), and gait analysis. Unlike demographic data that can be shared by multiple people, biometric identifiers are (for practical purposes) unique to individuals. Including fingerprints or retinal scan data in a 'de-identified' dataset would make it trivially re-identifiable. As biometric authentication becomes common in healthcare (fingerprint login, iris scanners for controlled substance access), ensuring these identifiers are excluded from analytical datasets is increasingly important.
Question 14: Under HIPAA, what is the primary risk of using 'quasi-identifiers' in a dataset claimed to be de-identified?
- Quasi-identifiers increase storage costs
- Multiple quasi-identifiers combined can uniquely identify individuals even when no direct identifiers are present (Correct answer)
- Quasi-identifiers are permitted in all de-identified datasets
- There is no risk associated with quasi-identifiers
Correct answer: Multiple quasi-identifiers combined can uniquely identify individuals even when no direct identifiers are present
Quasi-identifiers (age, ZIP code, gender, dates) can be combined to uniquely identify individuals through linkage attacks, making data that appears de-identified actually re-identifiable.
Quasi-identifiers are data elements that don't directly identify individuals but can be combined with other data to achieve identification. Latanya Sweeney's landmark 2000 research demonstrated that ZIP code + birthdate + gender could identify 87% of Americans using publicly available census data. Healthcare data quasi-identifiers include: age, sex, race, ZIP code, admission/discharge dates, diagnosis codes, and procedure codes. The Expert Determination method specifically addresses quasi-identifier risk by modeling the probability of re-identification from the specific combination of retained variables. The Safe Harbor method addresses this by requiring removal of dates and geographic specificity below state level.
Question 15: A covered entity's dataset contains dates of service but not patient names or identifiers. Under HIPAA Safe Harbor, what must be done with these dates?
- Service dates may be retained as they are not personal identifiers
- All dates related to an individual must be removed except year, including service dates, admission dates, and discharge dates (Correct answer)
- Only birth dates must be removed; service dates may be retained
- Dates may be retained if they are more than 3 years old
Correct answer: All dates related to an individual must be removed except year, including service dates, admission dates, and discharge dates
Safe Harbor requires removal of all dates (except year) directly related to an individual, including dates of service, admission, discharge, and procedures.
Identifier #3 in Safe Harbor (45 CFR §164.514(b)(2)(i)(C)) requires removal of: all dates (other than year) directly related to an individual, including birth date, admission date, discharge date, date of death, and all ages over 89 (which must be aggregated). Specific dates can be identifying — for example, the combination of 'admitted January 15, 2024, discharged January 18, 2024, treated at St. Mary's Hospital, age 45' can uniquely identify an individual when cross-referenced with hospital admission records or social media. Retaining only the year substantially reduces re-identification risk while preserving temporal information for research purposes.
Question 16: Under HIPAA, which method of de-identification is generally considered more flexible but requires more expertise to implement properly?
- Safe Harbor method, because it provides a clear checklist
- Expert Determination method, because it allows retention of more data elements through statistical risk assessment (Correct answer)
- Limited Data Set approach, because it doesn't require any statistical analysis
- The Minimum Necessary standard, which applies to all de-identification methods
Correct answer: Expert Determination method, because it allows retention of more data elements through statistical risk assessment
Expert Determination is more flexible as it uses statistical risk analysis to justify retaining data elements that Safe Harbor would require removing, but requires qualified expert analysis.
Expert Determination (45 CFR §164.514(b)(1)) is more powerful and flexible than Safe Harbor. It allows retention of data elements that Safe Harbor prohibits (such as specific dates, ZIP codes in small populations, or combinations of quasi-identifiers) if a qualified expert can demonstrate through statistical analysis that the re-identification risk is very small. This flexibility enables more analytically useful datasets for research. However, it requires engaging qualified biostatisticians or epidemiologists, documenting methodology, and the analysis must be defensible. Organizations using Expert Determination should retain their expert's documentation as evidence of due diligence.
Question 17: What is the 'Cell Size Suppression' technique used in healthcare data de-identification?
- Removing all data from patients under 18
- Suppressing data cells where the count is so small that individuals could be identified (typically fewer than 5) (Correct answer)
- Using asterisks to mask specific numbers in reports
- Suppressing all data that contains ZIP codes
Correct answer: Suppressing data cells where the count is so small that individuals could be identified (typically fewer than 5)
Cell size suppression removes or masks data cells where counts are too small to prevent identification of individuals in small groups, commonly suppressing cells with fewer than 5 individuals.
Cell size suppression is a statistical disclosure limitation technique used in healthcare data analytics. When aggregate data cells (e.g., 'patients with diagnosis X in ZIP code Y in year Z') contain very few individuals — typically fewer than 5 — those individuals can often be identified. Common practice is to suppress (replace with '<5' or 'N/A') any cell with fewer than 5 individuals. This prevents the dataset recipient from knowing that, for example, a particular small community had exactly 2 cases of a sensitive diagnosis in a given year, which could reveal specific individuals. Expert Determination practitioners apply cell suppression along with other statistical techniques to achieve overall low re-identification risk.
Question 18: Under HIPAA, may a covered entity share properly de-identified health data with a competitor for research purposes without a BAA?
- No, all health data sharing requires a BAA regardless of de-identification status
- Yes, properly de-identified data is not PHI and HIPAA does not restrict its sharing or use, including with competitors (Correct answer)
- Only sharing for public health purposes is permitted without a BAA
- De-identified data may be shared only with non-profit organizations
Correct answer: Yes, properly de-identified data is not PHI and HIPAA does not restrict its sharing or use, including with competitors
Properly de-identified data is not PHI and HIPAA imposes no restrictions on its sharing — though other legal considerations (trade secrets, state law, contract terms) may apply.
Once health information meets HIPAA's de-identification standards (Safe Harbor or Expert Determination), it is no longer PHI and falls entirely outside HIPAA's regulatory framework. No BAA is needed. The covered entity may share it with anyone — competitors, commercial analytics firms, academic researchers, data brokers — without HIPAA restriction. This is why de-identified healthcare data has become a significant commercial asset, with healthcare organizations licensing de-identified claims, EHR data, and wearable data for pharmaceutical research, AI development, and population health analytics. However, organizations should consider state privacy laws (CCPA, etc.), contractual obligations, and reputational risks before sharing de-identified data broadly.
Question 19: What is 'k-anonymity' in the context of healthcare data de-identification?
- A HIPAA-required certification for de-identification specialists
- A technique ensuring each individual's record cannot be distinguished from at least k-1 other individuals in a dataset (Correct answer)
- The minimum number of identifiers that must be removed under Safe Harbor
- A federal standard for data encryption key management
Correct answer: A technique ensuring each individual's record cannot be distinguished from at least k-1 other individuals in a dataset
k-anonymity is a privacy model ensuring that each record in a dataset is indistinguishable from at least k-1 other records based on quasi-identifiers, reducing re-identification risk.
k-anonymity (Samarati and Sweeney, 1998) is a formal privacy model often applied in Expert Determination de-identification. A dataset satisfies k-anonymity if every combination of quasi-identifiers in the dataset appears in at least k records. For example, with k=5: if '50-year-old female in ZIP 12345' appears in only one record, that person is uniquely identifiable (k=1 violation). Increasing k means each person 'hides in a crowd' of at least k similar records. HIPAA's Expert Determination approach often uses k-anonymity analysis (plus extensions like l-diversity and t-closeness that address the shortcomings of k-anonymity for sensitive attributes) to statistically certify low re-identification risk.
Question 20: Under HIPAA, what does 'actual knowledge' mean in the context of the Safe Harbor de-identification method?
- The Privacy Officer must certify each de-identified dataset personally
- The covered entity must have no actual knowledge that the information, in combination with other information, could be used to identify an individual (Correct answer)
- Knowledge that a breach has occurred, triggering notification obligations
- Documentation that staff know how to apply Safe Harbor requirements
Correct answer: The covered entity must have no actual knowledge that the information, in combination with other information, could be used to identify an individual
Safe Harbor requires not only removal of 18 identifiers but also that the organization has no 'actual knowledge' that remaining information could re-identify individuals.
45 CFR §164.514(b)(2)(ii) adds a second requirement to Safe Harbor beyond removing the 18 identifiers: 'The covered entity does not have actual knowledge that the information could be used alone or in combination with other information to identify an individual who is the subject of the information.' This means if the covered entity knows that — despite removing the 18 identifiers — some combination of remaining variables in their specific dataset could identify individuals (perhaps because of unique local characteristics or known available external data), they cannot claim Safe Harbor compliance. This provision requires some judgment beyond mechanical checklist compliance.
Question 21: Under HIPAA's Safe Harbor method, what must be done with telephone and fax numbers associated with patients?
- Telephone numbers may be retained if they are not cell phones
- All telephone numbers and fax numbers associated with individuals must be removed (Correct answer)
- Only direct lines to patients must be removed; facility main numbers may be retained
- Phone numbers may be retained if more than 2 years old
Correct answer: All telephone numbers and fax numbers associated with individuals must be removed
Both telephone and fax numbers are among the 18 Safe Harbor identifiers and must be removed during de-identification.
Telephone numbers (identifier #5) and fax numbers (identifier #6) are explicitly listed in HIPAA's Safe Harbor de-identification requirements (45 CFR §164.514(b)(2)). Any number that could be used to contact or identify a specific individual must be removed. This includes: home phone, cell phone, work direct line, and personal fax numbers. General facility or department phone numbers that are not personally associated with a specific patient are not covered. In practice, any number appearing in a patient record as belonging to the patient, their emergency contact, or their personal physician should be treated as a potential identifier requiring removal.
Question 22: What is 'data suppression' as a de-identification technique and when is it applied?
- Encrypting data before sharing it
- Removing specific records or data cells that would allow identification of individuals, particularly in small groups (Correct answer)
- Converting all data to aggregate statistics
- Adding false data to confuse potential re-identification attempts
Correct answer: Removing specific records or data cells that would allow identification of individuals, particularly in small groups
Data suppression removes specific records or cells — particularly small-population cells — that could enable re-identification of individuals.
Data suppression is used when specific data combinations would make individuals identifiable even after removing direct identifiers. Common applications: suppressing records where geographic + demographic combinations uniquely identify a person; removing cells in statistical tables where the count is below a minimum (typically 5) to prevent inference about specific individuals; suppressing records of individuals who are outliers in quantitative measures that could identify them. Suppression trades some analytical utility for privacy protection. In Expert Determination analyses, suppression is one of the tools applied alongside generalization, rounding, and top-coding to achieve the required very small re-identification risk.
Question 23: Under HIPAA, can dates of birth be retained in a de-identified dataset?
- Yes, dates of birth are standard demographic information and are not PHI identifiers
- No, dates of birth must be removed; only the year of birth may be retained for individuals 89 and under (Correct answer)
- Dates of birth may be retained if the individual is over 18
- Birth dates may be retained in research datasets approved by an IRB
Correct answer: No, dates of birth must be removed; only the year of birth may be retained for individuals 89 and under
Under Safe Harbor de-identification, birth dates (as specific dates) must be removed; only the year of birth may be retained for individuals 89 and under.
Birthdate is identifier #3 in HIPAA's Safe Harbor list (45 CFR §164.514(b)(2)(i)(C)). Specific birthdates (month/day/year) must be removed from de-identified datasets. The year of birth alone may be retained for individuals 89 years old and under. For individuals 90 and older, even the year cannot be retained — they must be categorized as '90 or older.' The reason: precise birthdates are highly identifying, particularly when combined with other quasi-identifiers like ZIP code or gender. The famous '87% identifiability' research by Sweeney used ZIP + gender + birthdate. Removing the specific date while retaining the year preserves some analytical value (age calculations) while dramatically reducing re-identification risk.
Question 24: What is 'generalization' as used in health data de-identification?
- Making all health records generic by removing diagnosis-specific information
- Replacing specific values with broader categories (e.g., replacing exact age with an age range) to reduce re-identification risk (Correct answer)
- Summarizing individual records into population statistics
- Replacing real patient names with fictional generalized names
Correct answer: Replacing specific values with broader categories (e.g., replacing exact age with an age range) to reduce re-identification risk
Generalization replaces specific values with broader categories — such as converting exact age '34' to age range '30-39' — to reduce uniqueness and re-identification risk.
Generalization is a data transformation technique that replaces specific values with broader, less identifying categories. Examples in healthcare de-identification: replacing exact age with 5-year or 10-year age bands; replacing 5-digit ZIP codes with 3-digit ZIP prefix or region; replacing specific diagnosis dates with year-only or year-quarter; replacing specific drug dosages with dose ranges; replacing exact procedure times with general time-of-day categories. Generalization trades some analytical precision for privacy protection. The Expert Determination method uses generalization selectively — applying it where specific values create re-identification risk while retaining specificity where combinations remain safe. Safe Harbor's removal of dates and geographic details below state level is essentially a mandated form of generalization.
Question 25: Under HIPAA, what is required when a covered entity wants to share a de-identified dataset with a researcher who intends to combine it with another dataset?
- The covered entity must pre-approve all dataset combinations before sharing
- The covered entity should contractually prohibit re-identification attempts and consider whether the planned combination creates re-identification risk before sharing (Correct answer)
- Once data is de-identified, the covered entity has no further responsibility
- Combining de-identified datasets is prohibited under HIPAA
Correct answer: The covered entity should contractually prohibit re-identification attempts and consider whether the planned combination creates re-identification risk before sharing
While de-identified data is not PHI, covered entities should consider whether planned data combinations could re-identify individuals and contractually prohibit re-identification.
When a covered entity shares de-identified data and knows the recipient plans to link it with other datasets, the risk of re-identification increases significantly (called a 'linkage attack'). Best practices and emerging privacy standards recommend: informing the Expert Determination analysis of the specific planned combination; requiring the data use agreement to prohibit re-identification attempts; considering whether the combined dataset maintains the very small re-identification risk standard; and designing de-identification specifically to be robust against known planned combinations. While HIPAA technically releases the covered entity from obligations upon proper de-identification, contractual protections and consideration of downstream use represent responsible data stewardship and reduce reputational risk.
Question 26: What is 'top-coding' and 'bottom-coding' in the context of health data de-identification?
- Encrypting headers and footers of data files
- Capping extreme values at a threshold to prevent identification of outliers — e.g., capping age at '65+' or capping charges at '$100,000+' (Correct answer)
- Removing the first and last records in a sorted dataset
- Applying code-based de-identification to the top and bottom of each field
Correct answer: Capping extreme values at a threshold to prevent identification of outliers — e.g., capping age at '65+' or capping charges at '$100,000+'
Top-coding and bottom-coding cap extreme values at a threshold, preventing identification of outliers who might be identifiable by unusual values.
Outliers in healthcare data can be uniquely identifying — a person who is 97 years old, had a $2 million hospitalization, or has a diagnosis code combination seen in only 3 patients nationally may be identifiable despite having direct identifiers removed. Top-coding caps high values above a threshold (age shown as '90+', charges shown as '$500,000+'); bottom-coding caps low values below a threshold. HIPAA's Safe Harbor specifically requires top-coding of ages above 89. Expert Determination analyses use top-coding and bottom-coding more broadly based on the actual distribution of values in the specific dataset. These techniques are standard statistical disclosure limitation methods applied in federal statistical agency data releases.
Question 27: Under HIPAA, which type of biometric data NOT explicitly listed in the 18 Safe Harbor identifiers might still need to be considered for de-identification?
- Height measurements
- Genomic/DNA sequences, which contain unique individual identifiers not yet explicitly listed but clearly identifying (Correct answer)
- Blood type, which is not unique and can be safely retained
- BMI calculations, which are derived measures rather than identifiers
Correct answer: Genomic/DNA sequences, which contain unique individual identifiers not yet explicitly listed but clearly identifying
Genomic data is uniquely identifying (more so than fingerprints) and covered entities should treat it as identifying even though it wasn't enumerated in the original Safe Harbor list.
When HIPAA's original Safe Harbor identifiers were drafted (circa 2000), genomic data was not as prevalent in healthcare. DNA sequences are extraordinarily unique identifiers — more so than fingerprints or SSNs — and can identify individuals even from a small DNA fragment. The Safe Harbor list includes 'biometric identifiers including finger and voice prints' (#17) and the catch-all 'any other unique identifying number, characteristic, or code' (#18). Genomic sequences fall under identifier #18's catch-all provision. Additionally, the GINA (Genetic Information Nondiscrimination Act) intersects with HIPAA for genetic information. Any dataset containing genomic sequences should treat them as high-risk identifiers requiring removal or specialist de-identification.
Question 28: What is the purpose of HIPAA's 'Safe Harbor' method's requirement for no 'actual knowledge' of re-identification risk?
- It creates a subjective safe harbor that any organization can claim
- It prevents organizations from knowingly releasing re-identifiable data while hiding behind the Safe Harbor checklist (Correct answer)
- Actual knowledge only applies to healthcare clearinghouses, not providers
- The actual knowledge requirement only applies if fewer than 500 individuals are in the dataset
Correct answer: It prevents organizations from knowingly releasing re-identifiable data while hiding behind the Safe Harbor checklist
The actual knowledge requirement prevents gaming of the Safe Harbor checklist — if the organization knows the stripped data can still identify individuals, they cannot claim Safe Harbor compliance.
45 CFR §164.514(b)(2)(ii) requires that after removing the 18 identifiers, 'the covered entity does not have actual knowledge that the information could be used alone or in combination with other information to identify an individual.' This provision addresses a known limitation of the Safe Harbor list: it was designed for general use, but specific datasets in specific contexts may retain re-identification potential even after removing all 18 identifiers. For example, a covered entity releasing data about a rare disease population of 15 patients in a specific city knows that even with all 18 identifiers removed, the data could identify individuals. Claiming Safe Harbor compliance in this situation would violate the actual knowledge standard.
Question 29: Under HIPAA, may a covered entity charge a fee for de-identifying patient data before releasing it to a researcher?
- No fees may be charged for de-identification activities
- Covered entities may charge reasonable cost-based fees for the labor, materials, and supplies involved in de-identification (Correct answer)
- Fees for de-identification are not addressed in HIPAA and therefore are unlimited
- De-identification services must be provided free of charge to academic researchers only
Correct answer: Covered entities may charge reasonable cost-based fees for the labor, materials, and supplies involved in de-identification
HIPAA permits covered entities to charge reasonable cost-based fees for de-identification activities as part of releasing health information for research purposes.
HIPAA's fee provisions (45 CFR §164.524(c)(4)) allow cost-based fees for providing PHI and by extension for de-identification work. Covered entities may charge fees that represent the actual costs of: labor for de-identification; materials and supplies; postage; and preparation of an explanation if requested. Fees must be cost-based — not a profit center. Many academic medical centers and health systems have established data services offices with defined fee schedules for de-identification and data sharing. Organizations should document their actual costs to defend fee structures if challenged. Note that fees for responding to patient access requests specifically have additional restrictions under the 2020 right of access rule updates.
Question 30: What is 'l-diversity' in the context of healthcare data de-identification, and why was it developed?
- A HIPAA requirement for diverse representation in healthcare datasets
- An extension of k-anonymity ensuring that sensitive attributes within each equivalence class have sufficient diversity to prevent sensitive attribute disclosure (Correct answer)
- A statistical measure of how many languages a dataset covers
- A certification level for de-identification software platforms
Correct answer: An extension of k-anonymity ensuring that sensitive attributes within each equivalence class have sufficient diversity to prevent sensitive attribute disclosure
l-diversity extends k-anonymity by requiring that sensitive attributes (diagnoses, treatments) within groups of indistinguishable records are sufficiently diverse to prevent inference attacks.
k-anonymity prevents re-identification by ensuring each individual hides in a group of at least k others with identical quasi-identifiers. But k-anonymity doesn't prevent attribute disclosure attacks: if a group of 5 people with identical quasi-identifiers (same age, ZIP, gender) all have the same diagnosis, knowing someone is in that group reveals their diagnosis even without identifying who they are. l-diversity (Machanavajjhala et al., 2006) extends k-anonymity by requiring that the sensitive attribute values within each k-anonymous group have at least l 'well-represented' values. Expert Determination analyses often apply l-diversity principles when datasets contain sensitive health conditions to prevent both re-identification and diagnosis disclosure attacks.
Question 31: Under HIPAA, what happens to the de-identified status of information if additional data is added that re-identifies individuals?
- The information remains de-identified as long as it was originally properly de-identified
- When de-identified information is combined with identified data, the resulting combined dataset is PHI subject to full HIPAA protections (Correct answer)
- Only the newly added identifying data becomes PHI; the original de-identified portion remains non-PHI
- HIPAA does not address what happens when de-identified data is later re-identified
Correct answer: When de-identified information is combined with identified data, the resulting combined dataset is PHI subject to full HIPAA protections
When de-identified information is combined with identifying information, the entire resulting dataset is treated as PHI because the data can now be linked to individuals.
45 CFR §164.514(a) defines de-identified information as health information that 'does not identify an individual and with respect to which there is no reasonable basis to believe that the information can be used to identify an individual.' When de-identified information is re-linked to identifiers (whether through a covered entity's own code system, a linkage attack by an external party, or deliberate re-identification), the resulting information meets the definition of PHI and is subject to full HIPAA protections. This is why covered entities that maintain re-identification keys must apply appropriate safeguards to the key mechanism and should contractually prohibit recipients of de-identified data from re-identifying it.
Question 32: What is the 'mosaic effect' in the context of health data de-identification?
- A graphic design technique for displaying health statistics
- The phenomenon where multiple pieces of seemingly harmless information, when combined, create a complete picture that identifies an individual (Correct answer)
- A method for visually representing de-identified patient populations
- The pattern of removing identifiers from a dataset one at a time
Correct answer: The phenomenon where multiple pieces of seemingly harmless information, when combined, create a complete picture that identifies an individual
The mosaic effect describes how combining multiple individually non-identifying data pieces creates an identifying picture — similar to how individual tiles form a recognizable image.
The mosaic effect (or mosaic theory, originally from national security law) in data privacy refers to the way multiple individually innocuous data points combine to create identifying information. In healthcare de-identification: a patient's age (not unique), ZIP code (not unique), race (not unique), diagnosis (not unique), and service date (not unique) — none of these alone identifies the person. But the combination creates a profile that may be unique. Expert Determination practitioners model this explicitly, computing the probability that a person's specific combination of retained attributes appears in only one person in the population (or a very small number). The Safe Harbor method addresses the mosaic effect by removing the most sensitive quasi-identifiers.
Question 33: Under HIPAA, what is required documentation for the Expert Determination de-identification method?
- A letter from the Privacy Officer certifying de-identification was performed
- Documentation of the methods applied, statistical analysis performed, and conclusions reached regarding the low probability of re-identification (Correct answer)
- Registration with HHS as a qualified de-identification provider
- An IRB-approved research protocol covering the de-identification analysis
Correct answer: Documentation of the methods applied, statistical analysis performed, and conclusions reached regarding the low probability of re-identification
Expert Determination requires documented methodology, analysis, and conclusions — the expert's written certification and supporting analysis are the compliance evidence.
45 CFR §164.514(b)(1) requires documentation for Expert Determination: the expert must document 'the methods and results of the analysis that justify such determination.' This documentation should include: description of the dataset; quasi-identifier analysis; assessment of external data available for linkage attacks; statistical methods applied (k-anonymity analysis, uniqueness computations, modeling of specific threat scenarios); transformation decisions and rationale; and the expert's conclusion with supporting evidence. This documentation is the covered entity's evidence of HIPAA compliance if the de-identification is later challenged. The covered entity should retain the expert's report as a record of the de-identification determination, maintaining it under HIPAA's 6-year documentation retention requirement.
Question 34: What is 'noise addition' or 'perturbation' as a de-identification technique in healthcare data?
- Adding audio noise to prevent voice recognition of recorded patient consultations
- Adding small, random statistical noise to numerical values to prevent exact-match re-identification while preserving analytical utility (Correct answer)
- Randomizing the order of records in a dataset to prevent sequential identification
- Adding false records to a dataset to confuse re-identification attempts
Correct answer: Adding small, random statistical noise to numerical values to prevent exact-match re-identification while preserving analytical utility
Noise addition/perturbation modifies numerical values by small random amounts to prevent exact-value matching for re-identification while preserving statistical distributions.
Perturbation (noise addition) is a statistical disclosure limitation technique that adds small random values to continuous variables (age, lab values, charges) to prevent exact-match linkage attacks. For example, if a patient's exact age and lab values were published and an adversary has access to a clinical database with the same patient's exact values, they could link the records. Adding ±1-2 years to ages and small percentages to lab values prevents exact matching while preserving statistical properties for population-level analysis. Perturbation must be calibrated: too little provides insufficient protection; too much destroys analytical utility. Expert Determination practitioners apply perturbation when sensitive continuous variables would create re-identification risk even after removing categorical identifiers.
Question 35: Under HIPAA, what special consideration applies to geographic data for patients who live in very rural areas?
- Rural patients receive less privacy protection because geographic information is publicly known
- Rural patients are at higher re-identification risk because small population sizes make geographic data more identifying, requiring more aggressive geographic generalization (Correct answer)
- HIPAA treats urban and rural geographic data identically
- Rural patient ZIP codes are automatically treated as public information under HIPAA
Correct answer: Rural patients are at higher re-identification risk because small population sizes make geographic data more identifying, requiring more aggressive geographic generalization
Rural patients face higher re-identification risk from geographic data because small populations mean geographic specificity uniquely identifies very few individuals.
Geographic specificity is much more identifying in rural populations than urban ones. A ZIP code in Manhattan with 75,000 residents provides little uniqueness; the same-digit-count ZIP in rural Wyoming with 200 residents could identify individuals easily. HIPAA's Safe Harbor recognizes this with its 20,000-person threshold for 3-digit ZIP code retention. Expert Determination analysts must be especially careful with rural patient data: even removing specific ZIP codes, the combination of geographic region + demographic data + diagnosis may uniquely identify individuals in sparse populations. Additional geographic generalization (county level, multi-county regions, or only state) may be necessary. This disparity means rural healthcare data requires more aggressive de-identification than urban data to achieve equivalent privacy protection.
Question 36: Under HIPAA, why must Social Security Numbers be removed in the Safe Harbor de-identification method?
- SSNs are government data and not healthcare information
- SSNs are unique national identifiers that can definitively link de-identified health records back to specific individuals through cross-referencing (Correct answer)
- Only the last 4 digits of SSNs need to be removed
- SSNs are already encrypted in healthcare databases
Correct answer: SSNs are unique national identifiers that can definitively link de-identified health records back to specific individuals through cross-referencing
SSNs are unique national identifiers that can be cross-referenced with government and commercial databases to link apparently de-identified health records to specific individuals.
Social Security Numbers (#8 in the Safe Harbor list) must be removed because they are de facto universal identifiers in US administrative systems — present in IRS records, DMV records, credit databases, Social Security Administration records, and many other systems. An adversary with a de-identified health record containing an SSN can trivially identify the individual by cross-referencing any of these databases. Even partial SSNs (last 4 digits) are sometimes sufficient for re-identification when combined with other quasi-identifiers. Despite being phased out of Medicare cards and many healthcare contexts, SSNs often appear in legacy health records and must be identified and removed during de-identification processing.
Under HIPAA, what is the primary significance of de-identifying health information?