Research · 11 min read

What De-Identified Actually Means

The word has a legal test behind it, and at least three different tests are in circulation. Two of them are satisfied partly by a promise the company makes about itself, which is why a reader can read the claim but cannot check it.

Key takeaways

  • De-identified is a defined legal term with specified methods, and meeting a method is not the same as being anonymous.
  • The federal rule allows either a documented expert determination that the identification risk is very small, or removal of eighteen listed identifier categories.
  • The list method has a second condition: the entity must also lack actual knowledge that what remains could identify someone.
  • The same section lets an entity keep a re-identification code, provided the code is unrelated to the person and the mechanism is not disclosed.
  • A limited data set is still protected health information by the regulation's own words, and it requires a data use agreement.
  • Both state consumer definitions rest partly on a public commitment and a contract term, so two of their three conditions are promises rather than properties.

Answer first: it is a defined term, not a description

When a health company says data is de-identified, it is usually pointing at a legal standard rather than describing a state of affairs.

The federal health privacy rules set out one standard. Two state privacy statutes set out another, close in wording and different in structure.

None of them promises that a person cannot be recognized. They set tests, and a test is passed by doing specified things, not by achieving anonymity.

That distinction is the whole subject. The rest of this explains what each test asks for and what each one leaves alone.

The federal standard, and the two ways to meet it

The federal regulation states the standard first. Health information that does not identify an individual, and about which there is no reasonable basis to believe the information can be used to identify an individual, is not individually identifiable health information.

It then says a covered entity may determine that information meets that standard only in one of two ways.

The first is an expert determination. A person with appropriate knowledge of and experience with generally accepted statistical and scientific principles applies those methods. That person determines that the risk is very small that the information could be used to identify an individual. The judgment covers use alone or in combination with other reasonably available information, by an anticipated recipient.

That person also has to document the methods and results of the analysis that justify the determination.

Notice how much of that sentence is about circumstances rather than about the data. The risk is judged against other reasonably available information and against an anticipated recipient. Change the recipient and you have changed the question.

The second way is a list, and the list is the education

The alternative method removes eighteen categories of identifiers, of the individual and of relatives, employers and household members.

Names. All geographic subdivisions smaller than a state, including street address, city, county, precinct and zip code. There is a narrow exception for the first three digits of a zip code, where the area those digits cover holds more than twenty thousand people.

All elements of dates except year for dates directly related to an individual, including birth date, admission date, discharge date and date of death. All ages over eighty-nine, which the regulation allows to be collapsed into a single category.

Telephone numbers, fax numbers, email addresses, social security numbers, medical record numbers, health plan beneficiary numbers, account numbers, certificate and license numbers.

Vehicle identifiers and serial numbers. Device identifiers and serial numbers. Web addresses. Internet protocol address numbers. Biometric identifiers including finger and voice prints. Full face photographic images and comparable images.

And then the catch-all that carries the weight: any other unique identifying number, characteristic or code, except a re-identification code permitted by the section's own next paragraph.

The method has a second half that is easy to miss. Removing the list is not enough on its own. The entity must also not have actual knowledge that the remaining information could be used, alone or in combination with other information, to identify an individual.

The holder may keep the key

The same section permits a covered entity to assign a code or other means of record identification so that de-identified information can be re-identified by that entity later.

Two conditions attach. The code must not be derived from or related to information about the individual, and must not otherwise be capable of being translated so as to identify the individual.

And the entity must not use or disclose the code for any other purpose, and must not disclose the mechanism for re-identification.

So under the federal rule, de-identified is compatible with the holder being able to put the name back. What the rule restricts is who else can do it and what the code may be used for.

A limited data set is a different thing with a similar sound

The same regulation defines a limited data set, and the definition begins by saying what it is. A limited data set is protected health information that excludes a list of direct identifiers.

That opening clause is the point. A limited data set has not left the rulebook. It is still protected health information, and it may be used or disclosed only where the entity enters into a data use agreement with the recipient.

Its exclusion list is also shorter than the de-identification list. It removes names, most address information other than town or city, state and zip code, and telephone and fax numbers. Email addresses, social security numbers, record and beneficiary and account numbers, and certificate and license numbers come out too. So do vehicle and device identifiers, web addresses, internet protocol addresses and biometric identifiers.

Dates survive it. So does the three-digit-plus zip code. A dataset described as stripped of direct identifiers may be this, and this is not de-identification.

The consumer-law version, and the part that is a promise

California's consumer privacy statute defines deidentified as information that cannot reasonably be used to infer information about, or otherwise be linked to, a particular consumer, provided that the business that possesses it does three things.

It takes reasonable measures to ensure the information cannot be associated with a consumer or household.

It publicly commits to maintain and use the information in deidentified form and not to attempt to reidentify it, with one narrow exception. The business may attempt re-identification solely to determine whether its own deidentification process satisfies the requirement.

And it contractually obligates any recipients of the information to comply with all of those provisions.

Washington's health data statute uses the same three-part structure and adds one word. Its definition asks whether the data can reasonably be used to infer information about, or otherwise be linked to, an identified or identifiable consumer, or a device linked to such a consumer.

Read the three conditions together and the shape becomes clear. One is a technical measure. The other two are a public commitment and a contract term, which means a good part of the definition is satisfied by what a company undertakes rather than by what the data is.

Why the classification changes which rulebook applies

This is not a labeling exercise. Meeting one of these definitions moves information out of a regime.

The Washington act states plainly that personal information does not include deidentified data, and its definition of consumer health data is built on personal information.

The same act's exemption section goes further in one direction. It does not apply to information that is deidentified in accordance with the requirements set out in the federal privacy regulations and that is derived from the health-care information the exemption section lists.

It also carves out information that is part of a limited data set and is used, disclosed and maintained in the manner the federal limited data set provision requires. That is a rare instance of one statute borrowing another's machinery by name.

So a determination made under a federal method can decide whether a state statute reaches the same records. The methods are not interchangeable, but the outcomes connect.

What a reader can and cannot verify

You can read whether a policy uses the word at all, and whether it says which standard it means.

You can look for the public commitment the consumer statutes require. That commitment is meant to be public, so its absence from a published document is something you can notice.

You can check whether the policy says anything about obligating recipients, since the third condition in both consumer definitions is a contract term imposed on whoever receives the data.

You cannot verify the technical measures. Neither can you verify an expert determination, which is documented internally and is a judgment about risk rather than a certificate.

And you cannot tell from outside which dataset a sentence refers to. A single policy can describe an identified record, a limited data set and a de-identified extract in three consecutive paragraphs without saying which is which.

What this does not decide

It does not say that any company's de-identification is adequate or inadequate. That is an assessment of specific data and specific methods, and nothing here reaches it.

It does not say that de-identified information has been re-identified in this market. No such claim is made and none is implied.

It does not describe the standards in any state other than the two named, and other states define the term differently.

And it is not legal advice. It reports what one federal section and two state definitions say.

Sources

  1. 45 CFR 164.514, "Other requirements relating to uses and disclosures of protected health information"Department of Health and Human Services, via the Electronic Code of Federal Regulations · Current text as displayed · Retrieved September 2026The de-identification standard in paragraph (a): health information that does not identify an individual and with respect to which there is no reasonable basis to believe the information can be used to identify an individual is not individually identifiable health information. The two permitted methods in paragraph (b): the expert determination, its very small risk formulation referring to other reasonably available information and an anticipated recipient, and its documentation requirement; and the removal of the eighteen identifier categories listed at (b)(2)(i)(A) through (R), including the three-digit zip code exception and its twenty thousand person threshold, the treatment of dates and of ages over eighty-nine, device identifiers, web addresses, internet protocol address numbers, biometric identifiers, full face images, and the catch-all for any other unique identifying number, characteristic or code. The additional condition at (b)(2)(ii) that the entity lack actual knowledge that the remaining information could be used to identify an individual. The re-identification code provisions at paragraph (c), including that the code must not be derived from or related to information about the individual and that the mechanism must not be disclosed. The limited data set provisions at paragraph (e), including the definition of a limited data set as protected health information excluding the listed direct identifiers, and the data use agreement requirement.
  2. California Civil Code section 1798.140, "Definitions"California Legislative Information, Office of Legislative Counsel · Amendment credit printed on the section: Amended by Stats. 2025, Ch. 67, Sec. 27 (AB 1170), effective January 2026 · Retrieved September 2026The definition of deidentified information as information that cannot reasonably be used to infer information about, or otherwise be linked to, a particular consumer, subject to three conditions on the business that possesses it: taking reasonable measures to ensure the information cannot be associated with a consumer or household; publicly committing to maintain and use the information in deidentified form and not to attempt to reidentify it, except solely to determine whether its deidentification processes satisfy the subdivision; and contractually obligating any recipients to comply with all provisions of the subdivision.
  3. RCW 19.373.010, "Definitions"Washington State Legislature · Session law credit printed on the section: 2023 c 191 s 3 · Retrieved September 2026The definition of deidentified data as data that cannot reasonably be used to infer information about, or otherwise be linked to, an identified or identifiable consumer, or a device linked to such consumer, subject to three conditions: reasonable measures to ensure the data cannot be associated with a consumer, a public commitment to process the data only in a deidentified fashion and not attempt to reidentify it, and a contractual obligation on any recipients to satisfy the same criteria. The statement in the definition of personal information that the term does not include deidentified data.
  4. RCW 19.373.100, "Exemptions"Washington State Legislature · Session law credit printed on the section: 2023 c 191 s 12 · Retrieved September 2026The exemption for information that is deidentified in accordance with the requirements for deidentification set forth in 45 CFR Part 164 and that is derived from the health-care-related information listed in that subsection. The separate exemption for information that is part of a limited data set, as defined, and is used, disclosed and maintained in the manner required by 45 CFR 164.514.

Frequently asked questions

Does de-identified mean anonymous?

Not under these definitions. The federal regulation sets a standard and two permitted methods for meeting it. One of those methods is a documented expert judgment that the risk of identification is very small. That judgment is made against other reasonably available information and an anticipated recipient. The same section also lets the entity keep a re-identification code, so long as the code is not derived from information about the individual and the mechanism is not disclosed. Meeting the standard is about following a specified method, not about achieving anonymity in the ordinary sense.

What are the eighteen identifiers?

They are the categories the federal safe harbor method requires to be removed, for the individual and for relatives, employers and household members. Names. Geographic subdivisions smaller than a state, with a narrow exception for the first three digits of a zip code where the covered area holds more than twenty thousand people. All elements of dates except year for dates directly related to a person, and ages over eighty-nine. Telephone and fax numbers, email addresses and social security numbers. Medical record numbers, health plan beneficiary numbers, account numbers, and certificate and license numbers. Vehicle identifiers, device identifiers, web addresses, internet protocol addresses and biometric identifiers. Full face images. And any other unique identifying number, characteristic or code.

Is removing my name enough?

The safe harbor method does not stop at names, and it does not stop at the list either. After the eighteen categories are removed, the entity must also not have actual knowledge that the remaining information could be used, alone or in combination with other information, to identify an individual. That second condition is part of the method rather than an afterthought, and it is written as a duty on the entity holding the data.

What is a limited data set?

It is a defined category in the same federal section, and its definition begins by calling it protected health information. It excludes a shorter list of direct identifiers than the de-identification method does, and dates survive it. It may be used or disclosed only where the entity enters into a data use agreement with the recipient. A dataset described as having had direct identifiers stripped out may be this, and this has not left the federal rulebook.

Why do the state definitions include a promise?

Because they were written that way. California defines deidentified information as information that cannot reasonably be used to infer information about, or be linked to, a particular consumer. Three conditions then attach to the business that holds it. It takes reasonable measures, publicly commits to maintain and use the information in that form without attempting reidentification, and contractually obligates recipients to do the same. Washington's health data statute uses the same three conditions and extends the linkage question to a device linked to a consumer. Two of the three are undertakings rather than technical properties.

Does calling data de-identified change which law applies to it?

It can. The Washington act states that personal information does not include deidentified data, and consumer health data under that act is built on personal information. Its exemption section separately excludes information deidentified in accordance with the federal privacy regulations, where that information is derived from the health-care information the section lists. It also excludes information that is part of a limited data set, used and maintained as the federal provision requires. A classification decision can therefore move records outside a statute.