Redacted Medical Records: HIPAA's 18 Identifiers and How to Remove Them
By PDFized Team·Published on ·8 min read
A redacted medical record is a patient record. However, the difference is that in a redacted piece, PHI (protected health info) is removed for good. The redaction is done before the document is shared, filed, or published.
If we take into account the HIPAA's Safe Harbor standard, you have to remove 18 specific categories of identifiers from the medical record. The categories include different stuff, like names, dates, medical record numbers, full-face photos, and so on. If you remove all that, you can then safely disclose the document. If you fail, a breach becomes the problem.
When Medical Records Need Redaction
Unfortunately, we have to share medical records more often than we can imagine. When it happens, these docs simply leave the safety of a provider's systems. The bad news is that every time it happens, the redaction should be part of the process:
- Litigation/court filings. When it comes to the records that are used as evidence in injury, malpractice, or disability cases, it is crucial to remove all PHI. Thus, all people who are not related to the case won't see it.
- Insurance and billing disputes. When you deal with claims reviews and appeals, you have to regularly send records to third parties. But they don't need to see everything in those docs.
- Research and case studies. Publishing clinical findings requires de-identified source material.
- Sharing of docs by patients. They send their own records to advocates or look for second opinions from other experts. But in most cases, there is no need to give full docs, just a fraction of some content visible is enough.
- Staff training/audits. Real docs can be the best training material only if they do not identify real people.
Unfortunately, everything mentioned above is not just some theory. In 2025, 772 large healthcare data breaches were reported to the HHS Office for Civil Rights. As a result, the records of nearly 138.5 million people were shared with people absolutely not related to them. It is important to mention that healthcare has held the highest average breach cost of any industry for over a decade, at $7.42 million per incident. Most of those breaches are network attacks. But improper disclosure of unredacted documents is the category an organization controls completely. It is the preventable one.
Need to redact a medical record?
PDFized detects and permanently removes patient names, MRNs, dates, and every other HIPAA identifier for you – free.
What HIPAA Actually Requires
You are not going to see a lot of use of the word "redaction" in HIPAA's Privacy Rule. What you see is that it talks about de-identification and the minimum necessary standard. And redaction is how both are done on documents.
The Privacy Rule (45 CFR §164.514) provides two ways to remove identifying information from health data:
- Safe Harbor. Remove all 18 categories of identifiers listed in §164.514(b)(2). Plus, you have to be sure the remaining info cannot be used to identify this or that person. This method can be seen in all redaction workflows. Besides, we talk about it in the guide.
- Expert Determination. A qualified statistician applies accepted methods and documents that the re-identification risk is "very small". This works best for large research datasets. It's much better than sharing docs day to day.
Alongside de-identification sits the minimum necessary standard: even when disclosure is permitted, you share only the information the recipient needs. In practice, that means redacting a record down to the relevant visit, diagnosis, or date range – not handing over the full chart because one page of it matters.
Getting this wrong is expensive. Following the penalty update published in the Federal Register on January 28, 2026, HIPAA civil monetary penalties now range from $145 to $2,190,294 per violation depending on culpability, and state attorneys general can bring separate actions. Individual OCR settlements regularly land in six and seven figures.
The 18 HIPAA Identifiers
Safe Harbor requires removing all of the following for the patient and for the patient's relatives, employers, and household members:
| # | Identifier | Where it hides in real records |
|---|---|---|
| 1 | Names | Headers, signatures, "reviewed by" lines, family history |
| 2 | Geographic data smaller than a state | Addresses, city, county, ZIP (first 3 digits of a ZIP may remain if the area exceeds 20,000 people) |
| 3 | All dates (except year) tied to the individual | Birth, admission, discharge, procedure, death dates; all ages over 89 |
| 4 | Phone numbers | Contact blocks, emergency contacts, provider callbacks |
| 5 | Fax numbers | Transmittal headers and cover sheets |
| 6 | Email addresses | Portal correspondence, contact details |
| 7 | Social Security numbers | Intake forms, billing, insurance sections |
| 8 | Medical record numbers | Every page header, labels, barcodes |
| 9 | Health plan beneficiary numbers | Insurance cards, claims, EOB references |
| 10 | Account numbers | Billing statements, payment records |
| 11 | Certificate/license numbers | Driver's licenses on intake, professional licenses |
| 12 | Vehicle identifiers and license plates | Accident and ambulance reports |
| 13 | Device identifiers and serial numbers | Implant records, pacemaker and pump logs |
| 14 | Web URLs | Portal links tied to the patient |
| 15 | IP addresses | Telehealth logs, audit trails |
| 16 | Biometric identifiers | Fingerprints, voiceprints |
| 17 | Full-face photographs and comparable images | ID photos, wound and dermatology photos showing identifying features |
| 18 | Any other unique identifying number, characteristic, or code | Study IDs, tattoos noted in exams, rare job titles in small towns |
The last category is the one that catches people: "the only firefighter in a town of 400 with this diagnosis" identifies someone as surely as a name does. Safe Harbor also requires that you have no actual knowledge the remaining data could identify the patient – the list is the floor, not the whole test.
Redacted Medical Record Example: Before and After
A properly de-identified record keeps the clinical content intact while removing the patient's identifiers:

How to Redact Medical Records Correctly
Step 1 – Work on a copy, inventory the document
Medical records are compound documents: typed notes, scanned faxes, lab PDFs, images. Note which pages are scanned, because image-based pages need OCR before any text-based tool can find identifiers in them.
Step 2 – Find every instance of the 18 categories
Manual review means reading every page against the checklist above, including headers, footers, margins, and stamps – identifiers like MRNs repeat on every page, and dates hide inside narrative text ("seen three days after her March 12 admission"). Automated detection reads the full text layer and flags names, dates, numbers, and addresses pattern-wide, which is the practical difference on a 200-page chart. A dedicated healthcare redaction tool is tuned to exactly these categories (MRNs, member IDs, dates of service) and processes the document in one pass.
Step 3 – Redact, don't cover
True redaction deletes the underlying text and image data. Drawing black rectangles in an editor or annotator leaves the text selectable underneath – the most common failure mode in published court documents. If the visual layer changes but the text layer doesn't, it isn't redacted.
Step 4 – Strip the metadata
PDF metadata, embedded file properties, and revision history can carry patient names and provider details even after the visible content is clean. Flatten and sanitize the file as part of the export.
Step 5 – Verify before release
Search the finished file for known identifiers (try selecting and copying "redacted" areas – if text copies out, it's still there), check every page including scanned ones, and confirm the file properties are clean. For a deeper look at what failure looks like, see what happens when redaction fails.
Redaction vs. De-Identification vs. Anonymization
The terms overlap but aren't interchangeable. Redaction is the document-level act of removing specific content. De-identification is HIPAA's legal status – a record that satisfies Safe Harbor or Expert Determination is no longer PHI and falls outside the Privacy Rule. Anonymization is the stronger, mostly GDPR-associated concept of making re-identification impossible by any party. Redacting the 18 identifiers is how a document typically achieves de-identified status; whether it's "anonymous" in the European sense is a separate, stricter question. If your documents also contain financial or general personally identifiable information, the same pass should cover those categories too – a hospital bill mixes PHI with account and card data on one page.
// faq
FAQ
No. HIPAA defines what must be removed (the 18 identifiers under Safe Harbor), not the software or technique. What matters is that the information is actually unrecoverable – from the visible layer, the text layer, and the metadata.
On paper, a marker plus a photocopy of the marked page works. Digitally, drawing black shapes over text is not redaction – the text remains beneath and can be extracted in seconds. Use a tool that removes the content itself.
Under Safe Harbor, all date elements except the year – including admission and discharge dates – must be removed, and ages over 89 must be aggregated. Many people redact only the birth date and leave service dates behind; that record is not de-identified.
If all 18 identifiers are properly removed and no actual knowledge of re-identifiability remains, the record is de-identified and no longer regulated as PHI. If even one identifier survives – a MRN in a page footer, a date in a sentence – it remains PHI, and disclosing it is a reportable event.