Redacting a PDF properly means removing the sensitive content from the file, not just hiding it from view. A black rectangle drawn with a highlighter or shape tool only covers the text visually. The words are still in the document, and anyone can copy them out, search for them, or delete the rectangle. Real redaction deletes the underlying text, images, and vector graphics in the marked area, so there is nothing left to recover.
This guide explains why the black box method fails so often, what a proper redaction workflow looks like, and how to check your work before a document leaves your hands.
Why drawing a black box is not redaction
A PDF is built in layers. Text, images, and shapes are separate objects, placed on the page one after another. When you draw a black rectangle in a viewer or an annotation tool, you add a new object on top. The text underneath is unchanged.
That leaves several easy ways to recover the "hidden" content:
- Select and copy. Drag across the black area and paste into a text editor. The covered words often appear in full.
- Search. Press Ctrl+F and search for a name you suspect is under the box. The viewer will happily highlight it.
- Move or delete the box. If the rectangle is an annotation, anyone with a PDF editor can click it and press Delete.
- Extract the text. Any text extraction tool reads the text layer directly and ignores what is drawn on top.
- Convert the file. Exporting to Word or another format often brings the text along while the overlaid shape gets dropped.
Public reporting over the years has included court filings and official documents where "redacted" passages were recovered simply by copying and pasting. The lesson is not that people were careless. It is that the tools they used looked like they did one thing while actually doing another.
Other places sensitive data hides
Even with the visible text handled, a PDF can leak information in less obvious places:
- Document metadata, such as the author's name, the organization, the software used, and editing dates.
- Comments and annotations, including resolved review notes.
- Form fields, whose values may sit in the file separately from what is printed on the page.
- Bookmarks and outlines, which can contain section titles that name people or projects.
- Hidden layers in some design and engineering exports.
- Embedded attachments placed inside the PDF.
- Earlier revisions, because some editors save changes by appending them to the end of the file, leaving older content present.
A careful redaction workflow addresses the visible content and these extras.
What proper redaction involves
There are two broad approaches to redaction that actually removes content.
Object-level removal
A redaction engine can find every text character, image fragment, and vector path inside the marked area and delete them from the page's content, then paint a solid box in their place. Done well, this keeps the rest of the page as selectable text. Done imperfectly, it can leave fragments behind, such as a character whose outline crosses the edge of the box.
Rasterizing the page
The other approach is to render the page to an image, paint the redaction boxes into that image, and replace the original page with the flattened picture. Because the whole page is now pixels, nothing from the old text layer survives on that page. The trade-off is that the redacted page is no longer selectable text, and the file may grow somewhat.
Our free Redact PDF tool uses the rasterizing approach. Pages that contain a redaction box are rendered to an image with the black areas burned in, so the text, images, and graphics that were under the boxes are permanently gone from those pages. Pages without any redactions are left as they were.
A step-by-step redaction workflow
Treat redaction as a small process rather than a single click. It saves embarrassment later.
1. Work on a copy
Keep the original, unredacted file somewhere safe and clearly named. Redaction is destructive by design, and you may need the original again.
2. Decide what must go
Before touching the file, list what you are removing. Common categories include:
- Names, signatures, and initials of people not party to the matter
- Home addresses, phone numbers, and personal email addresses
- Government ID numbers, account numbers, and dates of birth
- Medical or financial details unrelated to the purpose of sharing
- Internal project names, prices, or negotiation positions
If you are redacting for a legal, regulatory, or public records process, the rules about what must be withheld and how the redactions should be marked vary by jurisdiction and context. Follow the instructions you were given, and ask a qualified professional when you are unsure.
3. Mark every instance
Search the document for each term on your list, because the same name or number often appears in a header, a footer, a table, and a signature block. Draw boxes generously, with a little margin around the text, so that descenders and accents are fully covered.
Using the Redact PDF tool:
- Open the tool and choose your PDF.
- Draw a box over each area you want removed, on every page where it appears.
- Choose the output resolution: 150, 200, or 300 DPI. Higher values keep the rest of a redacted page sharper, especially small print, at the cost of a larger file.
- Apply the redactions and download the new file.
Everything runs locally in your browser, so the unredacted document is never uploaded to a server. For material sensitive enough to need redaction, that is a meaningful difference.
4. Clean the metadata
Redaction handles the page content. Next, strip the document properties with our free Remove metadata tool, which removes the author, title, software information, and XMP data. A report can be perfectly redacted and still carry the name of the person who wrote it in its properties.
5. Flatten forms if relevant
If the document had fillable fields that you did not redact, consider running it through Flatten PDF so the values become part of the page and cannot be edited by the recipient.
6. Verify before sending
This is the step most people skip. Open the final file and try to break your own redactions:
- Press Ctrl+F and search for each term you removed. There should be no results on redacted pages.
- Drag-select across each black box and paste into a text editor. Nothing from the covered area should appear.
- Run the file through PDF to Text and read the output for anything that should not be there.
- Check the document properties in your viewer to confirm the metadata is gone.
- Zoom in at high magnification on box edges to make sure no partial characters peek out.
If any check fails, go back to the original, fix the marking, and redo the process rather than patching the output.
Mistakes that undo good redaction
Redacting in one place but not another. A name removed from the body text but left in a running header is a common leak. Search for every term.
Relying on a viewer's "hide" or "highlight" feature. Anything that can be toggled off was never removed.
Changing the text color to white. White text on a white page is still text.
Cropping instead of redacting. Most cropping tools hide content outside the visible area without deleting it. Our Crop PDF tool is great for trimming margins, but cropping is not a substitute for redaction.
Leaving context clues. A box that exactly fits a short word, surrounded by revealing context, can sometimes let a reader guess what was removed. Consider redacting the surrounding phrase as well.
Forgetting attachments and comments. If the original file had them, check that the final version does not.
Native options worth knowing
Some desktop PDF editors, often in their paid tiers, include dedicated redaction features that mark content and then apply removal as a separate step. If you use one, make sure you actually apply the redactions, since marked-but-unapplied redactions typically just display as boxes. Basic viewers and many free markup tools only draw shapes, which is exactly the black box problem described above. When in doubt, verify with the copy-and-search test.
Printing to paper, redacting with a marker, and rescanning also works in principle, but marker ink can be translucent under a scanner's bright light. If you go that route, check the scan carefully.
Key takeaways
- A black box drawn on top of a PDF hides text visually but leaves it in the file.
- Proper redaction removes the underlying content; rasterizing redacted pages guarantees nothing from the old text layer remains on them.
- Search every term, mark every occurrence, and draw boxes with a margin.
- Remove metadata and flatten forms after redacting page content.
- Always verify by searching, copying, and extracting text from the final file.
Redaction is one of those tasks where the cost of a small mistake is high. Build the verification step into your routine, and use a tool like Redact PDF that removes content instead of covering it.