Open data · CC0 public domain

Download the benchmark 10.5 MB MANIFEST.csv table only

The Redaction Benchmark

73 test documents for checking whether a PDF redaction tool actually works.

Version 1, published 23 September 2026. Free to use for any purpose, including commercially, including by people who sell competing tools.


What this is for

When somebody covers a name or a number on a document by drawing a black box over it, the words are very often still sitting underneath, and anybody who receives the document can get them back out. It has embarrassed courts, governments, law firms and newspapers.

Plenty of programs claim to fix this. There has been no shared set of documents to test them against, so every claim has had to be taken on trust.

These are those documents. Run any tool against them and you can see for yourself what it catches and what it misses.

It is deliberately hard in both directions. Roughly a third of these documents contain nothing wrong at all. A tool that shouts about a problem on a clean document is as useless as one that stays quiet on a broken one, and false alarms are the reason people stop checking.


What is in it

Group one — the basic failures (17 documents)

The ways a document leaks, one per file. A black rectangle drawn over the words. Words printed white on white. A scan with the words typed out and stored behind the picture. An earlier saved version still inside the file. A comment stuck to the page. A second file tucked inside. Filled-in form boxes. Hidden instructions. A contents list. The file's own label.

Five of the seventeen are clean and are there to catch false alarms.

Group two — the awkward ones (11 documents)

Built to defeat a tool that takes shortcuts. Black words printed on a black block, which looks like a solid bar. A leak on page three only, to see whether every page is checked. A rotated page. A picture laid over live words.

Five of the eleven are clean, and they are the ones that separate a careful tool from a noisy one: a background graphic behind clean text, a photograph with a caption, a shaded table row drawn after the text. Each of these looks like a cover to a naive check and is not one.

Group three — scripts and layouts (45 pairs, 90 documents)

A different question. Not *did it leak*, but *can a tool reliably find a word and cover it* when the page is unusual.

Japanese written down the page. Arabic and Hebrew read right to left. Thai, which has no spaces between words. Devanagari. Joined letter pairs. Type too small to read and type filling the page. Words at the very edge of the paper. Two and three columns. Table cells and table headers. Letters spaced out. Justified text with wide gaps. Superscript. Rotated blocks. Landscape pages. A word split by a hyphen across two lines. A term appearing only on page three of many. Two hundred pages. A page that is only a photograph.

Each comes as a pair: the original page, and the same page after a tool marked the term. The pair is what lets you see whether the mark landed in the right place.


How to use it

One. Take the documents in group one and group two and put each one through the tool you are testing.

Two. Compare what it said against MANIFEST.csv, which lists every document, what is in it, and what a correct tool should report.

Three. Count four numbers and publish all four:

A tool that reports every document as a problem scores perfectly on the first number and is worthless. Publish all four or the result means nothing.

Four. For group three, run the tool's search-and-cover function over each original looking for the term the manifest names, and compare the result against the supplied .mask.pdf. What you are looking for is whether the mark landed on the word, on part of the word, or somewhere else entirely.


What this benchmark cannot tell you

Whether the picture on the page still shows something. If somebody photographed a page that had a name visible on it, no software can tell you that the name should not have been there. Only a person looking at the page can.

Whether a tool is safe. Passing every document here means it handled 73 documents correctly. It does not mean it will handle yours.

Anything about speed, price, or whether the tool uploads your document. Those are real questions and this measures none of them.

Whether the set is representative. These documents were built on purpose to contain specific failures. They are not a random sample of real documents, and the proportions here say nothing about how often each failure occurs in the world.


Where the documents came from

All 73 are synthetic and were generated by a script, which is included as generate-basic-documents.py so you can rebuild them or add your own.

No document here contains real information about any real person. The names are Alice Example, Blake Example and Casey Example. The companies are Acme Corp and Hidden Client LLC. The identity numbers are structurally invalid and cannot belong to anybody. The passwords are jokes.


Licence

The documents are released into the public domain under Creative Commons Zero (CC0 1.0). No permission is needed and no attribution is required. Use them in a product, in a paper, in a blog post, in a sales page, or to demonstrate that a tool you dislike fails.

That includes people who sell tools that compete with the one that made this. A benchmark only works if the people being measured can use it too.

Attribution is welcome and never required. If you do cite it, "The Redaction Benchmark v1, 2026" is enough.


Contributing a document

The set is short of real-world variety, and the gaps are known:

If you add one, include what it contains, what a correct tool should report, and confirmation that it holds nothing about a real person. Anything that carries real personal information will be refused, whoever it belongs to.


A note on what to do with a result

If you test a tool and it misses something, say what you tested and how, and show the document. The document is here and anybody can repeat it, which is the entire point.

Say what the tool did. Do not say what the people who made it knew or intended.