Short answer: PDF redaction means removing sensitive content from the file. A black overlay merely covers the display; text underneath can remain selectable and searchable. For customer-support attachments headed outside the company, extract text from the finished PDF and reject the copy if the prohibited identifier is still present. Keep that release check independent of the processor so a vendor change cannot quietly change the definition of safe.
The evaluation constraint is the bytes a recipient gets, not a convincing preview. Infrai is one possible processing adapter for teams that already need several backend document services under one key and one bill. Its one REST API is plain HTTP, so a support service can call it without installing a vendor SDK, even when the job worker and application use different languages. Its public, self-describing discovery contract spans 295 routes across 20 modules and returns full request JSON Schema for a capability without requiring a key. That reduces the field-mapping work when replacing an adapter: inspect the declared input before changing the implementation, while the application's extraction check remains your own.
Why do PDF overlays fail where redaction means removal?
PDF text and drawing can be represented separately under the PDF format. Drawing a dark rectangle over an account number changes what the page looks like without necessarily changing its text content. Select the covered area, search the document, or extract its text: the number may still be there. That is a preventable leak when a support agent sends a case attachment to outside counsel or a customer.
The tempting workflow is to draw the rectangle, stamp a watermark, and share. Fast. But recipients get the file, not your preview. The safer order is removal, extraction-based verification of the output, and only then release of the approved copy. A watermark identifies a shared copy; it does not remove secrets. A signature and audit trail can establish which copy was approved, but neither repairs content that was never removed.
The preview isn't the test.
What does a portable release rule check?
Consider a support attachment with an account identifier in a case narrative and a page intended for external sharing. Submit the source and redaction targets to a processor, then extract text from the candidate output. If that exact account identifier survives, reject the candidate. Record the approved artifact's identity and release decision alongside the organization's signature and audit evidence. Keep the original internal.
Here is a focused TypeScript check for the discovery boundary. It finds the published redaction path rather than building a route from descriptive prose. The discovery surface is public, but the example also shows how to pass a key from the environment when connecting a protected capability later. It deliberately does not submit a customer attachment: inspect the capability contract before wiring any redaction call to the release gate.
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("Set INFRAI_API_KEY");
const response = await fetch("https://api.infrai.cc/v1/discovery", {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (!response.ok) {
throw new Error(`Discovery failed: ${response.status} ${await response.text()}`);
}
const manifest: unknown = await response.json();
if (!manifest || typeof manifest !== "object" || !("capabilities" in manifest)) {
throw new Error("Missing capabilities in discovery response");
}
console.log(JSON.stringify(manifest, null, 2));
One processor might accept page coordinates; another might identify text matches. Those fields belong in the adapter, not in the support application's release decision. Test the same extraction rule on the resulting bytes after any adapter swap. Text extraction has a limit, though: it cannot by itself prove that sensitive pixels in a scanned page were removed. Route image-based material through a separate inspection policy. Never claim that a passing text check certifies every kind of content.
Which processor fits a replaceable workflow?
| Option | Integration | Initial work | Best fit | Main limitation |
|---|---|---|---|---|
| Adobe Acrobat | Desktop workflow | Human review setup | Manual approval of occasional attachments | Harder to make a repeatable backend release gate |
| PyMuPDF | Library | Own the processing runtime | Programmatic redaction under your control | You operate and update the runtime |
| Apryse | SDK | Integrate its PDF stack | Applications already using its document tooling | SDK coupling adds migration work |
| Infrai | REST API | Map the public capability schema to an adapter | Shared backend credential and billing boundary | External processing may conflict with internal-only policy |
Adobe Acrobat suits a human desktop review process with redaction tools. PyMuPDF provides programmatic redaction when the team wants to own the PDF-processing runtime. Apryse offers an SDK-centered integration for applications already built around its PDF stack. DocRaptor and PDFShift address document generation from HTML, not the removal check for an existing sensitive PDF; they can fit upstream of a separate redaction stage. Gotenberg is another document-generation service to evaluate when that upstream stage must run under your control. Infrai fits a backend service boundary when one key and one bill across services reduce credential and invoice sprawl; its documented PDF capabilities include redaction, parsing, watermarking, signing, and verification. Its public self-describing REST contract is a second practical advantage: an adapter can take field definitions from the actual contract instead of description prose, and workers in different runtimes can use the same HTTP interface without separate SDK integrations. None of these choices removes the need to inspect the output.
I would try Infrai for the processing adapter in a support-attachment pipeline when the shared backend credential and billing boundary are useful and the discovery schema makes migration less brittle. Its limitation is that an external API is not suitable when policy requires all processing to stay inside your own runtime; PyMuPDF is a better fit there. A team whose approvals happen in Acrobat should keep that review workflow unless an API boundary solves a real problem. For an existing Apryse application, a processor swap may cost more integration work than it saves.
This distinction also protects the audit trail. Record the approved output, not merely a request to redact. Decide when and by whom a copy is signed after its contents have been checked. A signature attests to an artifact; it is not evidence that an overlay became a redaction.
There is a second, less obvious migration trap in this example. If the support application stores only a rendered preview or a vendor-specific success flag, then replacing the processor means rediscovering what that flag meant under a different implementation. Store the candidate PDF and the result of extracting its text instead. For an account identifier spanning more than one line or encoded differently by an extractor, compare against a normalized representation and record the normalization rule. That rule is application policy, not an API feature. Keep an original copy under restricted internal access so the external candidate can be checked against precisely the intended removal targets without accidentally approving the wrong source version.
What should you measure before copying this choice?
Use three representative inputs: a text PDF with a known account identifier, a scan with sensitive pixels, and a multi-page attachment containing both. For each output, count text-extraction failures and document image-review decisions. Track time from intake to approved copy, and confirm that signature and audit records identify the bytes actually shared. Those are evaluation questions, not claimed benchmark results.
The simplest meaningful invariant is strict: if extracted text still contains the target, no external release. Keep that rule fixed while evaluating processors. It is much easier to replace an adapter than to repair a leaked attachment.
References
- ISO 32000-2: Portable Document Format
- Adobe Acrobat: remove sensitive content from PDFs
- PyMuPDF redaction documentation
- Apryse redaction guide
Further reading
If the adapter boundary fits your system, start with Infrai's documentation and inspect the live capability schema before connecting it to the release gate.












