PDFSanitize Overview and Operating Model

Applies to: All editions

Purpose

Use this page to understand what PDFSanitize does, what happens during analysis and sanitization, and how the source PDF, sanitized output, verification record, and exported evidence relate to each other.

What PDFSanitize processes

PDFSanitize works with PDF files selected individually or discovered under a selected folder tree. Multiple PDFs can be included in one working list and processed as a batch. Each file maintains its own analysis, sanitization, and verification state.

PDFSanitize does not edit the source file in place. A successful sanitization creates a new PDF in the source file’s directory. The original remains available unless it is later removed under a separate retention procedure.

Operating model

A normal workflow has four stages:

  1. Add PDFs or folders to the PDF list and select the documents to include.
  2. Analyze the selected documents to identify relevant metadata, structure, interactive features, and selected forms of hidden content.
  3. Apply a predefined profile or permitted custom options and create a sanitized output.
  4. Verify the output against the policy that was actually used for that sanitization.

If sanitization is requested for a selected PDF that has not yet been analyzed, PDFSanitize can run the missing analysis first and continue only if the analysis succeeds.

Preserve Quality and Maximum Assurance

Preserve Quality removes selected PDF structures while retaining normal source text, vector graphics, and page-level PDF content wherever possible. When hidden or completely off-page text is selected for removal, only affected pages can be reconstructed from rendered content.

Maximum Assurance reconstructs every page from rendered content. This creates a new document whose pages no longer depend on the original page-level text, vector objects, forms, hyperlinks, or similar structures. The tradeoff is that selectable source text, original vector objects, and interactive document behavior are not retained in their original form.

Analysis, sanitization, and verification are separate states

Analysis reports what PDFSanitize detected before sanitization. Sanitization records the exact profile, mode, options, timestamps, output path, page reconstruction information, and available hashes. Verification then reopens the sanitized output and checks the requirements associated with the options that were actually applied.

Changing the current options after a file has been sanitized does not change the verification policy for that existing output. Verification uses the policy snapshot stored with the sanitization result.

Local document processing

PDF parsing, rendering, sanitization, verification, and report generation occur on the Windows computer. The PDF content itself does not need to be sent to an online service. Pro and Enterprise licensing can require network access for activation and periodic validation, and the application can contact Probativa services for executable-integrity checks.