Supprimer les pages en double v1.0

Trouvez et supprimez les pages répétées

Remove Duplicate Pages is a free browser tool that finds pages repeated inside a single PDF and keeps one copy of each. It exists for the ordinary scanner accident: a sheet pulled through the feeder twice, producing two pages that look the same but match in almost nothing else, because noise, skew, and exposure differ on every pass. The DocPivot version runs entirely on your own machine, shows which pages it plans to drop before anything is deleted, and reports the count once the work is done. Nothing is uploaded, which matters for the medical records, court bundles, and bank statements where double feeds usually turn up.

What DocPivot Remove Duplicate Pages Does

DocPivot Remove Duplicate Pages compares every page in one PDF against the pages it has already decided to keep, then rebuilds the file without the repeats. Each page is read twice: once for its text, normalized so that spacing and line breaks do not create false differences, and once for its appearance, rendered on a white ground and reduced to a 32 by 32 grid of average brightness values. Those two signatures are what the comparison runs on, not the raw bytes of the file, which is why a page scanned twice can still be recognized as one page.

The people who need this are working with paper that became a PDF. Legal assistants assembling discovery, clinic staff digitizing patient files, accountants handling a year of receipts, and anyone who has run Merge PDF across a stack of overlapping exports all end up with the same problem. Batch scanners double feed. Merged files repeat cover sheets. A 300 page bundle can carry 40 pages of pure repetition that nobody has time to find by eye.

Before this tool, the options were to scroll thumbnails and delete by hand, which is slow and unreliable past about 50 pages, or to run a desktop plug-in. Automatic detection turns a half hour of squinting into a few seconds of work, and the file shrinks with it, so a later pass through Compress PDF starts from a smaller document. If the source scan has no text layer at all, running OCR PDF first gives the comparison more to work with, although it is not required.

How DocPivot Remove Duplicate Pages Decides Two Pages Match

Text decides, and the picture only confirms. If two pages carry different words, they are different pages, and no amount of visual similarity overrides that. This veto is the safety property of the whole tool, because the pages most likely to look alike are the ones where a single changed field carries all the meaning: invoice numbers on a template, dates on a form, account balances on a statement.

The visual fingerprint gets a say in two situations. When the words match, it has to agree before a page is dropped, which catches the case of a form repeated with one field altered in a way the text layer records poorly. When there are no words at all, because the page is a photograph of paper, the fingerprint becomes the only evidence available. A page is removed only when both signals point the same way, and it survives whenever either one objects.

Grouping works against the survivors rather than pair by pair. A page is tested against the copies already marked as keepers, so a sheet that was fed through three times collapses to one page instead of leaving a stray second copy behind. The same logic means the tool reports its work honestly: you see a line such as "Found 2 duplicates. Pages 2, 3 would be removed, leaving 1" rather than a silent file swap. Once the cleanup is finished, a tool such as PDF to Text gives you a quick way to confirm that nothing you needed left the document.

How DocPivot Remove Duplicate Pages Works

  1. Open the file locally. The PDF is read with pdf.js and displayed as a page grid. It is never sent anywhere.
  2. Read every page twice. The tool extracts normalized text and renders the brightness fingerprint for each page.
  3. Compare and group. Text is checked first, appearance second, and repeats are grouped against the pages already kept.
  4. Review the marks. Candidates appear highlighted in the grid and listed by page number in the results panel.
  5. Rebuild and download. pdf-lib writes a new PDF containing the surviving pages in their original order, and the download starts only when you ask for it.

Reading a 300 page document takes roughly 6.5 seconds, most of it spent rendering pages for the fingerprint step. Because the work happens on your own processor, older hardware and very large scans will run slower, and a phone will be slower than a laptop.

Settings in DocPivot Remove Duplicate Pages

Three controls decide how aggressive the cleanup is, and the DocPivot defaults suit the double feed case without further adjustment.

ControlOptionsDefault
How alike counts as duplicateIdentical copies only, Same page re-scanned, Same page poor scanSame page re-scanned
Which copy to keepThe first one, The last oneThe first one
Pages to checkAny range, for example 1-4, 8, 15-Empty, meaning check them all

Strictness. Identical copies only is for digital files where a page was pasted in twice and the two copies really are the same. Same page re-scanned is the setting for feeder accidents. Same page poor scan widens the tolerance further for faded carbon copies, skewed sheets, and low contrast faxes, at the cost of more false matches to review.

Which copy survives. Keeping the first copy preserves the original order of a document. Keeping the last one is useful when somebody rescanned a page because the first attempt came out crooked or half cut, since the later version is usually the better one. Either way, page order in the output matches the input, so anything you later do in Organize PDF starts from a predictable sequence.

Page range. A range fences off the part of the document the tool is allowed to touch. Pages outside it are never examined and never removed, whatever they look like. A mistyped range examines nothing rather than deleting something, which is the safer direction for a destructive operation to fail in.

When to Use DocPivot Remove Duplicate Pages

Reach for this tool when a PDF is longer than its content, and the extra length is repetition rather than clutter. It is most valuable on documents you did not create yourself, where you cannot know in advance which pages repeat.

  • Batch scanner output. Feeders grab two sheets at once often enough that long scans usually contain at least one double.
  • Merged exports. Combining monthly statements or report versions repeats cover pages and summary sheets.
  • Received case files. Medical and legal records passed between offices accumulate the same page many times over.
  • Archive cleanup. Shrinking a stored bundle before it goes into long term storage saves real space across hundreds of files.
  • Print preparation. Removing repeats before printing saves paper on documents that run to hundreds of pages.
  • Pre-review passes. A shorter bundle costs less to read when somebody bills by the hour for reading it.

Skip the tool when you already know exactly which pages to drop, since Remove Pages from PDF does that in one step, and when the repetition you want gone is a recurring header or footer rather than a whole page.

Practical Workflows for DocPivot Remove Duplicate Pages

Cleaning a Double-Fed Contract Scan

Context: A 120 page signed contract comes back from the scanner at 137 pages.

  • Open the file and leave strictness at the re-scanned default.
  • Check the marked pages in the grid against their neighbors.
  • Keep the last copy, since repeat scans here were corrections.

Result: The rebuilt file returns to 120 pages with the cleaner version of each rescanned sheet retained.

Deduplicating a Merged Medical Record

Context: Three clinics send overlapping histories for the same patient.

  • Combine the files, then run the duplicate check across the whole result.
  • Set the range to skip the first few pages if each source starts with an index.
  • Review flagged lab reports carefully before confirming.

Result: One consolidated record without three copies of the same discharge summary, ready for Rotate PDF if any source arrived sideways.

Trimming an Archive Before Storage

Context: A finance team stores 400 scanned invoice batches per year.

  • Run each batch through the check at the default setting.
  • Note the reported count as a rough quality signal for the scanning process.
  • Pull anything genuinely unique into its own file with Extract Pages from PDF.

Result: Smaller archives and a running record of which scanning stations jam most often.

What DocPivot Remove Duplicate Pages Cannot Do

Duplicate has no definition in the PDF format, so every answer this tool gives is a judgment rather than a fact. That honesty shapes several real limits worth knowing before you press Download.

  • It works on one file at a time. Cross-file comparison is not supported, so overlapping documents have to be combined first.
  • It cannot tell you two pages belong to different documents. A repeated cover sheet or a blank separator is a duplicate by every measure the tool has, and removing it is the tool doing its job. Pages you want kept should sit outside the range you check, or be restored afterward.
  • Poor scan mode produces false matches. Widening tolerance to catch faded pages also groups pages that merely resemble one another, which is why review comes before removal.
  • There is no undo inside the tool. The output is a new file and your original is untouched on disk, so recovery means going back to the source rather than reversing a step. Keeping that original is worth doing, and Repair PDF is the right stop if the source itself is damaged.
  • Very large files depend on your device. Browser memory sets the ceiling, and a thousand page scan on a phone may struggle where a laptop does not.

How DocPivot Remove Duplicate Pages Compares With Other Options

Most well known online PDF services do not detect duplicates at all. Smallpdf states plainly that its tool cannot do this automatically and recommends merging files first, then removing repeated pages by hand with thumbnail previews and zoom to check them. iLovePDF and PDF24 work the same way: you open the page grid, you spot the repeats yourself, and you click them away, with PDF24 noting that deletion happens on its servers rather than in your browser. Adobe routes the same task through an account sign-in. Automatic detection has traditionally meant a commercial plug-in for Adobe Acrobat, such as Evermap AutoSplit, which compares pages by text or by visual appearance and presents the matches for review.

A small group of dedicated deduplication sites has appeared more recently offering automatic near-duplicate detection in the browser with adjustable similarity thresholds. Client-side processing and near-duplicate matching are therefore no longer rare on their own. What differs in the DocPivot tool is the decision rule and the guardrails around it.

CapabilityMainstream online PDF toolsRemove Duplicate Pages
Finds repeats for youNo, manual thumbnail reviewYes, automatic
Where the file is processedCommonly uploaded to a serverIn your browser
Re-scans of one sheetFound only by eyeMatched by text and appearance
Different words on similar pagesHuman judgmentNever removed, text vetoes
Limit the search areaNot applicableAny page range
Says what it didSilentReports the count, including zero

The text veto is the part worth weighing. Tools that score similarity as a single percentage can remove a page that differs by one number, because one changed field barely moves a similarity score. Requiring the words to match before appearance is consulted trades a little recall for the kind of mistake that is expensive to discover later. Reporting a zero result matters for the same reason, since a tool that returns your file unchanged without comment looks broken rather than thorough, and the count gives you something to check against a quick pass in PDF Editor.

Frequently Asked Questions About DocPivot Remove Duplicate Pages

Signaler un bug

NOUS CONTACTER

info@toolspivot.com

ADRESSE

Ward No.1, Nehuta, P.O - Kusha, P.S - Dobhi, Gaya, Bihar, India, 824220

Nos outils les plus populaires