OneDrive · How it works

By Pedro Albaladejo · ClearCopies

Updated August 28, 2026 · 7 min read

OneDrive Duplicate File Benchmark: 102,610 Files Scanned

See what a read-only ClearCopies scan found across 102,610 OneDrive files, how exact groups were classified, and what the result does not prove.

Large-drive duplicate claims are easy to inflate. Counting repeated names, equal sizes, or every item after the first candidate can produce a dramatic number without proving that the underlying bytes match. This benchmark records the narrower result produced by ClearCopies' exact-evidence rules on a real six-figure-file OneDrive library.

The scan was read-only. It enumerated provider metadata and used comparable OneDrive fingerprints together with exact byte size; it did not download original file bodies for hashing. The published figures are aggregate and anonymized: they do not expose account identity, filenames, paths, sharing details, or individual file records.

Decision snapshot

What the production scan measured

Each figure describes one stage of the same read-only metadata inventory.

Observed valueMeasurementInterpretation
102,610Files indexedThe inventory scale before duplicate grouping
584 GBMetadata inventoryCombined reported size of the indexed file library
16,718Exact groupsGroups requiring equal byte size and comparable fingerprint evidence
22.4 GBExtra-copy estimateStorage beyond at least one retained file in each exact group

Reviewed August 28, 2026. This is one anonymized library, not a universal duplicate-rate forecast.

Anonymized production benchmark · August 2026

Validated on a six-figure-file OneDrive library

A completed read-only production scan indexed 102,610 files across 584 GB. Exact size-and-fingerprint matching identified 16,718 groups and 22.4 GB of recoverable extra copies. Similar or uncertain files were kept outside that exact total.

102,610
files indexed
584 GB
metadata inventory
22.4 GB
exact recovery estimate

This benchmark exposed the difference between broad possible-match totals and exact recoverable storage, so the product now keeps those classifications visibly separate.

Written by ClearCopies Editorial · Technical review by ClearCopies Engineering

A practical checklist

  1. 1Inventory the complete scope and record whether enumeration finished successfully.
  2. 2Use byte size only to narrow candidates, never as exact proof by itself.
  3. 3Require a comparable provider fingerprint before creating an exact group.
  4. 4Calculate extra-copy storage only after retaining at least one item per group.
  5. 5Keep uncertain, unsupported, protected, or operationally required files outside the recovery estimate.

How the benchmark classified exact copies

The pipeline first grouped candidates by byte size, then compared a OneDrive-supplied content fingerprint such as quickXorHash or SHA-1 when the provider returned one. Algorithm identity travels with the value, so unlike fingerprint types are never compared as though they were interchangeable. A filename, timestamp, thumbnail, or equal size can help a reviewer investigate but cannot promote a candidate into the exact total.

Pagination, missing hashes, changing files, incomplete searches, and provider errors must remain visible. A clean-looking result is not useful if the inventory silently omitted part of the drive. The benchmark therefore reflects completed enumeration and the conservative exact grouping that the running product exposes to its review interface.

  • Candidate: matching size or contextual clue
  • Exact group: matching size plus comparable fingerprint
  • Uncertain item: insufficient or incompatible evidence

What 22.4 GB of recoverable storage means

The estimate counts only copies beyond the retained baseline in each exact group. If three byte-identical items appear together, at least one remains; the estimate considers at most two extra copies. It is therefore different from the total size of every file participating in a duplicate group.

Exact content equality still does not decide which file item is redundant. A separate path may own the active shared link, useful version history, a retention obligation, or a canonical project location. ClearCopies keeps these items visible so the user can select the keeper rather than accepting a filename-based automatic choice.

What this benchmark does not claim

One account cannot establish a universal duplicate percentage for OneDrive users. Libraries differ by camera imports, collaboration, migrations, backups, creative workflows, file types, retention requirements, and age. The result should be treated as evidence that the method operates at six-figure-file scale, not as a promise that another account will recover the same share of storage.

It also does not measure visual near-duplicates, revised documents, recompressed media, exported Google Workspace files, or cross-provider copies. Those items normally have different bytes and require a different comparison method. Keeping them outside the exact total prevents a broad similarity claim from being mistaken for safe cleanup evidence.

How to reproduce a defensible measurement

Record the account scope, scan completion state, file count, total byte count, supported fingerprint algorithms, and unresolved items. Preserve aggregate output before cleanup so a later reviewer can understand how the estimate was produced. If the library changes substantially, scan again rather than applying an old decision to new files.

For a cleanup plan, add ownership, sharing, location, protected-format, and retention checks after exact detection. Revalidate every selected item immediately before a later move to recoverable trash. Detection proves content equality; review and revalidation make the operational decision safer.

Limits and risks to check

  • The benchmark is one library and must not be presented as a universal duplicate rate.
  • Provider fingerprints can be absent or unsupported for some items.
  • Exact content equality does not remove sharing, ownership, retention, or version-history obligations.
  • Storage dashboards and recycle-bin behavior can delay visible quota changes after cleanup.

Official references

Frequently asked questions

Did ClearCopies download 584 GB to run this scan?

No. The initial scan used OneDrive metadata, byte sizes, and provider-supplied fingerprints. Original file bodies were not downloaded for duplicate detection.

Does 22.4 GB mean every listed file can be deleted?

No. It is an aggregate extra-copy estimate after retaining a baseline item in each exact group. Every candidate still needs keeper, sharing, ownership, history, protection, and retention review.

Can this result predict savings for another account?

No. It demonstrates the method and production scale. Another library needs its own completed read-only scan because duplicate patterns and operational requirements vary widely.

Scan first. Decide with evidence.

ClearCopies reads OneDrive metadata and groups exact copies by byte size plus quickXorHash or SHA-1. Original file bodies are not downloaded for the scan. You review the result and export a plan before any separate write step.