OneDrive · How it works
By Pedro Albaladejo · ClearCopies
Updated August 28, 2026 · 7 min read
OneDrive Duplicate File Benchmark: 102,610 Files Scanned
See what a read-only ClearCopies scan found across 102,610 OneDrive files, how exact groups were classified, and what the result does not prove.
Large-drive duplicate claims are easy to inflate. Counting repeated names, equal sizes, or every item after the first candidate can produce a dramatic number without proving that the underlying bytes match. This benchmark records the narrower result produced by ClearCopies' exact-evidence rules on a real six-figure-file OneDrive library.
The scan was read-only. It enumerated provider metadata and used comparable OneDrive fingerprints together with exact byte size; it did not download original file bodies for hashing. The published figures are aggregate and anonymized: they do not expose account identity, filenames, paths, sharing details, or individual file records.
Decision snapshot
What the production scan measured
Each figure describes one stage of the same read-only metadata inventory.
| Observed value | Measurement | Interpretation |
|---|---|---|
| 102,610 | Files indexed | The inventory scale before duplicate grouping |
| 584 GB | Metadata inventory | Combined reported size of the indexed file library |
| 16,718 | Exact groups | Groups requiring equal byte size and comparable fingerprint evidence |
| 22.4 GB | Extra-copy estimate | Storage beyond at least one retained file in each exact group |
Reviewed August 28, 2026. This is one anonymized library, not a universal duplicate-rate forecast.
Anonymized production benchmark · August 2026
Validated on a six-figure-file OneDrive library
A completed read-only production scan indexed 102,610 files across 584 GB. Exact size-and-fingerprint matching identified 16,718 groups and 22.4 GB of recoverable extra copies. Similar or uncertain files were kept outside that exact total.
- 102,610
- files indexed
- 584 GB
- metadata inventory
- 22.4 GB
- exact recovery estimate
This benchmark exposed the difference between broad possible-match totals and exact recoverable storage, so the product now keeps those classifications visibly separate.
Written by ClearCopies Editorial · Technical review by ClearCopies Engineering
A practical checklist
- 1Inventory the complete scope and record whether enumeration finished successfully.
- 2Use byte size only to narrow candidates, never as exact proof by itself.
- 3Require a comparable provider fingerprint before creating an exact group.
- 4Calculate extra-copy storage only after retaining at least one item per group.
- 5Keep uncertain, unsupported, protected, or operationally required files outside the recovery estimate.
How the benchmark classified exact copies
The pipeline first grouped candidates by byte size, then compared a OneDrive-supplied content fingerprint such as quickXorHash or SHA-1 when the provider returned one. Algorithm identity travels with the value, so unlike fingerprint types are never compared as though they were interchangeable. A filename, timestamp, thumbnail, or equal size can help a reviewer investigate but cannot promote a candidate into the exact total.
Pagination, missing hashes, changing files, incomplete searches, and provider errors must remain visible. A clean-looking result is not useful if the inventory silently omitted part of the drive. The benchmark therefore reflects completed enumeration and the conservative exact grouping that the running product exposes to its review interface.
- Candidate: matching size or contextual clue
- Exact group: matching size plus comparable fingerprint
- Uncertain item: insufficient or incompatible evidence
What 22.4 GB of recoverable storage means
The estimate counts only copies beyond the retained baseline in each exact group. If three byte-identical items appear together, at least one remains; the estimate considers at most two extra copies. It is therefore different from the total size of every file participating in a duplicate group.
Exact content equality still does not decide which file item is redundant. A separate path may own the active shared link, useful version history, a retention obligation, or a canonical project location. ClearCopies keeps these items visible so the user can select the keeper rather than accepting a filename-based automatic choice.
What this benchmark does not claim
One account cannot establish a universal duplicate percentage for OneDrive users. Libraries differ by camera imports, collaboration, migrations, backups, creative workflows, file types, retention requirements, and age. The result should be treated as evidence that the method operates at six-figure-file scale, not as a promise that another account will recover the same share of storage.
It also does not measure visual near-duplicates, revised documents, recompressed media, exported Google Workspace files, or cross-provider copies. Those items normally have different bytes and require a different comparison method. Keeping them outside the exact total prevents a broad similarity claim from being mistaken for safe cleanup evidence.
How to reproduce a defensible measurement
Record the account scope, scan completion state, file count, total byte count, supported fingerprint algorithms, and unresolved items. Preserve aggregate output before cleanup so a later reviewer can understand how the estimate was produced. If the library changes substantially, scan again rather than applying an old decision to new files.
For a cleanup plan, add ownership, sharing, location, protected-format, and retention checks after exact detection. Revalidate every selected item immediately before a later move to recoverable trash. Detection proves content equality; review and revalidation make the operational decision safer.
Limits and risks to check
- — The benchmark is one library and must not be presented as a universal duplicate rate.
- — Provider fingerprints can be absent or unsupported for some items.
- — Exact content equality does not remove sharing, ownership, retention, or version-history obligations.
- — Storage dashboards and recycle-bin behavior can delay visible quota changes after cleanup.
Official references
Frequently asked questions
Did ClearCopies download 584 GB to run this scan?
No. The initial scan used OneDrive metadata, byte sizes, and provider-supplied fingerprints. Original file bodies were not downloaded for duplicate detection.
Does 22.4 GB mean every listed file can be deleted?
No. It is an aggregate extra-copy estimate after retaining a baseline item in each exact group. Every candidate still needs keeper, sharing, ownership, history, protection, and retention review.
Can this result predict savings for another account?
No. It demonstrates the method and production scale. Another library needs its own completed read-only scan because duplicate patterns and operational requirements vary widely.
Scan first. Decide with evidence.
ClearCopies reads OneDrive metadata and groups exact copies by byte size plus quickXorHash or SHA-1. Original file bodies are not downloaded for the scan. You review the result and export a plan before any separate write step.