What is ROT data, and how to clean it up
ROT stands for redundant, obsolete and trivial. It is the data nobody uses but everybody pays for, in storage cost, compliance risk and carbon. Here is what ROT data is, why it matters, and a practical way to find and remove it.
What ROT data actually means
ROT data is an acronym for three overlapping kinds of low-value information that accumulate in every organization:
- Redundant: duplicates and near-duplicates, the same file saved in five places, backups of backups, re-shared attachments, "final_v2 (copy)".
- Obsolete: data that was useful once but no longer is, superseded documents, finished-project workspaces, files from people who left years ago.
- Trivial: content with no business value at all, personal photos, system junk, cache files, throwaway drafts.
Individually each file looks harmless. In aggregate, ROT can make up a large share of everything you store, and unlike valuable data, it only ever costs you.
Where ROT data comes from
ROT is not a sign of a badly run organization: it is the natural by-product of everyday work. People save copies "to be safe". Systems keep every version and every draft by default. Email attachments get downloaded, re-saved and re-shared. Migrations lift entire folder trees into a new platform without ever retiring the old one. Employees leave, and their files stay behind with no owner and no review date.
Because no single action creates a visible problem, ROT accumulates silently. There is rarely a moment where someone decides "we will now store half a petabyte of data nobody uses": it just grows, one reasonable-looking decision at a time, until the storage bill and the risk register force the question.
How big is the problem? The dark-data number
ROT overlaps heavily with what analysts call dark data: information that is stored but never used or even classified. The widely cited Veritas Databerg study found that around 52% of the data organizations store is dark: more than half of it invisible, unmanaged and, for the most part, ROT. (Source: Veritas Databerg report.)
If half of what you keep serves no purpose, then a large part of your storage bill, backup windows and data-protection effort is being spent on nothing.
Why ROT data costs you three ways
1. Money
Every gigabyte you keep is a gigabyte you pay to store, replicate, back up and (in the cloud) egress. ROT inflates all of it. Worse, it hides the growth of your real data, so capacity planning and licensing decisions get made on padded numbers. Removing ROT is one of the few IT actions that cuts cost with essentially no downside.
2. Risk
ROT is a compliance liability. Old files full of personal data you no longer have a reason to hold sit in direct tension with the GDPR principle of data minimization: you should not keep personal data longer than you need it. Forgotten, unclassified files are also exactly what attackers exfiltrate and what turns a breach into a reportable incident. The less stale personal data you hold, the smaller your attack surface and your regulatory exposure.
3. Carbon
Stored data is not weightless. It lives on disks in datacenters that draw power around the clock for compute, cooling and redundancy. Datacenters account for roughly 2.5% of global emissions (Nature Climate), and every terabyte of ROT you keep contributes its share. Deleting data you never use is one of the simplest ways to shrink your digital carbon footprint, and, increasingly, something you may need to measure and report under frameworks like the CSRD. Data you were going to delete anyway becoming a reportable sustainability win is about as close to a free lunch as IT gets.
How to find and remediate ROT data
Cleaning up ROT is a repeatable process, not a one-off purge:
- Discover. Scan across every repository (SharePoint, OneDrive, cloud object storage, NAS and local drives) and inventory what you actually hold. You cannot clean up what you cannot see.
- Classify. Separate the redundant, obsolete and trivial from the data that genuinely matters, so you know what is safe to act on.
- Prioritize. Rank findings by the cost, risk and carbon they carry, so you tackle the biggest wins first instead of nibbling at the edges.
- Decide and act. Apply clear rules (keep, archive or delete) and document the reasoning so the cleanup is auditable.
- Repeat. ROT regrows. Rescan on a schedule so it never builds back up to where you started.
Automate the whole thing with TerraBytes
Doing this by hand across a real estate is slow and error-prone. TerraBytes runs the discover-classify-prioritize loop in a single serverless scan. Your data never leaves your environment and stays encrypted throughout. It flags duplicate, outdated, oversized and risky files across SharePoint, OneDrive, NetApp, NAS and local drives, and puts a euro and CO2 figure on each finding so you fix what matters most. Nothing is deleted automatically: you define the keep, archive and delete rules and stay in control. To tackle the risk side, see GDPR data cleanup; to clean up cost and carbon across your whole estate, start with the storage cleaner and optimizer.
Keep reading
How to find duplicate files →
Tackle the "redundant" in ROT: spot duplicates across cloud and on-prem.
GDPR data cleanup →
Find and remediate risky, personal and unencrypted files before they become a liability.
The hidden cost of storage →
Why unused data quietly drives up both cost and carbon emissions.