← all posts

July 13, 2026 · 5 min read

The hidden tax of manual data cleanup

The visible cost is the spreadsheet fix. The hidden cost is every person who learns to verify the fix forever.

Researchers spent six months embedded inside an industrial company with 18 subsidiaries, roughly 8,500 SKUs, and more than 10,000 customers. What they found wasn't a data-quality incident. It was two full-time business-intelligence managers whose entire job was cleaning data and pricing activities, month after month, indefinitely. The company's own COO put a number on the waste around them: duplicated information work and parallel registration equivalent to two more full-time marketing salaries, on top of the two people already doing cleanup full-time.

That's what "a little manual cleanup" looks like once it becomes part of the operating model instead of an occasional afternoon. The visible cost is easy to point at: somebody corrects rows, reconciles two exports, rebuilds a report. The cost that doesn't show up on anyone's task list is what happens after: an analyst checks the number against a second source out of habit. A manager adds a reviewer nobody asked for. A team keeps its own shadow spreadsheet because the official one surprised them once, two years ago, and they never forgot it.

Cleanup fixes a record; controls fix the pattern

A one-time correction doesn't tell the next person whether the same defect will come back, which is exactly why trust doesn't recover just because the spreadsheet got fixed. Fixing a number is an act of diligence. People don't extend trust to an act of diligence, no matter how careful it was. They extend trust to a process they've watched hold up more than once.

The process that actually retires the tax is visible and repeatable: the same checks run every time a dataset lands, not just when someone remembers to run them. Issues carry their provenance instead of arriving as an unexplained correction. Fixes are reversible. Exceptions stay readable months later, not buried in a Slack thread. None of that promises zero imperfections, because zero isn't achievable and chasing it is its own kind of waste. What it promises is knowing which imperfections actually matter, so the same surprise doesn't get rediscovered by three different people in three different quarters.

The cost of a bad dataset is not just the cleanup. It is the permanent shadow review that follows it.

Clean your data.
Trust your forecasts.