Every revenue team has sat through this meeting. RevOps asks for a quarter to clean the CRM. The CFO asks what it is worth. Someone quotes a vendor study saying bad data costs companies millions a year, the CFO points out that the company is not the Fortune 500 company in that study, and the project goes back to the bottom of the list. Six months later the same conversation starts again.
The request failed for an architectural reason, not a persuasive one. Nobody could trace the cost from a specific field on a specific record to a specific line in the P&L. "Duplicates" was a count of rows. The cost of those rows lived in rep calendars, routing logs, campaign reports and renewal forecasts, and no system joined them. Until you build that join, data quality stays an opinion.
Those numbers establish that the problem is real. They do not tell you what it costs you. Gartner's widely cited 2020 estimate, that poor data quality costs organizations at least $12.9 million a year on average, describes large enterprises across every data domain. Thomas Redman's 2017 estimate in MIT Sloan Management Review put the cost of bad data at 15% to 25% of revenue for most companies, a range so wide it invites the CFO to pick the bottom and discount it further. Validity's 2025 survey of 602 CRM users found 76% saying less than half of their CRM data is accurate and complete. That confirms the pattern, not your figure.
This is why pricing data decay is a systems problem. The inputs already exist in your stack: activity logs, routing timestamps, campaign membership, account hierarchy, renewal dates. What is missing is a model that joins a record-level defect to a revenue event, and a pipeline that recomputes it every week, so the number moves when the data moves. Treat the cost as a derived metric with lineage, not a slide statistic.
Where it breaks
Duplicate and decayed records do not cost money by existing. They cost money when a downstream process reads the wrong record. Four failure modes account for most of the dollars, and each one has a specific object where the cost can be measured.
Rep time lost to record ambiguity
A rep finds two contacts with the same name, one with an email that bounced last month, and an activity history split across two accounts. They reconstruct context before the call, then log it on whichever record they found first, splitting the history again. The cost appears in the Task and Event objects as time with no selling outcome, and in email bounce fields as sequences sent to people who left. Salesforce's 2024 State of Sales survey found reps spending 70% of their time on non-selling tasks; record ambiguity is one of the few pieces of that time you can measure from system logs rather than survey memory.
Misrouted inbound leads
A demo request arrives from someone at an existing target account. Lead-to-account matching fails because the account's domain field is blank on the record the owner works, or because the lead matches a duplicate with a different owner. Round robin assigns it to a rep who has never spoken to the account, and the SLA clock runs while ownership is sorted out. The cost lives in routing logs: the assignment rule that fired, the owner change history, and the time between created date and first activity.
Attribution that splits one buyer into three
Plauti's analysis of more than 12 billion Salesforce records processed in 2021, published in January 2022, found more than 45% of new records entering CRMs were duplicates. When a buyer exists as three contacts, their webinar attendance sits on one, their content downloads on another, and the opportunity contact role on a third. Campaign influence reports read Campaign Member records attached to contacts with a role on the opportunity, so two-thirds of that buyer's journey drops out of the model. Programs that worked look like they did not. Where those extra contacts come from, and how to stop them being created, is the subject of why your CRM has three versions of every account.
Segmentation drift that hides churn risk
Customer segments are built on fields: employee count, industry, plan tier, owner. When an account is split into a parent and an orphaned duplicate, product usage lands on one record and the CSM assignment on the other. The health score reads the record without usage and flags a healthy customer as at risk, or reads the one without the owner and flags nothing. People move too: the US Bureau of Labor Statistics counted 38.0 million quits and 62.8 million total separations in 2025, so champion contacts go stale continuously, not once a year. A renewal forecast built on stale segments is a forecast of the wrong customers.
Reference architecture
A cost model that is recomputed from live data needs five layers. None of them require new platforms, only a clear contract between the ones you have. Tools named are examples, not endorsements.
Components: CRM objects (Account, Contact, Lead, Opportunity, Task, Campaign Member), routing logs, email bounce and engagement events, product usage, billing and renewal dates.
Example tools: Salesforce or HubSpot, a routing tool, a sales engagement platform, a product analytics pipeline, a billing system.
Contract to the next layer: Raw records land with their native IDs and timestamps, unmodified. The cost model must be able to replay any week from source data.
Components: Duplicate clustering on normalized domain and email, staleness flags (hard bounce, no engagement in a set window, title or company change), and completeness checks on the fields each downstream process reads.
Example tools: Native duplicate reports, a dedupe or matching tool, or SQL or dbt models in a warehouse. The matching rules behind the duplicate flag are covered in identity resolution for B2B RevOps.
Contract to the next layer: Every record carries a defect flag set (duplicate cluster ID, stale, incomplete, orphaned) with the rule that produced each flag.
Components: The cost model: four joins that connect a defect flag to a consumption event (a rep activity, a routing decision, an attribution touch, a renewal), each with a written pricing assumption.
Example tools: Warehouse models, a BI semantic layer, or a scheduled job in a workflow tool such as n8n.
Contract to the next layer: Each cost line outputs dollars, the record IDs behind them, and the assumption version used, so any number can be traced back to rows.
Components: A decay-cost table with one row per cost line per week, plus an assumptions table that finance has signed off.
Example tools: A warehouse table, synced to a CRM custom object if leaders live in the CRM.
Contract to the next layer: Historical rows are never overwritten. When an assumption changes, a new version is written and prior weeks are restated side by side.
Components: A weekly cost trend for leadership, fix queues ranked by dollars rather than row count, and alerts when a cost line jumps.
Example tools: BI dashboards, CRM list views, a monitoring agent posting to Slack or Teams.
Contract to the next layer: Every fix queue item carries its dollar weight, so the team works the most expensive records first.
The model itself is four lines. The first two and the fourth are losses. The third is reported separately as exposure, because misattributed spend is a decision risk rather than money that has already left the building, and mixing the two is the fastest way to lose a CFO's trust.
# Four-line data decay cost model (weekly or annualized)
rep_time_cost = reps x hours_lost_per_rep_week x selling_weeks x loaded_hourly_cost
misrouting_cost = inbound_volume x misroute_rate
x (conv_clean - conv_misrouted) x win_rate x avg_acv
segmentation_cost = arr_on_misclassified_accounts x churn_uplift
decay_loss = rep_time_cost + misrouting_cost + segmentation_cost
attribution_exposure = program_spend x share_of_touches_unlinked # report separately
An illustrative example shows how the lines add up. Every number below is a made-up round figure chosen to show the arithmetic, not a benchmark or a client result. Take a B2B SaaS company at $15M ARR with 20 quota-carrying reps. If activity logs show each rep losing two hours a week to duplicate and stale records across 46 selling weeks at a $75 loaded hourly cost, rep time costs $138,000. If 1,600 inbound requests a year see 12% misrouted, and misrouted leads convert to opportunities at 15% instead of 25%, with a 20% win rate and $30,000 average contract value, misrouting costs $115,200. If $1.2M of ARR sits on accounts split or misclassified by duplicates, and those accounts churn five points more often, segmentation costs $60,000. That puts the illustrative decay loss at $313,200, about 2.1% of ARR, with a separate $90,000 attribution exposure if 15% of touches on a $600,000 program budget cannot be tied to an opportunity.
Build sequence
Build the number before the fix. Four weeks of baseline give you a ranked fix list and the before-and-after that funds the next project.
Define the defect flags
Write the rules for duplicate, stale, incomplete and orphaned in plain language and in SQL. Keep them narrow: a duplicate is two accounts sharing a normalized corporate domain, not two with similar names. The diagnose-before-you-build playbook covers how to run this read-only against production data.
Instrument the four consumption events
For each cost line, find the event that proves a defect was read: an activity logged on a record in a duplicate cluster, a routing decision on a lead whose domain matched an account with a different owner, a campaign touch on a contact with no opportunity role, a renewal on an account with a split usage record. If an event is not logged, log it now. The model is only as good as these joins.
Agree on the pricing assumptions with finance
Loaded hourly cost, conversion rates, win rate, ACV and churn uplift should come from your own historical data or from finance, not from a vendor study. Put them in a versioned assumptions table. A number finance helped build is a number finance defends in the budget meeting.
Backtest on cases your team already judged
Pull around twenty past examples per cost line, such as misrouted leads the team investigated or renewals lost after a split record, and check that the model flags them and prices them plausibly. Our bar for anything we ship is 85 percent agreement with what a senior operator would have concluded, or it does not go live.
Publish the weekly cost table
Schedule the model, write one row per line per week and show the trend. Report the three loss lines as the headline and attribution exposure beneath it, never summed together.
Rank the fix queue by dollars
Sort duplicate clusters and stale records by the cost they generated in the last 90 days. Expect a small share of records to drive most of the dollars, which turns an open-ended cleanup into a scoped one.
Build vs. buy: trade-offs
There are three realistic ways to produce this number. Most teams start with the first to win the argument and move to the second or third once the number has to stay current.
| Approach | Fit | Cost of ownership | Failure risk |
|---|---|---|---|
| One-off spreadsheet model from CRM exports and duplicate reports | A first budget case, one CRM, a leadership team that needs a number this quarter | Lowest. A few days of analyst time | Goes stale the week it is presented. Assumptions live in cells nobody audits, and the model cannot show whether the fix worked |
| Native CRM reports plus a data quality or dedupe tool's dashboards | Duplicates are the main defect, and leaders already work inside the CRM | Moderate. Tool license plus admin time | Counts defects well but rarely joins them to routing, attribution or renewal events, so it reports rows rather than dollars |
| Warehouse cost model with versioned assumptions and a monitoring agent | Several sources, product and billing data in play, a CFO who wants lineage | Highest. Needs a data or GTM engineer to own models and tests | Silent pipeline failures can freeze the number. Needs freshness checks and an owner for the assumptions table |
The deciding factor is not tooling but whether the number is recomputed from live data with assumptions finance signed. Who owns that model is an org design question as much as a technical one, and the GTM engineer vs. RevOps manager decision tree is a useful way to settle it.
Running it in production
Watch each cost line weekly, not just the total. A jump in misrouting cost with a flat duplicate count usually means a routing rule changed, not that data got worse. Track input freshness too: a routing log that stops updating looks like falling cost.
When a source is stale or an assumption is missing, the model should output "not computed" for that line rather than zero. A zero looks like success and gets screenshotted into a board deck. Keep every prior assumption version so a restated number can always be explained.
Lead with the loss number and the three lines behind it, then the exposure line on its own. Say which assumptions finance supplied, show the trend since the fix started, and name the records that drove the largest share. Leaders need to see that the number moves when the data improves, which makes it a metric rather than a claim.
Where this fits in the system
A decay cost model is not a system on its own. It is the pricing layer that tells you which system to build first and proves whether it worked. The Pipeline Hygiene Sentinel is its natural operator: it flags duplicate, stale and orphaned records as they appear, so the cost line falls instead of the backlog growing between cleanups. The misrouting line is the one Speed-to-Lead depends on, since routing in minutes is only useful if the lead attaches to the right account. The segmentation line feeds the Churn Signal Watchtower, which cannot watch a customer whose usage and ownership sit on different records. And the Board Report Engine is where the weekly cost table reaches leadership, next to pipeline and retention rather than in a separate data quality report nobody opens. The full map is on the systems page.
That sequence is the logic of forward-deployed engineering: price the leak on your own data, build the system that closes the most expensive line, and leave the model running so the next budget request starts with a number rather than an argument.
Sources: Gartner, data quality research (2020), as cited on Gartner's data quality topic page. Validity, The State of CRM Data Management in 2025 (n=602, July 2025). Salesforce, State of Sales (n=5,500, July 2024). Thomas C. Redman, "Seizing Opportunity in Data Quality," MIT Sloan Management Review (2017). Plauti, "80% of all new integration data in CRMs is duplicate" (analysis of more than 12 billion Salesforce records in 2021, January 2022; archived copy, the original page has been removed). US Bureau of Labor Statistics, Job Openings and Labor Turnover Survey, 2025 annual figures (March 2026).




