Every tool is free to use. Enter your email once and all five open.All resources

76% of CRM Users Say Most of Their Data Is Wrong: Auditing CRM Data Quality With the 45-Metric Framework Behind a Real GTM Health Check

Three stacked translucent glass panes on a cream surface, warm gold light passing down through them, with small bright points marking flaws in each layer.

The audit came back green. A dedupe tool had scanned the CRM, flagged 1,400 duplicate contacts, merged them over a weekend and produced a report showing a duplicate rate under 2 percent. Two weeks later the board deck still disagreed with the finance model by six figures of ARR, a third of inbound leads were still landing on round robin, and the CS team was still discovering renewals from the customer's procurement email.

Nothing in that audit was wrong. It was just measuring the only thing that is easy to measure. Duplicates are a symptom sitting in one table. The causes sit in the fields nobody owns, the sync jobs that write records without checking for existing ones, the stage changes nobody validates, and the systems that never reach the CRM in the first place.

76%of CRM users say less than half of their CRM data is accurate and complete (Validity, 2025, n=602)
35%of sales professionals completely trust the accuracy of their organization's data (Salesforce State of Sales, 2024, n=5,500)
$12.9Ma year, the minimum average cost of poor data quality per organization (Gartner, 2020)

The numbers explain why a duplicate count is not enough. Validity's State of CRM Data Management in 2025 found that 37% of respondents say their firm loses revenue as a direct result of poor data quality, and the same share say staff fabricate data to appease decision-makers. Salesforce's 2024 State of Sales survey found reps spending 70% of their time on non-selling tasks, and only 35% completely trusting their organization's data. Fabricated close dates and placeholder values pass every duplicate check. So does a lead that never reached the CRM because a form integration failed silently.

This is a systems problem, not a people problem or a tool problem. Data quality in a CRM is the output of every process and integration that writes to it. An audit that inspects only the output can tell you the data is bad; it cannot tell you why, which fix comes first, or what the bad data is costing. That requires measuring the writers as well as the records.


Where it breaks

Most CRM data-quality audits fail in four predictable ways, and each one lives in specific objects.

Auditing the table, not the writers

A dedupe pass inspects the Contact and Account tables at one moment in time. It does not ask which process created the duplicates. Plauti's analysis of more than 12 billion Salesforce records (published January 2022) found 80% of records created through API integrations were duplicates, against 19% for imports. If the marketing automation sync, the enrichment job or the product sign-up webhook creates records without a lookup against existing ones, the duplicate rate is back where it started within a quarter. The objects to audit are the integration users, the record source field and the created-by distribution, not just the duplicate clusters.

Fill rate as a vanity metric

Field completeness reports count non-null values. They cannot tell a real Industry value from "Other", a real close date from one set to the last day of the quarter by habit, or a phone number from 555-0100. Validity's 2025 finding that 37% of respondents say staff fabricate data to appease decision-makers is the reason completeness alone misleads: required fields produce filled fields, not true ones. The check has to test validity and plausibility, and weight each field by whether anything downstream actually reads it.

No link between a defect and a decision

An audit that reports "12% of opportunities missing a next step" gives leadership nothing to act on. The same defect framed as "the forecast call relies on 140 opportunities whose next step is empty, worth a stated amount of pipeline" gets a decision. Every metric in the audit should name the report, workflow or system that consumes the field it measures. Without that, the remediation list is sorted by what is easiest to fix rather than by what is leaking revenue. A method for that pricing is in the true cost of duplicate records.

Ignoring what never reached the CRM

The most expensive gaps are records that do not exist: form submissions that errored in a sync, product-qualified accounts with no CRM match, billing customers whose ARR never joined an account. A CRM-only audit is blind to all of them by construction. Measuring coverage means reconciling the CRM against its sources, form logs against Lead creation, billing customers against Accounts, product workspaces against Account IDs.

The pattern behind all four: an audit that only reads the CRM measures symptoms. The causes live in the processes that write to it and the systems that never reach it, so a real health check measures all three layers and ties every defect to the decision it corrupts.

Reference architecture

The framework has three metric layers of 15 metrics each, plus a scoring layer that turns defects into priorities. It follows the logic of our Revenue Leak Report, which scores 45 metrics across four revenue domains (GTM Operations, Sales Operations, CS Operations and Revenue Intelligence), but groups the checks by what each one tests rather than by domain, which is the view a GTM engineer needs to build the audit. Exact thresholds and peer benchmarks stay in the report; the categories and the reasoning are here. Tools are named as examples, not endorsements.

Layer 1 · Data quality (metrics 1–15)

What it tests: whether the records themselves are right.

Metrics: (1) account duplicate rate (see why your CRM has three versions of every account), (2) contact and lead duplicate rate, (3) lead-to-account match rate (see deterministic vs probabilistic matching), (4) orphaned contacts with no account, (5) completeness of fields that downstream systems actually read, (6) format and picklist validity for country, state and industry, (7) email validity and bounce rate, (8) contacts not verified in the last twelve months, (9) job-change decay among active contacts, (10) firmographic freshness, (11) parent-account hierarchy coverage, (12) free email domains stored in company domain fields, (13) records owned by inactive users, (14) cross-system ID coverage linking CRM, billing and product IDs, (15) placeholder and fabricated values such as "test", "n/a" or dummy phone numbers.

Why it matters: decay is continuous. The US Bureau of Labor Statistics reported median employee tenure of 4.1 years in January 2026, so a contact database loses accuracy on a schedule whether anyone touches it or not.

Contract to the next layer: each metric is computed per object with a numerator, a denominator, the query used and the extraction timestamp.

Layer 2 · Process health (metrics 16–30)

What it tests: whether the processes that write records produce correct data.

Metrics: (16) lead routing accuracy, (17) unassigned lead rate, (18) lead response time at the median and 90th percentile, (19) stage entry criteria compliance, (20) stage skip rate, (21) open opportunities with no activity in the review window, (22) close-date push rate, (23) opportunities missing amount, close date or next step, (24) activity logging coverage, (25) closed-lost reason completeness and specificity, (26) sales-to-CS handoff completeness, (27) renewal date coverage on active customers, (28) manual override rate on automations, (29) human edits to system-owned fields, (30) entry lag between an event and its record.

Why it matters: these are the leading indicators. A rising push rate or override rate predicts next quarter's data quality better than today's duplicate count.

Contract to the next layer: each metric carries the field history or event log it was derived from, so every number can be traced to the records behind it.

Layer 3 · System coverage (metrics 31–45)

What it tests: whether the CRM is connected to everything that should feed it and everything that reads it.

Metrics: (31) form-to-CRM integrity, (32) sync error rate with marketing automation, (33) share of duplicates created by integrations, (34) source and campaign attribution coverage, (35) product usage joined to accounts, (36) billing and ARR joined to accounts, (37) support and CS data joined to accounts, (38) unused custom fields, (39) conflicting or overlapping automations, (40) documented field ownership showing which system writes each field, (41) reporting lineage from board metrics back to fields, (42) metric agreement across reports, (43) merge and delete permission hygiene, (44) field history coverage on key fields, (45) AI data readiness, meaning the fields an agent would act on are populated and trusted.

Why it matters: Salesforce's State of Data and Analytics report (n=7,652, November 2025) found data and analytics leaders estimating 26% of their organization's data is untrustworthy. Coverage metrics show where that untrusted share enters.

Contract to the next layer: each coverage gap is expressed as a count of affected records or accounts, not just a yes or no.

Layer 4 · Scoring & prioritization

Components: a consumer map linking every metric to the reports, workflows and systems that depend on it, a severity weight, and a dollar exposure estimate built from your own pipeline and ARR.

Example tools: a warehouse with dbt tests or a notebook for computation, a BI layer for the scorecard.

Contract to the output: a ranked list of defects, each with its metric, its consumer, its exposure and the system that would fix it.

Design principle: measure the writers, not just the records, and price every defect by the decision it corrupts. A metric that cannot name its downstream consumer does not belong in the audit.

The scoring logic fits on one screen. The weights below are illustrative, a suggested starting point rather than a benchmark.

# Illustrative scoring: rank defects by exposure, not by count
for metric in audit_metrics:
    defect_rate = failing_records / eligible_records
    consumers   = consumer_map[metric]            # reports, routing, forecast, renewals
    severity    = max(c.weight for c in consumers) # board or forecast = 3, routing = 2, hygiene = 1
    exposure    = affected_records * value_per_record(metric)  # pipeline $, ARR $, or hours
    priority    = severity * exposure
rank(audit_metrics, by=priority)   # ties broken by fix effort

Build sequence

Run it in this order, read-only from start to finish.

Extract read-only with a dedicated user

Create an API user with read access to CRM objects, field history, the setup audit trail and integration logs, and pull a full snapshot to a warehouse or local database. Nothing in the audit should write to production. The diagnose-before-you-build playbook covers the access model and how to run it without interrupting the team.

Build the consumer map before computing anything

List every report, dashboard, workflow, routing rule and integration that reads a field, and the field it reads. This is the step most audits skip, and it decides which of the 45 metrics carry weight in your stack.

Compute Layer 1 and Layer 2 from snapshots and field history

Data-quality metrics come from the snapshot; process metrics come from field history and activity logs over a trailing window long enough to include at least one full quarter-end, where push rates and entry lag behave differently.

Reconcile the CRM against its sources for Layer 3

Match form logs to Lead creation, billing customers to Accounts and product workspaces to Account IDs. Count what is missing on each side. Expect this to surface the largest single gap.

Validate against past cases

Before trusting any metric, check it on real history: take around twenty closed deals, routed leads or renewals and confirm the metric describes what actually happened. That is the same bar we hold any system to before it ships: 85 percent agreement with a senior operator's judgment on the client's own past cases.

Score, price and hand over one fix

Rank defects by exposure, then pick the single highest-priority fix and the system that owns it. An audit that ends in a forty-item remediation list tends to end in nothing.


Build vs. buy: trade-offs

There are three common ways to run the audit, and they differ mostly in how much of the framework they can see.

ApproachFitCost of ownershipFailure risk
Native CRM reports and data quality dashboardsA first pass on completeness and duplicates in one CRMLowest. Admin timeCovers most of Layer 1 and little of Layers 2 or 3. Field history retention limits process metrics, and nothing reconciles against billing or product data
Dedicated data quality or dedupe toolOngoing duplicate and validity monitoring at volumeModerate. License plus an ownerStrong on record-level defects, weak on process and coverage. Scores are rarely tied to downstream consumers, so priorities follow counts, not exposure
Warehouse audit with SQL or dbt tests and a consumer mapMulti-system stacks where CRM, billing and product must reconcileHighest up front. Models and the consumer map need an ownerCan compute all 45 metrics, but drifts if the consumer map is not updated when reports and workflows change

Whichever approach runs the computation, the consumer map and the prioritization are judgment work. Who owns them is an org design question; the GTM engineer vs. RevOps manager decision tree helps settle it. Validity's 2025 research, as reported by MediaPost, found 34% of respondents do not know who is responsible for CRM data quality, which is the most common reason an audit's findings go unowned.


Running it in production

Monitor

An audit is a snapshot; the useful version runs on a schedule. Keep the Layer 2 metrics, especially push rate, override rate and entry lag, on a weekly job with alert thresholds, because they move before the data-quality metrics do. Rerun the full 45 each quarter, and after any new integration goes live.

Fail safe

When an extraction is partial or field history has aged out, mark the affected metrics as not measured rather than scoring them from incomplete data. A metric reported as zero defects because the log was empty is worse than a blank. Version the queries so every score can be reproduced.

Explain it to leadership

Leaders do not need 45 numbers. Show three layer scores, the top three defects with their dollar exposure, and the one fix you are starting with. Then report the same three scores next quarter, so the audit becomes a trend rather than a one-off verdict.


Where this fits in the system

Every metric in the framework maps to a system that would fix it. Lead-to-account match rate, routing accuracy and response time are what Speed-to-Lead depends on. Stage compliance, stale opportunities and push rate are the Pipeline Hygiene Sentinel's territory, and they feed the Forecast Assistant. Handoff completeness belongs to the Handoff Orchestrator, renewal date coverage to Renewal Radar, and reporting lineage and metric agreement to the Board Report Engine. The full map is on the systems page.

That is why the audit comes before the build. Forward-deployed engineering starts from the defect with the largest exposure, ships the one system that fixes it, and measures the same metric again afterward.

Sources: Validity, The State of CRM Data Management in 2025 (n=602, July 2025), including MediaPost's coverage of the report (July 2025). Salesforce, State of Sales (n=5,500, fielded March to April 2024, published July 2024). Gartner, data quality research (2020). Plauti, "80% of all new integration data in CRMs is duplicate" (analysis of more than 12 billion Salesforce records, January 2022; archived copy, the original page has been removed). Salesforce, State of Data and Analytics (n=7,652, November 2025). US Bureau of Labor Statistics, Employee Tenure news release (September 2026).

Read next