Every tool is free to use. Enter your email once and all five open.All resources

Only 17% of Teams Contract-Test Their APIs: Enrichment Waterfalls and the Provider-Agnostic Data Layer That Doesn't Break When Vendors Change

A beam of warm gold light passing through a row of interchangeable translucent glass prisms into one clear glass vessel on a pale white surface

The renewal conversation with your enrichment vendor goes badly, so procurement signs with a cheaper provider that claims better coverage in your segment. The switch is scheduled for a Monday. By Wednesday the enterprise queue is suspiciously quiet. The routing flow reads a field called Employees__c, which the old provider's connector populated from its own employee_count key. The new connector writes to NumberOfEmployees and reports headcount as a range string, so every rule that asked "Employees__c >= 200" now evaluates against a null and falls through to the SMB round-robin. Lead scoring, two sequencer audiences, the territory job and a board dashboard read the same dead field. Nobody gets an error.

The provider, the connector and the routing flow all worked. What failed was an architecture in which forty downstream objects knew the name of one vendor's field. Enrichment was built as a pipe from a provider into the CRM, when it should have been a layer with its own schema that providers plug into.

17%of API practitioners run contract testing, the check that catches silent schema changes (Postman, 2025)
76%of CRM users say less than half of their CRM data is accurate and complete (Validity, 2025)
4.1 yearsmedian time US workers have been with their current employer (BLS, 2026)

The pressure on enrichment is structural. The US Bureau of Labor Statistics reported in September 2026 that the median wage and salary worker had been with their current employer 4.1 years as of January 2026. Every contact record is a snapshot of someone who will eventually change title or company, and no single provider sees every move at the same speed. Validity's State of CRM Data Management in 2025 (602 CRM users and stakeholders) found 76% say less than half of their CRM data is accurate and complete, and 37% say they have lost revenue as a direct result of poor data quality. Gartner research from 2020 estimates that poor data quality costs at least $12.9 million a year per organization on average.

The integration side is no better protected. Postman's 2025 State of the API Report (more than 5,700 developers, architects and executives, October 2025) found 55% of respondents struggle with inconsistent or outdated API documentation, and only 17% practice contract testing, the discipline of checking that an API still returns the shape its consumers depend on. Enrichment providers are API vendors. When one renames a key or starts returning empty fields with a success code, most stacks find out from a rep, not a test.

Automation raises the stakes. Salesforce's State of Data and Analytics report (7,652 respondents, November 2025) found 89% of data and analytics leaders with AI in production had seen inaccurate or misleading AI outputs. An agent writing outbound copy from a stale title makes the mistake faster and at greater volume than a person would.

This is a systems problem, not a vendor problem. A better provider only resets the clock. The durable fix is to own the schema, treat every provider as a replaceable adapter and keep ordering, stopping and fallback logic in a layer you control.


Where it breaks

Enrichment failures show up weeks later as routing drift, falling reply rates or a segment report that no longer adds up. Five patterns account for most of them.

Provider field names are wired directly into logic

The connector writes to fields named after the provider's payload, and everything downstream reads them: record-triggered flows, validation rules, scoring formulas, sequencer filters, territory jobs, warehouse models and dashboards. Nothing in the CRM records which fields came from which vendor. The test: ask how many objects would need to change if you replaced your primary provider tomorrow. If the answer is "we would have to look", the provider owns your schema.

The waterfall has no stop rule and no provenance

A waterfall calls providers in sequence until a field is filled. Built naively, it calls every provider for every field, overwrites values a rep verified on a call and leaves no trace of which provider supplied the answer. Credits burn on complete records, and when a value is wrong nobody can say who put it there. Without a per-field source and timestamp, you cannot measure which provider earns its place in the sequence.

Taxonomies do not line up

Providers disagree on more than values. One returns industry as a proprietary label, another as a NAICS code, a third as a LinkedIn-style category. Employee counts arrive as integers, bands or range strings. When two providers feed the same field without a mapping step, the CRM ends up with "Computer Software", "Software Development" and "511210" describing one segment, and every report grouped by industry quietly splits your market.

Silent nulls and schema drift

The most expensive failure returns no error. A provider deprecates a field and starts sending null, renames a key in a minor version, changes a date format, or returns HTTP 200 with an empty body when it has no match. The workflow tool logs a successful run, and downstream logic treats null as a value (usually "not enterprise" or "not in territory") and routes accordingly.

The order was set once and never measured

Most waterfall orders were chosen during a vendor evaluation, on a sample that did not look like today's pipeline. Coverage varies by region, company size and persona, so an order that works for North American mid-market can be wrong for European enterprise. If the order is never re-tested against verified outcomes, you pay for the first provider's gaps and the second provider's credits.

The common thread: in each case the rest of the stack depends on a provider's format rather than on a contract you own. The waterfall is not the architecture. The normalized schema it writes into is.

Reference architecture

The design separates four concerns most stacks blend together: calling a provider, translating its answer, deciding which answer wins and writing the winner. Tools are examples, not endorsements, and every threshold is a suggested starting point, not a benchmark.

Sources · Providers and research

Components: firmographic, contact, technographic and phone providers called by API; the enrichment features built into your CRM or sales engagement tool; web research agents; and a human research queue.

Contract to adapters: each provider is called with the canonical identifiers you already hold (normalized domain, account_key, person_key, LinkedIn URL where you have it), never with a raw form submission. Identity resolution runs first; enrichment attaches to a known record.

Layer 1 · Provider adapters

Components: one small adapter per provider that handles authentication, rate limits, retries and pagination, and stores the raw response untouched with the provider name, API version and request time. Built in a workflow tool such as n8n or Workato, a platform such as Clay, or middleware code.

Contract to normalization: the raw payload plus a response status the rest of the stack can trust: matched, no_match, error or schema_changed. An empty body with a success code is no_match, never data, and a missing mapped key raises schema_changed.

Layer 2 · Normalization and canonical schema

Components: a canonical schema you own (employee_band, industry_code, hq_country, seniority, verified_email and so on), with one mapping file per provider that translates its fields and taxonomies into yours. Often lives in the warehouse as dbt models or in a mapping table the workflow tool reads.

Contract to the resolver: candidate values expressed only in canonical terms, each carrying provider, observed_at and a confidence the mapping assigns. Nothing past this layer sees a vendor field name, so adding a provider means one adapter and one mapping file.

Layer 3 · Waterfall resolver

Components: per-field configuration that sets provider order (optionally by segment, such as region or company size), the stop rule (stop at the first candidate above a confidence threshold, or require two agreeing providers for high-stakes fields such as verified email), the maximum age before re-enrichment, and the fallback when every provider misses: a manual research task with the record, the fields needed and the reason.

Contract to the system of record: one winning value per field with its source, confidence and resolved_at, plus the losing candidates for audit. Suggested starting point: routing-critical fields resolved for inbound within five minutes.

System of record · Write-back and activation

Components: canonical CRM fields on Account and Contact, per-field source and last_verified fields (or a related enrichment history object), an overwrite policy, and the flows, sequencers, scoring models, dashboards and agents that read them.

Contract: one integration user writes enrichment, and only into canonical fields. A value marked rep_verified or customer_provided is never overwritten by a provider, only flagged when a provider disagrees. Downstream logic reads canonical fields exclusively, so a provider swap is invisible to it.

The contract fits on one screen, and it is worth reviewing whenever anyone proposes a new vendor:

canonical field: employee_band      # values: 1-10 | 11-50 | 51-200 | 201-1000 | 1001-5000 | 5000+
  order:
    NA, EMEA  : [provider_a, provider_b, provider_c]
    APAC      : [provider_b, provider_a]
  stop_rule   : first candidate with confidence >= 0.8
  max_age     : 180 days
  never_overwrite_if: source in (rep_verified, customer_provided)
  on_all_miss : create research_task(record, field, reason="no provider match")

write: value, source, confidence, resolved_at, candidates[]
Design principle: own the schema, rent the data. Providers are adapters behind a canonical field list that belongs to you. Order, stop rules and fallbacks live in configuration you control, and nothing downstream is allowed to know which vendor supplied a value, only where to look it up.

Build sequence

Six steps, in order, each ending in a test.

Inventory every enrichment dependency

List every field a provider writes, then every flow, formula, report, sequencer filter, warehouse model and agent prompt that reads it. The diagnose-before-you-build playbook covers how to do this without touching production. Test: for your primary provider, you can state exactly which objects would break if it disappeared tomorrow.

Define the canonical schema and taxonomies

Decide your own field list, size bands, industry taxonomy and seniority levels as the only vocabulary downstream logic may use, with source and last_verified fields beside them. Test: every routing, scoring and segmentation rule can be rewritten against canonical fields without losing meaning.

Build one adapter and one mapping per provider

Wrap each provider in an adapter, write its mapping file, and assemble a golden test set of a few hundred hand-verified records across your real segments. Test: each provider's fill rate and accuracy per field and per segment are measured against the golden set, not taken from a vendor dashboard.

Configure the waterfall per field

Use the golden-set results to set provider order per field and segment, stop rules, maximum age and the research fallback. Test: on the golden set, the waterfall beats the best single provider on accuracy for routing-critical fields, and every miss produces a research task instead of a null.

Cut downstream logic over to canonical fields

Repoint every dependency at the canonical fields, retire the vendor-named fields and route all enrichment writes through one integration user. Test: a search of flows, formulas and reports finds no reference to a vendor-named field.

Replay real cases, then run a swap drill

Run around twenty recent inbound and outbound records through the new layer in a sandbox and compare routing, scoring and segmentation decisions with what an experienced operator says should have happened. We hold every system to the same bar: 85 percent agreement on the client's own past cases, or it does not ship. Then disable the primary provider in the sandbox and rerun. Test: agreement clears the bar, and the swap drill changes only the configuration file.


Build vs. buy: trade-offs

The choice matters less than whether the canonical schema and overwrite policy exist wherever the logic runs.

ApproachFitCost of ownershipFailure risk
Native CRM enrichment (built-in data services, one marketplace connector, duplicate and validation rules)One primary provider, moderate volume, a lean RevOps team, segments the provider covers wellLowest. Bundled or single-vendor licensing and admin skills you already haveVendor-shaped fields; no real fallback; a switch means remapping every dependency by hand
Waterfall platform or workflow tool (for example Clay for provider sequencing, n8n or Workato for orchestration)Several providers, a technical RevOps owner, segments with uneven coverage, a need to test new vendors quicklyModerate. Platform licensing, provider credits and an owner for every table or workflowEasy to write vendor-shaped values straight to the CRM; normalization and stop rules must be designed on purpose
Custom adapter layer (adapters in middleware or serverless functions, normalization and resolver in the warehouse, reverse ETL to the CRM)High volume, product-led motions, an existing data team, agents consuming enrichmentHighest. Engineering time, monitoring, on-call and provider contract managementMost control and provenance; the risk is a small team maintaining plumbing instead of improving decisions

A reasonable default for most teams between Series A and Series C: keep the canonical schema and the mapping files in a place you own (warehouse or a governed mapping table), let a platform handle provider calls and sequencing through an orchestration layer, and keep the CRM write narrow and policy-enforced. Who should own the custom pieces is an org design question the GTM engineer vs. RevOps manager decision tree works through.


Running it in production

Monitor

Track, per canonical field and provider: fill rate, accuracy against a refreshed golden set, which waterfall position supplied the winner, schema_changed and no_match rates, median field age and cost per filled field. A rising schema_changed count tells you before your reps do.

Fail safe

Put a circuit breaker on each adapter: when error or schema_changed responses cross a threshold, skip to the next provider. Keep the last known good value with its original timestamp rather than a null. Never overwrite a verified value. When every provider misses a routing-critical field, send the record to research with a deadline, never into a default queue.

Explain it to leadership

Leadership needs three statements, not the diagram: switching data vendors is now a configuration change measured in days, not a rebuild measured in quarters; we know the cost per usable record for each provider and can negotiate on it; and every value that drives routing can be traced to its source and date.


Where this fits in the system

Enrichment is the second layer of the GTM data stack: it sits after identity resolution has decided which account and person a record belongs to, and before signal and orchestration decide what to do about it. Speed-to-Lead needs size, region and segment resolved within minutes of a form fill, or a hand-raiser lands in the wrong queue. The Signal-Based Outbound Engine is only as good as the title, seniority and verified contact data it attaches to a buying signal. The Board Report Engine segments pipeline and retention by company size and industry, so a taxonomy split in enrichment becomes a wrong number in a board deck. The full map is on the systems page.

Build this layer before the next vendor decision, not after it. A forward-deployed engineering approach measures your current providers against your own verified records, builds the schema and waterfall around what that shows and proves the result on your past cases before moving to the systems that depend on it.

Sources: US Bureau of Labor Statistics, Employee Tenure in January 2026 (released September 24, 2026). Validity, The State of CRM Data Management in 2025 (602 CRM users and stakeholders, 2025). Gartner, data quality research (2020). Postman, 2025 State of the API Report (more than 5,700 developers, architects and executives, October 2025). Salesforce, State of Data and Analytics (7,652 respondents, surveyed June to August 2025, published November 2025).

Read next