The scene is a composite, and the details are illustrative. An outbound agent goes live on a Monday with access to the CRM, the sequencing tool and an enrichment API. By Wednesday it has enrolled 2,400 contacts. On Thursday someone notices that 300 of them belong to accounts with open opportunities, owned by AEs who never asked for help. Forty are at a customer in the middle of a renewal. Nobody can say which signal drove each enrollment, because the agent logged a summary rather than its inputs. Unwinding it takes two admins most of Friday, and the CRO then puts every agent action behind a manual review queue. A month later the queue is 900 items deep, reps ignore it and the agent is effectively off.
Both are governance failures. The first had no controls; the second had one, and it was a human.
The incident data is not reassuring. SailPoint's AI agent research, conducted by Dimensional Research with 353 IT professionals who hold enterprise security responsibilities and published in May 2025, found that 82% of organizations already use AI agents, but only 44% have policies securing them. Eighty percent said their AI agents had taken unintended actions, including accessing unauthorized systems or resources (39%). Deloitte's 2026 State of AI in the Enterprise (3,235 IT and business leaders across 24 countries) found that only 21% of organizations have a mature governance model for agentic AI, meaning clear decision boundaries, real-time monitoring and audit trails.
Sales is moving faster than its controls. Salesforce's State of Sales report (4,050 sales professionals, published February 2026) found that 54% of sellers have already used AI agents and nearly nine in ten plan to by 2027. Gartner's June 2025 prediction that over 40% of agentic AI projects will be canceled by the end of 2027 names inadequate risk controls among the three causes, alongside cost and unclear value. Gartner also predicts that "guardian agents", systems that oversee other agents, will account for at least 10 to 15% of the agentic AI market by 2030 (June 2025).
This is a systems problem, not a policy-document problem and not a people problem. A policy in a slide cannot stop an API call, and a reviewer who sees every action becomes the system's throughput ceiling. The only governance that scales with agent volume is governance the stack enforces on its own: at the permission layer, at the write path and in the log. OWASP's 2025 Top 10 for LLM Applications names the failure pattern directly as "Excessive Agency" and traces it to three root causes: excessive functionality, excessive permissions and excessive autonomy.
Where it breaks
Four anti-patterns recur in stacks that run agents without a governance layer.
The integration user with admin rights
The agent authenticates as a shared integration user, usually the one created years ago for the marketing automation sync, with a profile that can edit every object and field. The agent only needs to read Account, Contact and Intent_Signal__c and create Task records, but it can also update Opportunity.Amount, reassign OwnerId and delete records. The blast radius is the whole CRM, and every change shows up in field history as the integration user, indistinguishable from the sync. The same ownership problem, with people instead of agents, is why territory assignment needs its own audit trail.
Logs that record outcomes, not decisions
Most agent logs capture what happened: "enrolled contact in sequence." Few capture why: which trigger fired, which fields and signals were read, which policy version applied, what the model returned and with what confidence. Without the inputs, a bad action cannot be replayed or traced to the data, the prompt or the rule. A log entry without a run_id joining it to the inputs is a diary, not an audit trail. Agent reliability and the decision log covers what to instrument.
Approval gates on everything, or on nothing
Teams pick one of two settings. Either every action waits for a human, which recreates the bottleneck and trains reviewers to click approve without reading, or nothing does, which works until the first email goes to a strategic account mid-negotiation. Neither setting distinguishes a task creation from a discount change. Gates should be triggered by the risk of the action, not applied to the agent as a whole. The risk-reversibility matrix for human-in-the-loop design shows how to sort them.
No ceiling and no undo
An agent in a retry loop or a prompt change that loosens a filter can run thousands of actions before anyone looks. Without per-run and per-day caps on actions and spend, the only limit is the API quota, and the cost of oversight and error correction lands in the cost-per-qualified-meeting model anyway. And without a record of each field's prior value, rollback means restoring from a backup and losing every legitimate edit made since.
Reference architecture
The architecture has four controls: scoped permissions, action and spend caps, action logging and rollback capability. Approval gates sit on top of them as a policy, not as a fifth queue. The controls are declared once per agent, in a policy file that version control can diff, and enforced by the layers below. A minimal example for an outbound agent:
agent: signal_outbound_v3
identity: svc_agent_outbound # its own user, never shared
read: [Account, Contact, Intent_Signal__c, Opportunity.StageName]
write: [Task.*, Sequence_Enrollment__c.*, Contact.Agent_Last_Touch__c]
deny_if: Account.Open_Opportunity__c = true OR Account.Renewal_Window__c = true
caps: {enrollments_per_run: 50, enrollments_per_day: 300, llm_spend_per_day_usd: 40}
gates:
- action: send_first_email
when: Account.Tier__c = "Strategic" OR model.confidence < 0.80
route: account_owner, expires_after: 24h, on_expiry: drop
log: {fields: [run_id, trigger, inputs_hash, policy_version, output, confidence, prior_values]}
rollback: by run_id
Every value is illustrative; the structure is the point. Each layer of the stack enforces part of it.
Components: CRM record events, intent and enrichment feeds, product usage events, form fills.
Contract to identity & data quality: every trigger carries an event type, a timestamp, a source system and a record ID, so the log can later show exactly what started a run.
Components: one service identity per agent (a dedicated integration user or connected app with its own permission set), resolved account and contact IDs, and suppression flags such as Open_Opportunity__c, Renewal_Window__c and Do_Not_Contact__c computed before the agent runs. Those flags are only as reliable as the identity resolution layer beneath them.
Contract to orchestration: the agent receives only the fields on its read list, already resolved to canonical records. Suppression is a field the agent reads, not a judgment it makes.
Components: a policy check that runs before every write, evaluating deny rules, caps and gate conditions; counters for actions and model spend per run and per day; a circuit breaker that pauses the agent when a cap is hit or the error rate spikes.
Contract to system of record: no write leaves this layer without a policy decision attached: allow, gate or deny, with the policy_version that made it.
Components: field-level security on the agent's permission set, so even a buggy policy engine cannot write Amount or Discount__c; an approval object for gated actions; field history tracking, plus an Agent_Action_Log__c object or warehouse table that stores prior and new values per run_id.
Contract to activation: every agent-made change can be found by run_id and reverted to its prior value without touching edits made by people.
Components: the agent itself; sending and sequencing infrastructure; gated actions delivered to the account owner in Slack or the CRM with a one-click approve or reject, an expiry, and a safe default on expiry.
Contract to leadership: a weekly per-agent view of actions, gates, approval latency, overrides, cap hits and rollbacks.
Here is how the four controls map onto common GTM agent use cases, as an illustrative example rather than a prescription:
| Agent use case | Scoped permissions | Caps | Approval gate | Rollback |
|---|---|---|---|---|
| Signal-based outbound | Read accounts and signals; create tasks and enrollments only | Enrollments per run and per day; model spend per day | First email to strategic accounts or below a confidence threshold | Unenroll by run_id |
| Pipeline hygiene | Write next-step and flag fields; never stage or amount | Records touched per run | Any proposed change to CloseDate goes to the rep as a suggestion | Restore prior values by run_id |
| Handoff brief | Read the closed-won deal; create one brief record | One brief per opportunity | None; the CSM edits the brief | Delete the brief |
| Renewal pricing recommendation | Read only; write to a recommendation object | Recommendations per day | Always; a human sets any price | Nothing to roll back |
| Revenue questions and forecast commentary | Read only, with row-level access matching the asker | Queries per user per day | None for answers; any write is out of scope | Not applicable |
Build sequence
Six steps, each with a test.
Inventory every agent and the identity it runs as
List each agent, copilot action and AI-assisted automation that writes to a revenue system, with the user or token it authenticates as and the objects that identity can edit. The diagnose-before-you-build playbook applies. Test: no two agents, or an agent and a sync, share an identity.
Classify each action by reversibility and blast radius
For every action an agent can take, ask two questions: can it be undone without the customer noticing, and how many records or people can one bad run affect? Tasks and flags are low; emails, ownership changes and pricing are high. Test: every action on the list has a class, agreed by the person who owns the outcome.
Scope the identity and enforce it in the CRM
Give each agent its own permission set with only the objects and fields its action list needs, and deny the rest with field-level security, not with instructions in a prompt. Test: an attempted write to a field outside the list fails at the CRM, not in the agent's code.
Add caps, a circuit breaker and a decision log
Put the policy check in front of every write: deny rules, counters for actions and spend, and a pause when a cap is hit. Log each decision with its run_id, trigger, inputs, policy version, output, confidence and the prior value of every field it changed. Test: a deliberately looped run stops at the cap, and its log entries can reconstruct every action.
Gate by policy, with expiry and a safe default
Route only high-risk actions to a human, deliver them where the owner works, and give each gate an expiry and a default: drop the email, keep the old value, create a task instead. Test: a gated action that nobody answers resolves safely on its own, and the share of actions gated stays small enough that reviewers still read them.
Prove rollback and backtest before widening scope
Reverse a full test run by run_id and confirm human edits survive. Then backtest the agent on around twenty of your own past cases, labelled by your team. We hold every system to the same bar: at least 85 percent agreement with those labels and no uncaught unsafe action, or it does not ship. Test: rollback and backtest results are recorded before any cap is raised or any gate removed.
Build vs. buy: trade-offs
There are three common places to put the governance layer. Tools are examples, not endorsements.
| Approach | Fit | Cost of ownership | Failure risk |
|---|---|---|---|
| Native CRM controls (permission sets, field-level security, approval processes and field history in Salesforce or HubSpot, plus any guardrails in a native agent platform) | Agents that act only inside one CRM; strongest place to enforce scoped permissions | Low. Admin-maintained and already audited by the platform; field history retention limits may require exporting logs | Weak on cross-system caps and spend limits; decision logs rarely include model inputs and confidence |
| Workflow tool as policy layer (for example n8n, Make or Workato sitting between agent and systems) | Agents that act across CRM, sequencing and messaging tools; caps and gates in one place | Moderate. Someone must own policies, counters and log tables | Bypass risk if the agent also holds direct API credentials; the CRM must still enforce permissions underneath |
| Custom policy service with an agent framework (for example a small service in front of agents built on LangGraph or similar) | Several agents, high volume, or a need for versioned policy-as-code and replay | Highest. Engineering time to build, test and maintain | Becomes its own critical system; a bug in the policy engine can allow or block everything at once |
Running it in production
Watch five numbers per agent: actions, share gated, approval latency, override rate and cap hits. A rising gate share means the policy is too broad; a rising override rate means the agent is wrong more often; long approval latency means the gate is routed to the wrong person.
Every gate expires to its safe default, every cap trips the circuit breaker and every agent can be dropped to read-only with one setting, without a redeploy. If the policy engine is down, writes are denied, and rollback by run_id is rehearsed quarterly.
Three sentences carry it: each agent can only touch the records and fields its job needs, and the CRM enforces that; each agent has hard limits on how much it can do in a day, and every action can be traced and undone; humans approve only the actions that would be costly and hard to reverse, and we track how fast they do it.
Where this fits in the system
Governance is what lets each VANDFORT system run without a person watching every step. The Signal-Based Outbound Engine carries the heaviest controls: scoped enrollment rights, daily caps, suppression on open opportunities and renewals, a hard boundary on the segments you named, and no send without a person's approval. The Pipeline Hygiene Sentinel flags and escalates but never edits a deal, changes a stage or moves a close date. Renewal Radar escalates and schedules but never commits your team to a discount or a concession. Revenue Answers and the Forecast Assistant answer and flag rather than write, which removes most of the governance problem before it starts. Every one of them starts on read-only access to your systems of record until you widen it. The full map is on the systems page.
The governance layer also depends on the layers around it. Suppression flags are only as good as identity resolution, and caps only protect you if every agent runs through the same policy check. That is why a forward-deployed engineering engagement starts with the client's own data and ships one system at a time, each with its controls tested before go-live. The GTM engineer vs. RevOps vs. growth engineer decision tree helps decide who owns agent policies.
Sources: SailPoint, AI Agents: The New Attack Surface, research conducted by Dimensional Research (353 IT professionals; May 2025). Deloitte, State of AI in the Enterprise (3,235 IT and business leaders; 2026). Salesforce, State of Sales report (4,050 sales professionals; February 2026). Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (June 2025). Gartner, Gartner Predicts that Guardian Agents will Capture 10-15% of the Agentic AI Market by 2030 (June 2025). OWASP GenAI Security Project, LLM06:2025 Excessive Agency (2025).




