Every tool is free to use. Enter your email once and all five open.All resources

80% Have Seen Agents Take Unintended Actions: Governance for Autonomous GTM Agents, the Guardrails, Approval Gates and Audit Trails That Don't Slow Agents Down

A stream of glowing gold particles flows along a glass track lined with brass rails, passes through a square glass and brass gate with a round opening, and continues past a row of thin upright glass panels on a white background.

The scene is a composite, and the details are illustrative. An outbound agent goes live on a Monday with access to the CRM, the sequencing tool and an enrichment API. By Wednesday it has enrolled 2,400 contacts. On Thursday someone notices that 300 of them belong to accounts with open opportunities, owned by AEs who never asked for help. Forty are at a customer in the middle of a renewal. Nobody can say which signal drove each enrollment, because the agent logged a summary rather than its inputs. Unwinding it takes two admins most of Friday, and the CRO then puts every agent action behind a manual review queue. A month later the queue is 900 items deep, reps ignore it and the agent is effectively off.

Both are governance failures. The first had no controls; the second had one, and it was a human.

80%of organizations say AI agents have taken unintended actions (SailPoint, 2025)
44%have policies in place to secure their AI agents (SailPoint, 2025)
21%report a mature governance model for agentic AI (Deloitte, 2026)

The incident data is not reassuring. SailPoint's AI agent research, conducted by Dimensional Research with 353 IT professionals who hold enterprise security responsibilities and published in May 2025, found that 82% of organizations already use AI agents, but only 44% have policies securing them. Eighty percent said their AI agents had taken unintended actions, including accessing unauthorized systems or resources (39%). Deloitte's 2026 State of AI in the Enterprise (3,235 IT and business leaders across 24 countries) found that only 21% of organizations have a mature governance model for agentic AI, meaning clear decision boundaries, real-time monitoring and audit trails.

Sales is moving faster than its controls. Salesforce's State of Sales report (4,050 sales professionals, published February 2026) found that 54% of sellers have already used AI agents and nearly nine in ten plan to by 2027. Gartner's June 2025 prediction that over 40% of agentic AI projects will be canceled by the end of 2027 names inadequate risk controls among the three causes, alongside cost and unclear value. Gartner also predicts that "guardian agents", systems that oversee other agents, will account for at least 10 to 15% of the agentic AI market by 2030 (June 2025).

This is a systems problem, not a policy-document problem and not a people problem. A policy in a slide cannot stop an API call, and a reviewer who sees every action becomes the system's throughput ceiling. The only governance that scales with agent volume is governance the stack enforces on its own: at the permission layer, at the write path and in the log. OWASP's 2025 Top 10 for LLM Applications names the failure pattern directly as "Excessive Agency" and traces it to three root causes: excessive functionality, excessive permissions and excessive autonomy.


Where it breaks

Four anti-patterns recur in stacks that run agents without a governance layer.

The integration user with admin rights

The agent authenticates as a shared integration user, usually the one created years ago for the marketing automation sync, with a profile that can edit every object and field. The agent only needs to read Account, Contact and Intent_Signal__c and create Task records, but it can also update Opportunity.Amount, reassign OwnerId and delete records. The blast radius is the whole CRM, and every change shows up in field history as the integration user, indistinguishable from the sync. The same ownership problem, with people instead of agents, is why territory assignment needs its own audit trail.

Logs that record outcomes, not decisions

Most agent logs capture what happened: "enrolled contact in sequence." Few capture why: which trigger fired, which fields and signals were read, which policy version applied, what the model returned and with what confidence. Without the inputs, a bad action cannot be replayed or traced to the data, the prompt or the rule. A log entry without a run_id joining it to the inputs is a diary, not an audit trail. Agent reliability and the decision log covers what to instrument.

Approval gates on everything, or on nothing

Teams pick one of two settings. Either every action waits for a human, which recreates the bottleneck and trains reviewers to click approve without reading, or nothing does, which works until the first email goes to a strategic account mid-negotiation. Neither setting distinguishes a task creation from a discount change. Gates should be triggered by the risk of the action, not applied to the agent as a whole. The risk-reversibility matrix for human-in-the-loop design shows how to sort them.

No ceiling and no undo

An agent in a retry loop or a prompt change that loosens a filter can run thousands of actions before anyone looks. Without per-run and per-day caps on actions and spend, the only limit is the API quota, and the cost of oversight and error correction lands in the cost-per-qualified-meeting model anyway. And without a record of each field's prior value, rollback means restoring from a backup and losing every legitimate edit made since.

The common thread: each failure puts governance in the wrong place. Rights live in a shared profile instead of a scoped identity, review lives in a person's inbox instead of a policy, and history lives in a summary instead of a structured record. Move each control to the layer that can enforce it.

Reference architecture

The architecture has four controls: scoped permissions, action and spend caps, action logging and rollback capability. Approval gates sit on top of them as a policy, not as a fifth queue. The controls are declared once per agent, in a policy file that version control can diff, and enforced by the layers below. A minimal example for an outbound agent:

agent: signal_outbound_v3
identity: svc_agent_outbound          # its own user, never shared
read:  [Account, Contact, Intent_Signal__c, Opportunity.StageName]
write: [Task.*, Sequence_Enrollment__c.*, Contact.Agent_Last_Touch__c]
deny_if: Account.Open_Opportunity__c = true OR Account.Renewal_Window__c = true
caps:  {enrollments_per_run: 50, enrollments_per_day: 300, llm_spend_per_day_usd: 40}
gates:
  - action: send_first_email
    when: Account.Tier__c = "Strategic" OR model.confidence < 0.80
    route: account_owner, expires_after: 24h, on_expiry: drop
log:   {fields: [run_id, trigger, inputs_hash, policy_version, output, confidence, prior_values]}
rollback: by run_id

Every value is illustrative; the structure is the point. Each layer of the stack enforces part of it.

Sources · Triggers and signals

Components: CRM record events, intent and enrichment feeds, product usage events, form fills.

Contract to identity & data quality: every trigger carries an event type, a timestamp, a source system and a record ID, so the log can later show exactly what started a run.

Identity & data quality · Scoped identities and clean inputs

Components: one service identity per agent (a dedicated integration user or connected app with its own permission set), resolved account and contact IDs, and suppression flags such as Open_Opportunity__c, Renewal_Window__c and Do_Not_Contact__c computed before the agent runs. Those flags are only as reliable as the identity resolution layer beneath them.

Contract to orchestration: the agent receives only the fields on its read list, already resolved to canonical records. Suppression is a field the agent reads, not a judgment it makes.

Orchestration & logic · The policy engine

Components: a policy check that runs before every write, evaluating deny rules, caps and gate conditions; counters for actions and model spend per run and per day; a circuit breaker that pauses the agent when a cap is hit or the error rate spikes.

Contract to system of record: no write leaves this layer without a policy decision attached: allow, gate or deny, with the policy_version that made it.

System of record · Enforced at the write path

Components: field-level security on the agent's permission set, so even a buggy policy engine cannot write Amount or Discount__c; an approval object for gated actions; field history tracking, plus an Agent_Action_Log__c object or warehouse table that stores prior and new values per run_id.

Contract to activation: every agent-made change can be found by run_id and reverted to its prior value without touching edits made by people.

Activation / agents · Bounded action and asynchronous review

Components: the agent itself; sending and sequencing infrastructure; gated actions delivered to the account owner in Slack or the CRM with a one-click approve or reject, an expiry, and a safe default on expiry.

Contract to leadership: a weekly per-agent view of actions, gates, approval latency, overrides, cap hits and rollbacks.

Here is how the four controls map onto common GTM agent use cases, as an illustrative example rather than a prescription:

Agent use caseScoped permissionsCapsApproval gateRollback
Signal-based outboundRead accounts and signals; create tasks and enrollments onlyEnrollments per run and per day; model spend per dayFirst email to strategic accounts or below a confidence thresholdUnenroll by run_id
Pipeline hygieneWrite next-step and flag fields; never stage or amountRecords touched per runAny proposed change to CloseDate goes to the rep as a suggestionRestore prior values by run_id
Handoff briefRead the closed-won deal; create one brief recordOne brief per opportunityNone; the CSM edits the briefDelete the brief
Renewal pricing recommendationRead only; write to a recommendation objectRecommendations per dayAlways; a human sets any priceNothing to roll back
Revenue questions and forecast commentaryRead only, with row-level access matching the askerQueries per user per dayNone for answers; any write is out of scopeNot applicable
Design principle: enforce governance at the layer that can say no without a human, and save human attention for the actions whose cost of being wrong is high and hard to reverse. Permissions limit what is possible, caps limit how much, logs make every action explainable, rollback makes most mistakes cheap, and gates catch only what is left. Whether a task needs an agent at all is a separate call, covered in the agentic vs. static workflow decision framework.

Build sequence

Six steps, each with a test.

Inventory every agent and the identity it runs as

List each agent, copilot action and AI-assisted automation that writes to a revenue system, with the user or token it authenticates as and the objects that identity can edit. The diagnose-before-you-build playbook applies. Test: no two agents, or an agent and a sync, share an identity.

Classify each action by reversibility and blast radius

For every action an agent can take, ask two questions: can it be undone without the customer noticing, and how many records or people can one bad run affect? Tasks and flags are low; emails, ownership changes and pricing are high. Test: every action on the list has a class, agreed by the person who owns the outcome.

Scope the identity and enforce it in the CRM

Give each agent its own permission set with only the objects and fields its action list needs, and deny the rest with field-level security, not with instructions in a prompt. Test: an attempted write to a field outside the list fails at the CRM, not in the agent's code.

Add caps, a circuit breaker and a decision log

Put the policy check in front of every write: deny rules, counters for actions and spend, and a pause when a cap is hit. Log each decision with its run_id, trigger, inputs, policy version, output, confidence and the prior value of every field it changed. Test: a deliberately looped run stops at the cap, and its log entries can reconstruct every action.

Gate by policy, with expiry and a safe default

Route only high-risk actions to a human, deliver them where the owner works, and give each gate an expiry and a default: drop the email, keep the old value, create a task instead. Test: a gated action that nobody answers resolves safely on its own, and the share of actions gated stays small enough that reviewers still read them.

Prove rollback and backtest before widening scope

Reverse a full test run by run_id and confirm human edits survive. Then backtest the agent on around twenty of your own past cases, labelled by your team. We hold every system to the same bar: at least 85 percent agreement with those labels and no uncaught unsafe action, or it does not ship. Test: rollback and backtest results are recorded before any cap is raised or any gate removed.


Build vs. buy: trade-offs

There are three common places to put the governance layer. Tools are examples, not endorsements.

ApproachFitCost of ownershipFailure risk
Native CRM controls (permission sets, field-level security, approval processes and field history in Salesforce or HubSpot, plus any guardrails in a native agent platform)Agents that act only inside one CRM; strongest place to enforce scoped permissionsLow. Admin-maintained and already audited by the platform; field history retention limits may require exporting logsWeak on cross-system caps and spend limits; decision logs rarely include model inputs and confidence
Workflow tool as policy layer (for example n8n, Make or Workato sitting between agent and systems)Agents that act across CRM, sequencing and messaging tools; caps and gates in one placeModerate. Someone must own policies, counters and log tablesBypass risk if the agent also holds direct API credentials; the CRM must still enforce permissions underneath
Custom policy service with an agent framework (for example a small service in front of agents built on LangGraph or similar)Several agents, high volume, or a need for versioned policy-as-code and replayHighest. Engineering time to build, test and maintainBecomes its own critical system; a bug in the policy engine can allow or block everything at once

Running it in production

Monitor

Watch five numbers per agent: actions, share gated, approval latency, override rate and cap hits. A rising gate share means the policy is too broad; a rising override rate means the agent is wrong more often; long approval latency means the gate is routed to the wrong person.

Fail safe

Every gate expires to its safe default, every cap trips the circuit breaker and every agent can be dropped to read-only with one setting, without a redeploy. If the policy engine is down, writes are denied, and rollback by run_id is rehearsed quarterly.

Explain it to leadership

Three sentences carry it: each agent can only touch the records and fields its job needs, and the CRM enforces that; each agent has hard limits on how much it can do in a day, and every action can be traced and undone; humans approve only the actions that would be costly and hard to reverse, and we track how fast they do it.


Where this fits in the system

Governance is what lets each VANDFORT system run without a person watching every step. The Signal-Based Outbound Engine carries the heaviest controls: scoped enrollment rights, daily caps, suppression on open opportunities and renewals, a hard boundary on the segments you named, and no send without a person's approval. The Pipeline Hygiene Sentinel flags and escalates but never edits a deal, changes a stage or moves a close date. Renewal Radar escalates and schedules but never commits your team to a discount or a concession. Revenue Answers and the Forecast Assistant answer and flag rather than write, which removes most of the governance problem before it starts. Every one of them starts on read-only access to your systems of record until you widen it. The full map is on the systems page.

The governance layer also depends on the layers around it. Suppression flags are only as good as identity resolution, and caps only protect you if every agent runs through the same policy check. That is why a forward-deployed engineering engagement starts with the client's own data and ships one system at a time, each with its controls tested before go-live. The GTM engineer vs. RevOps vs. growth engineer decision tree helps decide who owns agent policies.

Sources: SailPoint, AI Agents: The New Attack Surface, research conducted by Dimensional Research (353 IT professionals; May 2025). Deloitte, State of AI in the Enterprise (3,235 IT and business leaders; 2026). Salesforce, State of Sales report (4,050 sales professionals; February 2026). Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (June 2025). Gartner, Gartner Predicts that Guardian Agents will Capture 10-15% of the Agentic AI Market by 2030 (June 2025). OWASP GenAI Security Project, LLM06:2025 Excessive Agency (2025).

Read next