NotFairNotFair
Start now
← Back to blog

AI Agent for Operations: How It Actually Works

Discover how an ai agent for operations transforms ad and campaign workflows using live reads, approval-gated writes, and measurable governance controls.

Tong Chen and Yuting Zhong16 min read
AI Agent for Operations: How It Actually Works

AI agents are moving into operations, but adoption is still uneven: 62% of organizations are at least experimenting with AI agents, while only 23% have scaled an agentic AI system in at least one business function. An AI agent for operations uses foundation models to read live business systems, propose or execute governed actions, and record auditable diffs so changes remain reversible and accountable.

The practical problem is familiar. A campaign manager notices that lead costs are rising, opens several advertising accounts, checks search terms in one dashboard, compares conversion data in another, then searches the CRM for sales quality. By the time the evidence is assembled, an automated rule may already have changed a bid, paused an ad, or moved budget without a clear record of why.

A production-grade agent shouldn't add another chat window to that process. It should connect the relevant systems, retrieve current data, explain the likely cause, produce a constrained change proposal, wait for approval when a write carries risk, and preserve the full history afterward. The useful distinction isn't whether an agent can call a tool. It's whether your team can trust what it read, understand what it wants to change, and undo the change when the result isn't acceptable.

Table of Contents

Why Ad and Campaign Ops Need Governed Agents

A paid media operator may start with a rising cost per lead, then trace the change through loose-match search terms, budget constraints, platform conversion data, CRM-qualified leads, and a campaign's learning phase. Those signals rarely sit in one system. The operator must assemble the chain across Google Ads or Meta Ads, analytics, search data, and the CRM before deciding what changed.

An automated rule can act faster while seeing less. It might pause an underperforming ad despite a new landing page, reduce spend on a campaign producing high-quality pipeline, or add a negative keyword that removes valuable traffic. The problem may come from incomplete data, stale exports, or permissions that treat a read, a recommendation, and a write as equally safe. A reasonable action can still produce an untraceable outcome.

A stressed office worker sitting at a cluttered desk surrounded by charts, financial documents, and laptops.

The operational gap behind the hype

McKinsey's 2025 Global Survey on the State of AI describes agentic systems as foundation-model-based systems that can plan and execute multiple workflow steps in production. For marketing teams, that distinction separates an ad-copy assistant from an operational agent. The latter reads account state, selects tools, evaluates connected evidence, and may change an external system.

The survey reports that scaled agent deployment remains concentrated in only one or two business functions, with no individual function exceeding 10% of respondents reporting scaled deployment. That limited adoption makes operational discipline a deployment requirement, not a presentation detail. Teams need clear ownership for every platform change and a record that supports causal review afterward.

A governed workflow should follow a defined sequence:

  • Read current evidence: Inspect spend, search terms, conversion paths, campaign settings, and CRM outcomes at query time.
  • Explain the diagnosis: Separate observed facts from assumptions and flag conflicting evidence.
  • Draft a bounded action: Show the exact keyword, budget, or status change before execution.
  • Obtain approval: Route material mutations to an accountable operator.
  • Record the result: Store the proposed diff, decision, platform response, and reversal path.

Live reads reduce stale-context decisions. Approval-gated writes limit the blast radius, while reversible changes let operators test an intervention and restore the prior state. Together, these controls make it possible to examine whether a platform change caused the performance movement that followed.

Practical rule: If an operator cannot reconstruct why an account changed, the workflow is not governed, even if performance improved.

Governance covers identity, access, retention, ownership, and tool permissions beyond the advertising platform. Teams can consult Menza's best practices guide for a broader governance reference. For implementation-level controls, the NotFair safety reference shows how permission boundaries, approvals, and reversibility can be made visible in an operational workflow.

Core Architecture Behind an Operational Agent

A budget edit can be technically valid and operationally wrong. The architecture behind an operational agent must show where data comes from, which component selects an action, and which control decides whether that action can reach a platform. Four responsibilities provide that separation: connectors, an MCP layer, the model and client, and a control plane.

The connector communicates with a business system through its supported interface. In marketing operations, it may retrieve Google Ads spend, search terms, keywords, and budgets; Meta Ads campaign status and budget data; Google Search Console queries and index coverage; GA4 traffic and conversions; or CRM lead and pipeline context. Each connector should expose defined tools and fields. Account-wide access makes permissions harder to limit and failures harder to trace.

The Model Context Protocol layer standardizes how an AI client discovers and calls those tools. A hosted MCP server can present a consistent interface to advertising, analytics, and CRM services. An operator can then use a compatible client such as Claude, Codex, Cursor, OpenClaw, or Hermes. Business-system permissions remain in an inspectable service layer, rather than being rebuilt separately for every client.

A diagram illustrating the core architecture of an operational AI agent, including retrieval, hosted model, action tools, and control plane.

The four layers that matter

Retrieval layer. This layer fetches current account state and relevant history. Structured results should give the agent enough context to distinguish an active campaign from an archived one, or a platform conversion from a CRM-qualified opportunity. Field definitions and timestamps help operators judge whether the evidence supports a proposed change.

Hosted model and client. The model interprets the operator's request, selects from the available tools, and produces a diagnosis or proposal. The client provides the conversational surface. Authorization needs enforcement outside that interface, so a different client cannot bypass the same permission rules.

Action tools. Read tools and write tools require different scopes. A read tool may inspect a search-term report. A write tool may pause a campaign, change a budget, or edit a keyword. Each write tool should define permitted fields, validation rules, and the expected resulting state before a request reaches the platform.

Control plane. This layer manages OAuth, organization identity, approval requirements, diff generation, logging, conflict checks, and rollback. It converts an intended action into a permitted operation with an accountable record. That record supports causal review when platform changes are followed by performance movement.

Teams building custom systems can consult AI agent development with Wonderment Apps for implementation context. Marketing operators should still require clear answers about permissions, tool schemas, approval boundaries, and audit reconstruction. The NotFair explanation of how its system works offers a concrete example of a hosted MCP pattern. The architecture is ready for production when these controls remain effective regardless of which client generated the request.

Live Reads and Approval-Gated Writes Explained

Static exports are useful for reporting, but they're a weak foundation for an agent diagnosing a live account. An export can omit a recent budget change, a newly added search term, a campaign status transition, or the current learning context. The agent may produce a coherent explanation that no longer describes the account an operator is about to change.

Live reads solve that timing problem by retrieving the relevant state at query time. The agent can inspect current spend, search terms, keywords, campaign settings, and available analytics or CRM context before forming a recommendation. That doesn't guarantee a correct diagnosis, but it reduces one common source of error: acting on a snapshot that has already gone stale.

Approval-gated writes solve a different problem. They control what happens after the diagnosis. A safe workflow doesn't translate “reduce wasted spend” directly into an unrestricted API call. It turns the recommendation into an explicit proposal that a human can review.

Workflow pattern What the operator sees Main operational weakness
Static export and chat A report or file interpreted after collection The account may have changed since the export
Raw automation A rule or agent writes directly to the platform The action can exceed intent and be hard to reconstruct
Live read with draft Current evidence and a proposed action Requires a deliberate approval step
Live read with approval-gated write Current evidence, exact diff, authorization, and result Needs reliable logging and tested rollback

What an explicit diff should contain

A diff should show the original value, the proposed value, and the object affected. “Optimize the account” isn't reviewable. “Pause this campaign,” “change this budget,” or “add these negative keywords” gives the operator something concrete to accept or reject.

The system should also record:

  • Original state: The values that existed before the request.
  • Changed fields: The exact properties the write will mutate.
  • Reason and evidence: The diagnosis or rule supporting the proposal.
  • Approval decision: The person or organization that authorized it.
  • Platform response: Whether the external system accepted, rejected, or partially applied the request.
  • Recovery path: The restore operation and any constraints on reversal.

Government guidance summarized in the Oklahoma AI agent standard treats reversibility and authorization as technical control properties. It distinguishes reversible, conditionally reversible, and irreversible actions, and recommends previews, confirmations, versioning, rollback procedures, and human approval where risk requires it.

That classification maps cleanly to paid media. A live read is generally lower risk than a budget mutation. A draft negative keyword list is lower risk than publishing it. A campaign pause may be reversible, but the commercial effect can still matter. The agent should therefore narrow its action space, ask for approval at the right boundary, and make undo a tested operation rather than a promise in documentation.

For Meta Ads workflows, the Meta Ads MCP reference shows the kind of platform-specific operation that benefits from this separation. The model can investigate and prepare an action, while the write remains constrained by the service's approval and logging rules.

Real Operations Workflows and Use Cases

A search account with rising cost per lead rarely has one clean cause. Search terms may drift, match behavior may broaden, and CRM outcomes may contradict the platform's conversion signal. An operations agent should turn that situation into a working case: identify mismatched queries, quantify their spend exposure, compare platform conversions with downstream pipeline, and record which account objects require attention.

The useful output is an ordered work queue, not a catalogue of suspicious terms. Rank findings by evidence strength and commercial exposure, then attach the exact campaign, ad group, keyword, landing page, or conversion path involved. Negative keyword proposals should include their intended scope. Budget and status changes belong in a separate review tier because they can alter delivery before the underlying diagnosis is settled.

A cute AI robot connected to various advertising dashboards representing automated marketing operations and campaign management.

Conflicting signals need triage rules

Diagnosis becomes harder when connected systems recommend different actions. Analytics may report weak engagement, the advertising platform may show efficient conversions, and the CRM may show that those leads rarely become qualified opportunities. The agent should not choose a winner from isolated metrics. It should label the conflict, identify the conversion stage each system measures, and route the case to an operator when the commercial interpretation remains uncertain.

A practical triage record includes the conflicting observations, the data windows used, identity or attribution limitations, and the decision that each team would make from its own system. For example, a search query can look inefficient in Google Ads while producing qualified pipeline in the CRM. The appropriate next action may be a landing-page or lead-quality investigation rather than an exclusion. Meta Ads budget recommendations require the same treatment, with delivery, conversion context, account constraints, and source and destination budgets reviewed together.

Partial success is a normal failure mode

External APIs can accept some requested changes and reject others. The agent should process each operation as an individual result, preserving the platform response, error category, and resulting object state. A batch that partly succeeds must not be reported as completed.

The recovery workflow should separate retryable errors, invalid requests, permission failures, and conflicts caused by a newer human change. Retry only the operations that remain safe and valid. If a campaign was changed between proposal and execution, stop the retry and return the case to an operator with the current state and the original intent. A rollback should restore the prior configuration only when it will not overwrite a later authorized change.

Operator hand-off needs a complete case

A hand-off should give the next person enough context to decide without reopening every dashboard. Include the question, affected objects, supporting observations, unresolved conflicts, attempted operations, platform responses, and the specific decision required. The record should distinguish the agent's recommendation from the human decision and the final account state.

Chat can provide the entry point, while linked findings keep the investigation auditable. The operator can sort work by spend at risk, lead quality, delivery impact, or urgency, then inspect the underlying objects before deciding.

The video below illustrates connecting AI clients with operational tools. It does not demonstrate safe modification of a live account without the controls required for production use.

Benefits, Risks, and What to Measure First

The strongest case for an AI agent for operations is that it gives human judgment better context and a more disciplined path from observation to action. An operator can ask a focused question, inspect connected systems, receive a ranked diagnosis, and review a concrete proposal without manually reconstructing every relationship.

The practical benefits include:

  • Faster investigation: The agent gathers related evidence from advertising, analytics, search, and CRM systems in one working session.
  • Clearer action review: Explicit diffs make proposed changes easier to challenge before publication.
  • Better operational memory: Logged decisions connect diagnosis, approval, platform response, and reversal.
  • More consistent triage: Teams can apply the same diagnostic questions and permitted-action rules across accounts.

The risks require equal attention. A fluent answer can conceal an incorrect tool selection, invalid arguments, an incomplete read, or a mistaken attribution assumption. The agent can also execute the wrong action consistently when permissions, validation rules, and acceptance tests are poorly designed. Governed deployment limits the blast radius, but it does not replace review of the diagnosis itself.

Repeatability is the first quality gate

A recent review of 12 agent benchmarks found that performance could fall from 60% on a single attempt to 25% when an eight-run consistency requirement was applied. Comparable workflows could vary by as much as 50× in inference cost, according to the benchmark review. These figures do not forecast every marketing agent. They establish a testing principle: one successful demonstration is not enough.

Replay representative diagnosis tasks across repeated, stateful runs. Check the full operation, not only the final answer:

  • Did the agent select the correct account and tool?
  • Were the arguments valid and within policy?
  • Did it distinguish a recommendation from an approved action?
  • Did it produce the same diff when account state was unchanged?
  • Did it avoid unauthorized writes?
  • Were latency and cost acceptable for the workflow?

A production acceptance test should reward stable behavior, not impressive improvisation.

Causal measurement needs a deliberate experiment. For a lead-generation search campaign, keep a comparable set of campaigns or geographic areas outside the change, while applying the approved adjustment to the treatment set. Define the treatment, holdout, observation window, primary outcome, and guardrails before execution. Qualified pipeline or incremental profit should carry more weight than clicks, conversions, or hours saved. Analysts should also record auction conditions, seasonality, sales follow-up, tracking changes, and attribution differences that could affect the result.

For example, if the agent proposes adding negative keywords, compare qualified outcomes for the changed group with the unchanged holdout, while checking delivery and lead-quality guardrails. If the holdout is contaminated by other optimizations, or the groups differ materially before launch, treat the result as directional rather than causal. The agent's recommendation, human approval, approved diff, execution record, and outcome analysis should remain separate evidence.

McKinsey's research on the state of AI highlights the broader measurement challenge: AI attribution can improve predictive accuracy while remaining causally unidentified without experimental calibration, and external validation remains under-tested. A more capable agent can increase optimization volume while making accountability harder if measurement begins only after automation is live.

Adoption Checklist for Marketing and Operations Teams

Start with a read-only workflow that has a clear operator owner and a defined success condition. Then evaluate the following before enabling writes:

  • Permission scope: Can the agent access only the accounts, fields, and actions it needs?
  • Approval boundary: Which changes require human confirmation, and does the interface show an exact diff?
  • Rollback: Does undo restore the original state, and has the team tested partial failures?
  • Audit reconstruction: Can someone trace every read, proposal, approval, write, platform response, and reversal?
  • Repeatability: Does the agent produce stable results across representative repeated tasks?
  • Causal design: Is there a baseline and, where feasible, a holdout or experiment for important changes?
  • Multi-agent ownership: Who resolves conflicts when analytics, advertising, and CRM agents recommend different actions?

Agent count isn't a maturity metric. Narrow permissions, explicit diffs, immutable logs, and a single accountable owner are usually more valuable than expanding autonomy before the operating model is ready.


NotFair connects AI clients with live advertising, analytics, search, and CRM data, while supporting approval-gated changes, explicit diffs, logged history, and one-call undo. Visit NotFair to evaluate a governed workflow for campaign diagnosis and operational changes.