The Markin ROI Report for Enterprise Growth TeamsRead now
MARKIN
Field notes
Guides13 min read

AI agents for marketing: what they are and how to evaluate them

AI agents for marketing explained: the four categories on the market, how to evaluate an agent, and the metrics that judge one.

Marc Sanchez
  • #AI agents
  • #Decisioning
  • #Guides
AI agents for marketing: what they are and how to evaluate them

AI agents for marketing are systems that do the work a growth team does between the data and the campaign: read signals, form a hypothesis, size the opportunity, propose an action, and measure the result against a holdout. They are not chatbots that write copy faster. The useful ones change how many decisions a company can make per week, and how many of those decisions are provably worth making.

This guide covers what an AI marketing agent actually is, the four categories on the market, how to evaluate one, and how to measure whether it earned its place in the stack.

1. What an AI agent for marketing is

An agent, in the technical sense, is a system that is given an objective, has tools it can call, and decides for itself which sequence of calls gets closest to the objective. Applied to marketing, the objective is commercial: incremental revenue, retained ARPU, margin per contact. The tools are the warehouse, the customer data platform, the model layer and the execution channels the company already runs.

The distinction that matters for a buyer is autonomy over the decision, not autonomy over the send. Most tools marketed as agents automate the send: a human decides who gets what, the tool executes it faster. An agent worth the name decides which opportunity is worth acting on at all, which is the part of the job that historically required a data scientist and a week.

2. The four categories on the market

  1. Content agents. Generate copy, subject lines, creative variants. High volume, low decision content. They raise the number of assets a team can ship, not the number of decisions it can prove.
  2. Workflow agents. Build audiences, assemble journeys and schedule campaigns from a prompt. They compress the operational cost of executing a plan a human already wrote.
  3. Service and conversation agents. Handle inbound conversation, resolve intent, sometimes upsell. Real value, but the surface is the conversation, not the revenue portfolio.
  4. Decisioning agents. Read first-party signals continuously, propose hypotheses about where revenue is leaking or unclaimed, size each one, select the action, and run it against a control group. This is the category that touches ARPU directly, and the one with the fewest credible entrants.

Most enterprise stacks already have some of the first three. The gap is almost always the fourth: nobody is deciding, at the scale of the whole base, which of a thousand possible interventions is worth the contact this week.

3. What AI agents change about growth capacity

A large B2C operator typically runs somewhere between four and twenty campaign decisions per week, bounded not by channel capacity but by analytical capacity. Each decision needs a hypothesis, a segment, a size estimate and a measurement plan. Five senior analysts can produce a few dozen of those a quarter, and most of that time goes to the plumbing rather than the thinking.

Agents raise the ceiling on the two steps that were never parallelisable: hypothesis authorship and opportunity sizing. When those become continuous, the constraint moves downstream, to how much holdout traffic the business is willing to spend and how quickly humans can adjudicate what the agents propose. That is a better constraint to have, because it is one you can price.

4. How to evaluate an AI agent for marketing

  1. Does it hold a control group? If impact is reported as opens, clicks or attributed conversions rather than treated-minus-holdout, the number is not causal and cannot be defended in a board review.
  2. Can it decide to do nothing? An agent that always finds a reason to send is optimising activity, not margin. Holding should be a recorded, auditable decision.
  3. Where does the data live? Agents that require a full migration into a vendor's own data platform push the decision layer behind a multi-quarter programme. Reading the warehouse in place is strictly better.
  4. Is every recommendation traceable? Each proposed action should resolve back to the signals, the model output and the evidence that produced it. Without provenance, review is opinion.
  5. What stays human? Customer-facing changes should require approval by policy, not by convention. The agent drafts and sizes; a human ships.

5. The metrics that actually judge an agent

  1. Decision volume. Distinct, measured decisions shipped per month. This is the throughput number, and the one that moves first.
  2. Win rate. Share of shipped experiments that move the target metric with significance. A rising decision volume with a falling win rate means the agent is generating noise.
  3. Incremental ARPU. Treated-minus-holdout revenue per customer over a fixed window. The only number that survives a CFO conversation.
  4. Cost per proven decision. Programme cost divided by decisions that produced a significant read. It falls fast when agents work and stays flat when they do not.

6. Failure modes to plan for

6.1 Activity inflation

The easiest way for an agent to look productive is to propose more contacts. Contact frequency caps and a margin-per-contact objective, set before the agent is switched on, are the mitigation.

6.2 Confirmation search

Agents that both propose a hypothesis and gather the evidence for it will find the evidence. Separating those two jobs across different agents, with the stress-testing agent blind to the original framing, kills a meaningful share of bad hypotheses before they consume holdout.

6.3 Measurement drift

When holdouts are dropped for a quarter to hit a number, the loop loses the only feedback it has. The holdout is not a tax on the programme; it is the programme.

7. Where to start

Pick one revenue outcome with a clean measurement path: cross-sell into an under-penetrated product, retention of a defined value band, or reactivation of a lapsed cohort. Give the agent the signals for that outcome, preserve a holdout, and read it for a full cycle before widening scope. Agents compound where the measurement is clean and stall where it is not.

For the decisioning vocabulary used throughout this guide, see next best action marketing. For how agents use external context safely, see agents on the open web.

Frequently asked

Questions readers ask about this.

What are AI agents for marketing?
AI agents for marketing are systems given a commercial objective, a set of tools such as the warehouse, models and execution channels, and the autonomy to decide which sequence of actions gets closest to that objective. In practice they read first-party signals, propose hypotheses about revenue opportunities, size them, select an action and measure the result against a holdout.
How are AI marketing agents different from marketing automation?
Marketing automation executes a plan a human wrote: rules, triggers and journeys defined in advance. An agent decides what the plan should be, per customer per moment, and can conclude that no action is worth taking. Automation compresses execution cost; agents raise the number of decisions a team can make and prove.
How do you use an AI agent for marketing?
Start with one revenue outcome that has a clean measurement path, such as cross-sell into an under-penetrated product or retention of a defined value band. Give the agent the signals for that outcome, preserve a randomised holdout, require human approval before anything customer-facing ships, and read incremental ARPU for a full cycle before widening scope.
Are AI agents for marketing safe to let run autonomously?
Customer-facing changes should require human approval by policy. The defensible split is that agents run autonomously on analysis, hypothesis generation, sizing and measurement, while a human adjudicates anything a customer would see. Every recommendation should resolve back to the signals and evidence that produced it.
How do you measure the ROI of AI marketing agents?
On four numbers: decision volume (measured decisions shipped per month), experiment win rate (share moving the target metric with significance), incremental ARPU (treated minus holdout over a fixed window), and cost per proven decision. Opens, clicks and attributed conversions are not causal and do not count.
What is agentic marketing?
Agentic marketing is an operating model where autonomous agents run the analytical loop continuously instead of a team running it on a campaign calendar: signals in, hypotheses generated, opportunities sized, actions chosen, experiments read against control, and learnings fed back into the next cycle.

See it in the product

This runs in Markin today.

The same loops this note describes run 24/7 against your customer base. Watch the workspace decide, experiment and execute 1:1.