---
title: The ROI of Growth Agents 2026: What agentifying growth actually returns, and how to prove it
url: https://markin.ai/roi-report
published: 2026-09-02
publisher: Markin Research
description: Agentifying growth means putting the whole loop under agents: reading the signal, forming the hypothesis, choosing the action, running the experiment and reading the result. The return on that is not the uplift it reports. It is the uplift that survives a randomised holdout, net of the margin given away to produce it. Almost nobody measures that number, which is why almost nobody can defend their growth budget twice.
---

# The ROI of Growth Agents 2026

> Agentifying growth means putting the whole loop under agents: reading the signal, forming the hypothesis, choosing the action, running the experiment and reading the result. The return on that is not the uplift it reports. It is the uplift that survives a randomised holdout, net of the margin given away to produce it. Almost nobody measures that number, which is why almost nobody can defend their growth budget twice.

## Method

- **6 findings.** Each built on independent, named research rather than on our own telemetry.
- **23 sources.** Peer-reviewed economics, consultancy research, sector data and government statistics.
- **2014–2026.** Publication window of the evidence reviewed, weighted towards 2025 and 2026.
- **0 aggregate claims.** We publish no industry-wide uplift figure, because we could not source one we would defend.

## What this report does not claim

- That agentic growth is inevitable, or that every base has the same headroom.
- Any uplift benchmark averaged across the industry. The credible number is the one your own holdout produces.
- That vendor-commissioned ROI studies are evidence of incrementality. They are evidence of value potential, and we label them.

## 01. Enterprise AI spend has scaled. Demonstrated return has not.

The gap between adoption and measurable impact is now the defining fact of enterprise AI, and growth is where it is most visible.

Two independent readings of the 2025 enterprise landscape agree on the shape of the problem. Deployment is near-universal, value capture is concentrated in a small minority, and the majority of programmes cannot produce a figure their own finance function will accept. This is not a technology failure. Nothing in the stack prevents a growth team from measuring incremental margin. What prevents it is that no one designed the programme to be measurable before it launched.

Growth is the acute case because the counterfactual is expensive to preserve. Holding out five percent of a base feels like leaving revenue on the table, so it is the first thing cut when a launch date slips. Twelve months later the programme has produced a great deal of activity, a dashboard that goes up and to the right, and no defensible answer to the only question that matters.

- **95%**, of enterprise generative AI pilots show no measurable P&L return, against 30 to 40 billion dollars of spend.. [independent: MIT Project NANDA, The GenAI Divide: State of AI in Business (2025)](https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf)
- **39%**, of organisations report a measurable EBIT impact from AI, although 88% now use it regularly.. [independent: McKinsey, The State of AI in 2025 (2025)](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-how-organizations-are-rewiring-to-capture-value)
- **~6%**, qualify as high performers capturing enterprise-scale value from AI.. [independent: McKinsey, The State of AI in 2025 (2025)](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-how-organizations-are-rewiring-to-capture-value)

**ROI consideration.** Before the next approval, write down the single number the programme will be judged on in twelve months, and the mechanism that will produce it. If that mechanism is a dashboard comparison rather than a control group, the number does not exist yet. Decide it now, while the design is still cheap to change.

## 02. Standard measurement cannot separate lift from cannibalisation.

This is not a nuance for analysts. It is the reason growth budgets get discounted by finance, and the discount is rational.

Attribution answers a question nobody asked: which touchpoint sat nearest the conversion. It cannot distinguish a customer who acted because of the message from a customer who was going to act anyway and happened to receive one. In a large installed base the second group is the majority, which is why attributed revenue routinely exceeds total revenue once the channels are added up.

The academic record on this is unusually blunt. Across twenty-five large field experiments, the median confidence interval on advertising ROI was more than one hundred percentage points wide: even with millions of subjects and real money at stake, the measured return was frequently indistinguishable from zero. A separate large-scale experiment found that switching off brand-term paid search had no statistically significant effect on new or infrequent buyers, whose behaviour attribution had been crediting to the channel for years.

The practical consequence is severe and specific. A programme optimised on attribution drifts towards the customers who are easiest to reach at the moment of purchase, which is precisely the population where the intervention was least necessary. The margin spent there is not a growth investment. It is a rebate.

- **>100 pts**, median width of the confidence interval on advertising ROI across 25 large field experiments.. [academic: Lewis and Rao, The Unfavorable Economics of Measuring the Returns to Advertising, Quarterly Journal of Economics (2015)](https://doi.org/10.1093/qje/qjv023)
- **Not significant**, effect of removing brand paid search on new and infrequent buyers in a large-scale eBay experiment.. [academic: Blake, Nosko and Tadelis, Consumer Heterogeneity and Paid Search Effectiveness, Econometrica (2015)](https://doi.org/10.3982/ecta12423)
- **20–40%**, of active next-best-action programmes deliver marginal or negative lift once incrementality-tested.. [independent: BCG, How Measurement Is Evolving in Next-Best Action (2026)](https://www.bcg.com/publications/2026/measuring-incrementality-in-next-best-action-programs)

**ROI consideration.** Take the last uplift figure your team reported and ask three questions of it: was the control group randomised before launch, was the window at least eight weeks, and was the incentive margin netted out. If any answer is no, the figure is a description of activity, not a measurement of return.

## 03. Personalisation programmes are abandoned for lack of proof, not lack of technology.

The forecast that eighty percent of investing marketers would walk away by 2025 was not about tooling. It was about the inability to show a return.

Gartner's prediction has aged into a description. The teams that retreated did not lack a platform: most had bought two. What they lacked was a way to demonstrate that the programme produced revenue the business would not otherwise have earned, and without that demonstration the budget line loses every argument it enters.

The counterweight is that the value is real where it is captured. Customers expect relevance and punish its absence, and the firms that operationalise it grow faster than their peers. Both things are true at once: the value pool exists, and most attempts to reach it fail on measurement rather than on modelling.

The honest reading is that personalisation was sold as a technology purchase when it is an operating-model change. The technology arrived on schedule. The measurement discipline, the suppression rules and the willingness to retire a losing programme did not.

- **80%**, of marketers who had invested in personalisation were predicted to abandon it by 2025, citing lack of ROI.. [independent: Gartner, Predicts 2020: Marketers, They Are Just Not That Into You (2019)](https://www.marketingdive.com/news/will-personalizations-role-in-marketing-shrink-as-challenges-grow/568607/)
- **71%**, of consumers expect personalised interaction, and 76% are frustrated when they do not get it.. [independent: McKinsey, The Value of Getting Personalization Right or Wrong Is Multiplying (2021)](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/the-value-of-getting-personalization-right-or-wrong-is-multiplying)
- **$217M / $385M**, annual incremental and retained revenue modelled for a composite enterprise running agentic decisioning at scale. Vendor-commissioned study; treat as value potential, not as verified incrementality.. [vendor-commissioned: Forrester Consulting, Total Economic Impact of Pega Customer Decision Hub (2025), commissioned by Pega](https://www.pega.com/forrester-tei-decision-hub)

**ROI consideration.** Audit your own stack for the abandonment pattern before it reaches you: count the personalisation use cases live today, then count the ones with a control group. If the second number is zero, the programme is one budget cycle from being cut, regardless of how well it performs.

## 04. The teams that measure hardest are the teams that grow fastest.

Rigour is not a tax on growth. In the available research it is the strongest observable correlate of it.

BCG's measurement work finds that organisations with mature incrementality practice see materially higher revenue growth than peers. The causal story is not mysterious. A team that can tell a real effect from a coincidence stops funding the coincidences, and the released budget compounds into the programmes that work.

The same research warns about the failure modes that make weak measurement look strong: novelty effects that inflate the first weeks of any new programme, halo and pull-forward effects that borrow revenue from next quarter, and contamination between overlapping experiments. Each of them biases in the same direction, which is why unaudited programmes are almost never revised downwards.

This is the argument to take into a budget meeting. Not that the ceiling is higher than finance believes, but that the floor is knowable, and that knowing it is what turns a pilot into an operating model.

- **up to 70%**, higher revenue growth reported for marketing measurement leaders versus their peers.. [independent: BCG, Six Steps to More Effective Marketing Measurement (2025)](https://www.bcg.com/publications/2025/six-steps-to-more-effective-marketing-measurement)
- **8–12 weeks**, recommended minimum window before drawing conclusions, because novelty effects inflate early results.. [independent: BCG, How Measurement Is Evolving in Next-Best Action (2026)](https://www.bcg.com/publications/2026/measuring-incrementality-in-next-best-action-programs)
- **Overstated**, programme-level ROI when halo, pull-forward and experiment contamination go uncorrected.. [independent: BCG, How Measurement Is Evolving in Next-Best Action (2026)](https://www.bcg.com/publications/2026/measuring-incrementality-in-next-best-action-programs)

**ROI consideration.** Institutionalise one number rather than a scorecard: incremental gross margin per contacted customer, against holdout, read monthly. Every other growth metric is diagnostic. This is the one that decides whether the programme continues.

## 05. Acquisition has stopped covering the leak. The installed base is where the growth is.

Two of the largest subscription categories in the world crossed the same threshold in 2025: the cost of a new customer now exceeds what the base loses without intervention.

Premium streaming subscriber growth halved in a single year, and the sector's own commentary has moved from acquisition to retention architecture and unit economics. Telecoms shows the same pressure from the other direction: roughly a third of customers globally are actively considering a switch, a rate no acquisition budget can outrun.

When acquisition stops being the lever, the economics of the base become the whole game, and they are governed by four questions asked millions of times a month. Who is worth contacting, about what, when, and at what cost. That is a decision problem, not a campaign problem, and its resolution is where incremental ARPU comes from.

The compounding matters more than any single lever. An activated customer feeds cross-sell; a retained high-ARPU customer keeps feeding both for years; a suppressed unnecessary discount returns margin without sending a single additional message. Business cases built on one lever are fragile precisely because these effects are multiplicative.

- **+7% vs +12%**, premium SVOD subscriber growth in 2025 versus 2024. The acquisition-led era has ended.. [independent: Antenna, State of Subscriptions: Premium SVOD 2025 Year in Review (2026)](https://www.antenna.live/insights/antenna-q126-state-of-subscriptions-report-premium-svod-2025-year-in-review)
- **30%**, of telco customers globally are actively considering switching provider.. [independent: Simon-Kucher, Global Telecommunications Study (2026)](https://www.simon-kucher.com/en/insights/global-telecommunications-study)
- **5 pools**, where incremental ARPU is available inside an existing base: activation, cross-sell, churn prevention, win-back and recovered promotional waste.. [independent: BCG, How Measurement Is Evolving in Next-Best Action (2026)](https://www.bcg.com/publications/2026/measuring-incrementality-in-next-best-action-programs)

**ROI consideration.** Split your growth plan into acquisition and base economics, and price them on the same basis: fully loaded cost per incremental gross-margin euro. In most large B2C bases the second column wins by a distance nobody has bothered to quantify.

## 06. The constraint is not campaign volume. It is how many decisions a team can defend.

A growth team's real capacity is the number of hypotheses it can generate, test honestly and retire per year. That number is small, and it is expensive.

Price the constraint properly. A senior data scientist costs well into six figures fully loaded, and the market rate for the causal-inference skills this work needs sits above the published median. That person's output is not dashboards: it is a bounded number of well-designed tests per year, most of which will fail, because most ideas fail everywhere they are measured properly.

The experimentation literature is clear that a high failure rate is the normal state of a healthy programme, not a symptom of a weak one. The implication is arithmetic. If most hypotheses fail, the value of the function is set by how many it can put through the loop, and by how cheaply it can kill the losers. A team that runs twelve tests a year is making twelve bets. A system that runs the same design over thousands of segments continuously is making a different kind of investment entirely.

This is the part of the case that is usually written as a headcount saving, and that framing is both wrong and weak. Nothing is removed. What changes is where the scarce hours go: away from rebuilding the same lift report every month, and towards deciding what is worth testing next.

- **$112,590**, median annual pay for data scientists, against a market band well above it for senior causal-inference skills.. [government: US Bureau of Labor Statistics, Occupational Outlook Handbook: Data Scientists (2024)](https://www.bls.gov/ooh/math/data-scientists.htm)
- **Most fail**, of ideas tested in mature online experimentation programmes, which is the expected outcome and the reason throughput decides value.. [academic: Kohavi and Longbotham, Online Controlled Experiments and A/B Tests (2023)](https://exp-platform.com/Documents/2023-03-11EncyclopeiaMLDSABTestingFinal.pdf)
- **8 weeks**, minimum honest read on a revenue metric, which caps how many sequential tests a single team can complete in a year.. [independent: BCG, How Measurement Is Evolving in Next-Best Action (2026)](https://www.bcg.com/publications/2026/measuring-incrementality-in-next-best-action-programs)

**ROI consideration.** Count the hypotheses your team put through a controlled test last year, and divide the fully loaded cost of the function by that number. That is your current cost per decision. It is the only baseline against which any agentic growth investment should be compared.

## The ROI equation

`( base × contactable % × ARPU × verified lift % × gross margin % ) − incentive margin − programme cost`

- **Contactable, not total.** The addressable base is the share you can legally, technically and tolerably reach in a given month. Fatigue caps belong inside the model, not in a footnote under it.
- **Margin, not revenue.** Incremental revenue at a 62% gross margin is not incremental revenue. Every scenario in this report is expressed in margin, because that is the currency the case is judged in.
- **Verified, not reported.** Apply the haircut inside the model, up front. BCG's band implies discounting reported lift by 20% to 40% before anyone else does it for you.
- **Net of the give-away.** Discounts, credits and save offers are the cost of producing the lift. Subtracting them is the difference between a business case and a press release.

## Scenarios

| Base | Incremental margin (verified) | Programme cost | Net | ROI |
| --- | --- | --- | --- | --- |
| 1M customers | $1,687,392 | $900,000 | $787,392 | 1.9× |
| 5M customers | $8,436,960 | $900,000 | $7,536,960 | 9.4× |
| 20M customers | $33,747,840 | $900,000 | $32,847,840 | 37.5× |

Assumptions: ARPU $24/month; 45% contactable; 62% gross margin; 3% reported lift on addressable revenue; programme cost $900K/year.

## The holdout protocol

1. **Hold out before you launch.** A randomised control group carved from the same population, at the same moment, under the same eligibility rules. A control group assembled after the fact is a matched sample, and a matched sample inherits every bias that made the treated group different.
2. **Size the test for revenue, not conversion.** Revenue per user is heavily right-skewed, so detecting a small relative lift needs far more exposures than a conversion test of the same effect size. Size it before you run it or you will read noise with confidence.
3. **Read the full window.** Eight weeks minimum, twelve where the purchase cycle is long. Anything read in week two is measuring novelty, and novelty decays in every published series.
4. **Net out the give-away.** Subtract every discount, credit and save offer from the incremental revenue before anyone calls it a result. This single step reverses the sign of more retention programmes than any other.
5. **Keep a permanent holdout.** A small standing control is the only instrument that tells you, a year later, whether the programme is still producing lift or merely producing activity. Budget for it as infrastructure, not as a test.

## The diligence checklist

- Can you show one uplift figure measured against a randomised holdout, with the window and the population disclosed?
- What share of the actions the system proposes are suppressed, and under what economic rule?
- How is incentive margin netted out of the reported lift?
- What happens to the number if we keep a permanent 5% holdout for twelve months?
- Which of the five value pools does this address, and which does it leave untouched?
- Who owns the decision when the model and the campaign calendar disagree?

## Conclusion

The pattern across the evidence is consistent and uncomfortable. Enterprises are spending at scale on growth intelligence, a minority can demonstrate a return, and the difference between the two groups is almost never the sophistication of the model. It is whether a control group existed before the launch date.

This is good news for anyone willing to be rigorous, because the discipline is cheap relative to the budget it protects. A five percent holdout, an eight-week window and an honest netting of incentive margin cost a rounding error and produce the only number that survives a second budget cycle.

Markin is built against this standard rather than around it. We publish no industry uplift benchmark, we report our own results as a range measured against holdout, and we consider the checklist above fair game to use on us.

Full report: https://markin.ai/roi-report