The Markin ARPU report for B2C enterprisesRead now
MARKIN

RESOURCES/Guide

Controlled growth experiments: proving an action caused the result

A controlled growth experiment measures what an action caused rather than what followed it. An eligible population is split into treatment and a randomised control, a primary metric and observation window are fixed in advance, guardrail metrics protect against harm, and the readout resolves to one of three decisions: scale, refine or stop.

Román Via-Dufresne, Co-founder, Markin

Updated 7 August 2026 · 9 min read

Book a demo

Most growth reporting is not evidence

A campaign report tells you what treated customers did. It cannot tell you what they would have done anyway, and in a large B2C base most of them would have done a great deal anyway. The gap between those two numbers is the entire question, and it is usually left unmeasured because measuring it means deliberately withholding the action from someone.

  • Attributed revenue counts customers who were going to convert regardless.
  • Pre and post comparisons absorb seasonality, price changes and traffic mix.
  • Uplift without a control is a description of the treated group, not an effect.

The six things fixed before an experiment runs

Every one of these is written down before the first customer is treated. Deciding any of them afterwards turns an experiment into a story.

  1. 01

    Eligible population

    Who could receive this action at all. Everything is measured inside this population, not across the whole base.

  2. 02

    Treatment and control

    Random assignment within the eligible population. The control receives no action, not a different action, unless the question is explicitly a comparison between two actions.

  3. 03

    Primary metric

    One metric, chosen in advance, at the level the business actually cares about: incremental margin or revenue per eligible customer, not open rate.

  4. 04

    Observation window

    Long enough to capture the behaviour and any pull-forward effect. A window that stops at the first good number measures timing, not value.

  5. 05

    Guardrails

    Metrics that can stop the experiment regardless of the primary result: margin floors, unsubscribe rates, complaint volume, support load.

  6. 06

    Decision rule

    What result leads to scale, to refine, or to stop, agreed before anyone sees the data.

The readout: scale, refine or stop

ResultWhat it meansDecision
Clear positive effect on the primary metric, guardrails intactThe action is worth more than doing nothing for this population.Scale, keeping a smaller permanent holdout to detect decay.
Positive in a subgroup, flat overallThe eligible population was too broad, not the action wrong.Refine the population and rerun rather than scaling as-is.
Flat or inconclusive with adequate powerThe action does not move the metric at this cost.Stop. This is a cheap, normal outcome, not a failure.
Underpowered, direction unclearThe base or the window could not support the question.Stop or redesign. Do not scale on a directional read.
Primary metric up, a guardrail breachedThe gain is being paid for somewhere else.Stop, then re-test with the constraint priced in.

Why the discipline is worth the cost of a holdout

Before you call a result a result

  • Was the control randomised inside the eligible population, or picked afterwards?
  • Was the primary metric chosen before the data arrived?
  • Was the observation window long enough to see pull-forward reverse?
  • Would the experiment have detected an effect worth acting on, given the base size?
  • Did any guardrail move while the primary metric improved?
  • Is the decision rule the one that was agreed at the start?

When a controlled experiment is the wrong instrument

  • Populations too small to reach power in a reasonable window. Judgement and qualitative evidence beat an underpowered test.
  • Changes that cannot ethically or legally be withheld, such as a security fix or a regulatory notice.
  • One-off structural changes with no comparable control group, where a time-series or matched-market method is more honest.
  • Questions about why something happens. An experiment measures effect, not mechanism.

Markin is an autonomous growth-science team for large B2C businesses. It investigates why revenue per customer is stuck, forms its own hypotheses across marketing, product, pricing and technical health, chooses the next best action for each customer, launches it through the systems the business already runs, and proves every one against a randomised holdout.

Decisioning tools choose between the actions your team already built. Markin decides what to build.

Questions people ask

What makes a growth experiment controlled?
A randomised control group drawn from the same eligible population, a primary metric and observation window fixed before the test runs, and guardrail metrics that can stop it. Without random assignment inside the eligible population, differences between the groups explain the result as well as the action does.
How large should a holdout be?
Large enough to detect the smallest effect worth acting on, and no larger. Oversized holdouts cost real revenue; undersized ones produce a number nobody can act on, which costs more. The size follows from the base, the baseline conversion rate and the effect size you would scale on.
Is a flat result a wasted experiment?
No. A flat result with adequate power tells you not to spend on that action, which is a decision with real value. The expensive outcome is scaling something that never worked, which is exactly what happens when there is no control.