RESOURCES/Methodology
How Markin measures incrementality
Markin measures incrementality with a randomised holdout attached to every decision, read over a full measurement window against a pre-declared primary metric. A result that has not beaten its control does not scale. Reported uplift, attributed conversions and post-hoc comparisons are not accepted as evidence at any stage.
Why attribution is not measurement
Attributed reporting answers a question nobody should care about: of the people who converted, how many touched this programme first. It counts customers who would have converted anyway, and it counts them in favour of whichever programme reached them. That is why programmes that look excellent for years often disappear when someone finally holds a control group.
- Attribution measures correlation with an outcome, not the outcome caused.
- The customers most likely to convert are also the easiest to reach, which biases every attributed number upwards.
- Contacting someone who would have bought anyway has a real cost that attribution reports as a win.
The design, step by step
01
Randomise at the customer, before the decision
Assignment happens before Markin chooses an action, not after. Assigning after the decision leaks the decision into the control group and quietly destroys the comparison.
02
Hold out a stable share of the eligible population
The holdout is drawn from the customers who were eligible for the decision, not from the base at large. Comparing treated customers against the whole base compares two different populations and produces a number that is always flattering.
03
Declare the primary metric before launch
One primary metric, expressed in revenue, fixed before any data arrives. Secondary metrics are recorded but cannot be promoted to primary after the fact.
04
Read over a full measurement window
The window is set by the business cycle being affected, not by impatience. A subscription upsell read at seven days measures novelty. Reading it across a billing cycle measures effect.
05
Apply the stopping rule
Stopping rules are set in advance. Peeking at a running test and stopping when it looks good manufactures significance, and it is the most common way a good measurement practice quietly becomes a bad one.
06
Scale or retire
A result that beats control across the window scales to the rest of the eligible population. A result that does not is retired. There is no third outcome where a programme keeps running while someone thinks about it.
What counts as evidence, and what does not
| Evidence | Accepted | Why |
|---|---|---|
| Randomised holdout, full window, pre-declared metric | Yes | Measures the effect caused by the decision. |
| Pre/post comparison on the same cohort | No | Cannot separate the decision from seasonality or anything else that changed. |
| Treated group against the unexposed rest of the base | No | Different populations. Selection alone produces a positive result. |
| Attributed conversions from the delivery platform | No | Counts customers who would have converted anyway. |
| Model-predicted uplift | No | A forecast, not a measurement. Useful for ranking, never for reporting. |
| Holdout stopped early because the result looked good | No | Optional stopping inflates the effect size. |
Why the haircut exists
BCG finds that when next-best-action programmes are incrementality-tested, 20% to 40% of the measured uplift does not survive. Markin applies that haircut to projections before they are shown, rather than after a customer discovers it.
Independent researchBCG, incrementality in personalisation programmes (2026)
Audit one of your own programmes
Pick the programme with the best-looking numbers. That is where the surprise usually is.
- Was a control group defined before the first send, or reconstructed afterwards?
- Was the control drawn from eligible customers, or from everyone else?
- Was the primary metric written down before launch?
- Was the reported window chosen in advance, or chosen once the data was in?
- Was the test stopped at a pre-agreed point?
- If you removed attributed conversions and used only the holdout comparison, does the programme still pay for itself?
Where holdouts are the wrong instrument
- Legally or contractually mandatory communications. You cannot withhold them, so there is no control group to hold.
- Effects that spill across customers, such as referral or network mechanics, where the control group is contaminated by the treatment.
- Very small eligible populations, where the test would never reach power and the honest answer is to decide on judgement.
- One-off structural changes with no repeatable unit, which belong in a before/after analysis with all its caveats stated.
Markin is an autonomous growth-science team for large B2C businesses. It investigates why revenue per customer is stuck, forms its own hypotheses across marketing, product, pricing and technical health, chooses the next best action for each customer, launches it through the systems the business already runs, and proves every one against a randomised holdout.
Decisioning tools choose between the actions your team already built. Markin decides what to build.
Questions people ask
- How large should a holdout be?
- Large enough to detect the smallest effect that would change your decision, at the significance level you are willing to act on. For large B2C bases that is usually a single-digit percentage of the eligible population, which is cheap. The right way to choose it is a power calculation on the minimum detectable effect, never a round number picked because it feels safe.
- How long should a measurement window be?
- As long as the business cycle the decision affects. For a monthly subscription that means at least one full billing cycle, usually two, so that a pulled-forward purchase is not counted as an incremental one. Reading a revenue effect at seven days measures novelty.
- What is the difference between uplift and incremental revenue?
- Uplift is usually the difference between treated customers and everyone else, which is contaminated by selection. Incremental revenue is the difference between a randomised treated group and a randomised control group drawn from the same eligible population. Only the second one supports a business case.
- Can you measure incrementality without a holdout?
- Sometimes, with quasi-experimental methods such as geographic splits, switchback designs or synthetic controls. They are legitimate when randomisation is impossible, and every one of them rests on assumptions that a holdout does not need. Use them as a fallback, state the assumptions, and do not present the result as equivalent.