The Markin ARPU report for B2C enterprisesRead now
MARKIN

COMPARE/Markin + data science

Markin and your data science team

+17–35% ARPU against holdoutObserved range across Markin deployments, measured on treated cohorts.

Markin runs on top of your data science team, not instead of it. The team owns features, models, economics and causal design. Markin behaves like an extra bench of scientists that never sleeps: it writes its own hypotheses, marketing, product, pricing or a technical anomaly holding growth back, sizes them, launches them in the systems you already run, and reads every one against a randomised control group, at a volume no analyst team can sustain by hand.

What is at stake

A decision layer is not a tool line item. It moves ARPU on the whole base, every month.

Installed base

4.0M

customers at $19 ARPU / month

Addressable revenue

$501.6M

per year, reachable base

Verified ARPU uplift

+17% to +35% ARPU

on treated cohorts, against holdout

What that is worth

$85.3M – $175.6M

incremental revenue per year

Same headcount, same models. The only variable changed is how many decisions those models drive and how many of them are verified.

Run it on your own numbers

What your stack does today.

01

Your data science team

The people who understand your customers quantitatively: features, propensity and uplift models, causal design, the economics behind every number.

02

Markin, decision + execution

The system that consumes those models at machine cadence: generating and sizing opportunities, choosing one action per customer, and proving it against a holdout.

Side by side

The differences that change outcomes.

DimensionYour data science teamMarkin, decision + execution
Question it answersWhat is likely to happen, and what actually caused it?For this customer, right now, what is the highest-value action?
Primary inputWarehouse data, event streams, domain knowledgeYour models, constraints, economics and outcome history
Primary outputModels, scores, experiment designs, readoutsOne sized, ranked decision per customer, with a control group
Usual ownerData science / analyticsSteered by data science, run continuously
How it's measuredModel quality, validity of the causal readIncremental ARPU against a randomised holdout

The unsolved part

The bottleneck was never modelling. It is what happens to the score.

A strong data science team can model anything you put in front of it. What it cannot do is spend a customer-level score at customer level across millions of people, every week, with a control group on each decision. So excellent models end their life in a monthly segment export, and the causal work is reserved for the flagship programmes. That is a throughput gap, not a competence gap.

  • Customer-level scores get collapsed into segments because that is the only unit downstream tooling can act on.
  • Sizing every candidate opportunity by hand costs more analyst time than most opportunities are worth, so prioritisation defaults to intuition.
  • Holdouts are wired manually, so most of what ships is never attributed to anything.
  • Senior analysts spend the week pulling data, building lists and reconciling reports instead of on causal judgement.
  • Model retraining is scheduled, not driven by what the last decisions actually proved.

The actual difference

Markin is not another decisioning engine.

Markin is not a decisioning engine. A decisioning engine ranks actions a human already defined. Markin works like a data science and growth team: it forms its own hypotheses about why ARPU is stuck, marketing, product, pricing or technical, sizes them, executes them inside the systems you already run, and reads each one against a holdout.

 A decisioning engineMarkin
Where the hypothesis comes fromA human authors it. The engine chooses between options someone already approved.Markin authors it. It reads the base, finds where revenue is leaking or unclaimed, and writes the hypothesis itself.
What it is allowed to questionMessage, offer, channel, timing, inside the campaign surface it was given.Anything that moves ARPU: onboarding friction, pricing and packaging, a feature nobody adopts, a payment failure spike, a broken deeplink.
Who does the analysisYour analysts, before and after. The engine optimises; it does not investigate.Markin does the analysis. Sizing, segment definition, experiment design and readout are automated end to end.
Where it stopsAt the recommendation. Someone still has to build and launch it.It launches. Markin executes inside your existing platforms and product surfaces, then closes the loop on the result.
ThroughputAs many hypotheses as your roadmap has room for, typically a handful per quarter.Hundreds in parallel, every one carrying a control group.
What happens when it is wrongThe programme keeps running until someone reviews it.It is retired automatically. Failing to beat control is a normal, cheap outcome.

A decisioning engine picks the best action from a list you wrote. Markin writes the list, and runs it in your stack.

Hypothesis space

Everything a human growth scientist would look at.

Most growth problems are not message problems. Markin is not restricted to the campaign surface: if something is holding ARPU back, it is in scope, and it gets tested the same way.

Marketing

The classic surface, but chosen per customer rather than per segment, and always against a holdout.

  • Which offer this specific customer is worth making
  • Channel and timing chosen per person, not per campaign
  • Contact pressure and fatigue arbitrated across every programme
  • Win-back economics: who is worth a discount and who is not

Product

Where the customer actually experiences the value, and where most silent revenue loss happens.

  • Onboarding steps that lose customers before first value
  • A feature with high retention correlation that half the base never discovers
  • Paywall and upgrade prompt placement
  • In-product surfaces used as a treatment arm, not just email and push

Commercial

Pricing, packaging and the shape of the offer itself, tested rather than argued about.

  • Plan and bundle structure by cohort
  • Discount depth against margin, not against conversion alone
  • Annual versus monthly framing per customer
  • Dunning and involuntary churn recovery sequences

Technical health

Anomalies nobody asked it to look for. This is the category no decisioning engine covers.

  • A checkout error rate that rose on one device and one region
  • Payment failures concentrated in a single issuer or method
  • A broken deeplink quietly killing a high-value journey
  • Latency or delivery degradation eating conversion before any message does

Think of Markin as a data science and growth team that never sleeps: it investigates, forms hypotheses, ships them into your own stack and proves each one against a control group, at a volume no human team can reach.

Architecture

How data science and the decision layer run together

Step 01

The team sets the frame

Features, models worth trusting, margin and contact economics, eligibility rules, and how a causal read must be designed to count. This stays with the people who know your business.

Step 02

Markin runs the volume

Inside that frame, Markin generates candidate actions, sizes the revenue behind each, ranks them per customer, chooses one, and holds a randomised control group back, continuously, across the whole base, without a ticket.

Step 03

Evidence comes back to the team

Every decision returns a measured result: what beat control, by how much, for whom. Data science reads the evidence, retrains on it, adjusts the frame and audits the system. The loop compounds instead of resetting each quarter.

The last step is execution, not a hand-off. Markin does not email a recommendation to someone who then has to build it: it launches the treatment inside your data science team and your product surfaces directly, with the holdout attached, and reads the result itself.

The loop

Execution is a step in the loop, not a hand-off.

  1. 01

    Observe

    Markin reads the behavioural, transactional and product signal you already collect, continuously.

  2. 02

    Hypothesise

    It writes the hypothesis itself, marketing, product, commercial or technical, and states the expected direction.

  3. 03

    Size

    Each opportunity is ranked by expected value, so the queue is ordered by money rather than by opinion.

  4. 04

    Design

    Segment, treatment, guardrails and a randomised holdout are set before anything ships.

  5. 05

    Execute

    It launches inside the systems you already run, your engagement platform, your product surfaces, your APIs. Nothing waits on a build queue.

  6. 06

    Read

    Results are measured against the holdout over a full window, so novelty is not mistaken for effect.

  7. 07

    Scale or retire

    What beats control is scaled across the base. What does not is switched off automatically.

Where Markin fits

Not a replacement. A growth-science team on top.

The split is simple: your team owns the science, Markin owns the arithmetic repeated millions of times. Modelling, economics and causal design stay human. Sizing, ranking, choosing, holding out and measuring stop being a roadmap item and become a system that runs while the team sleeps.

Your models finally get used at customer level

A propensity or uplift model built by your team feeds a per-customer decision that also knows margin, contact cost and eligibility. The score stops ending its life in a segment export.

Causal design becomes the default, not the exception

A randomised holdout on every decision means the model you shipped has a measured incremental effect, not a correlation you have to defend in a readout.

Analysts become owners of a decision system

Instead of servicing campaign requests, data science defines and audits the system that makes millions of decisions, a far larger surface of influence with the same headcount.

Retraining driven by outcomes, not by calendar

Each decision returns a labelled result under a known treatment assignment, which is the cleanest training data your models will ever get.

Operating model

The constraint is not ideas. It is how many you can test.

 Today, with the team aloneWith Markin on top
Where a model score is spentA monthly segment exportA live input to millions of decisions a week
Unit of decisionSegment, chosen in a planning meetingOne customer, chosen from expected value
Share of decisions with a control groupThe flagship programmes, when there is timeEvery decision, by default
Where senior analyst time goesPulling data, building lists, reconciling reportsEconomics, constraints, causal design, what to scale
Retraining signalScheduled, on observational dataContinuous, on randomised outcomes
Cost of testing the 500th hypothesisAnother analyst, another quarterEffectively zero

Markin does not replace your data science team. It removes the ceiling on how much of the base that team can act on, and how fast it finds out whether it worked.

What Markin does not replace.

To be explicit, because this is the question every analytics leader asks first:

  • Markin does not replace data scientists. Feature engineering, domain models and causal judgement stay with the people who own them.
  • Markin does not hide its reasoning. Every decision is auditable back to the inputs and the model outputs that produced it.
  • Markin does not force you off your models. Bring your own propensity and uplift models, or override any model in the loop.
  • Markin does not own your warehouse. It reads from where your data already lives; no migration, no new source of truth.
  • Markin is not a reporting tool. It produces decisions and the evidence that they worked; your analytics stack stays where it is.
  • Markin does not sit beside your data science team making suggestions. It drives it, the action is launched there, in the system your team already knows, and the result comes back into the loop.

Evidence standard

Most of this category reports its own lift.

None of the major engagement, CDP or personalisation vendors publishes an independently verified uplift figure for its decisioning product. Where numbers exist, they come from vendor-commissioned studies or single-customer case studies with no disclosed holdout methodology. The most rigorous public research in the category is not flattering to anyone, including us, which is exactly why we build against it.

How Markin holds itself to it

  • Every decision Markin makes carries a control group. Uplift is reported against that holdout, not against the customers who did not qualify.
  • Results are read over a full measurement window rather than in the first weeks, so novelty is not mistaken for effect.
  • Programmes that fail to beat control are retired automatically. Killing decisions that do not pay is part of the loop, not an annual review.
  • The one figure we quote about ourselves is a range, not an average: +17% to +35% ARPU on treated cohorts against a randomised holdout, across Markin deployments in large B2C bases. We publish no industry benchmark, because we could not source one we would be willing to defend. Your holdout is the number that matters.

Size it yourself

Size what your models could produce with the throughput removed

Preloaded for a large B2C business with a capable in-house data science function: good models, a full roadmap, and far more hypotheses than quarters to test them in. The figure below is not what a new team would produce. It is the incremental margin available from spending the scores you already build at customer level, continuously, measured against a randomised holdout.

Your base

4.0M

Accounts that generated revenue in the last 30 days. Not registered users.

$19

Recurring plus non-recurring revenue divided by active customers.

60%

Margin on the next unit sold, not blended company margin.

Your programme today

55%

Consented, non-fatigued, reachable on at least one channel.

2.4%

Revenue lost to cancellations each month, as a share of the base.

The bet

$1.2M

Licences, data, incentives and the people running it.

3%

Before any incrementality haircut. 2–4% is a defensible planning assumption.

Verified annual impact

$5.1M

Net incremental gross margin in the central case, after the programme cost and after the share of decisioning programmes that independent research finds deliver no real lift.

Reported uplift

$15.0M

What a before/after dashboard would claim, with no control group.

Verified uplift

$10.5M

What survives a holdout in the central case.

Return on programme cost

5.3×

Payback

3 mo

If 20–40% of it does nothing

Best case · 20% no lift$6.0M
Central case · 30% no lift$5.1M
Worst case · 40% no lift$4.2M

What it takes to prove it

To detect a 3% lift on revenue per customer you need roughly 40K customers in the control arm, about 1.8% of your addressable base, read over at least 8 weeks, so novelty is not mistaken for effect.

Addressable base

2.2M

Revenue at risk from churn

$230.6M

Annualised, at the current monthly rate.

Open the full calculator, with the method behind it

Time to value

90 days to a number that survived a holdout.

No replatform, no data migration, no rebuild of the channels you already run. If the first cohorts do not beat control, nothing scales and you have lost a quarter, not a roadmap.

  1. Weeks 0–2

    The team encodes the frame

    Models worth trusting, margin and contact economics, eligibility rules and measurement standards. Markin connects to the warehouse and channels you already run, no migration, no change of ownership.

  2. Weeks 3–6

    Scores start being spent per customer

    Hypotheses the team never had time to reach are generated, sized and put in test in parallel, each with a randomised holdout. Data science reviews the designs and the assignment.

  3. Weeks 7–12

    First verified incremental revenue, attributable to your models

    Results are read over a full measurement window. What beats control scales; what does not is retired. The number the team presents survived a holdout, and the system that produced it is theirs to steer.

When you don’t need Markin.

  • You have no outcome history and no channel to act in, so there is nothing to learn from or decide about yet.
  • The organisation is not willing to hold out a control group, in which case nothing here can be verified.
  • Your base is small enough that a person can reasonably reason about every customer segment in a meeting.

Questions buyers ask.

Does Markin replace my data science team?

No, and a team that tried to run Markin without data scientists would get less out of it. Markin removes the manual decisioning work: sizing candidates by hand, exporting segments, wiring holdouts, reconciling readouts. The scarce skill, knowing what to optimise, what to constrain and what a causal read actually proves, becomes more valuable, not less.

Who owns the models?

You do. Markin uses the features, propensity and uplift models your team already maintains, alongside its own opportunity sizing, and every decision is auditable back to the inputs that produced it. Your team can inspect, override or replace any model in the loop.

Can we bring our own uplift models?

Yes. Bring scores from the warehouse or serve them live; Markin combines them with margin, contact cost and eligibility to choose one action per customer. Where you have no model, Markin's own sizing fills the gap until you do.

How is this different from deploying our models to a campaign tool?

A campaign tool consumes a score to build a list. A decision layer compares every candidate action for a customer on expected value, applies constraints, chooses one or holds, and attaches a control group. The output is a decision with evidence, not an audience.

How do we know the lift came from the decision layer and not the models?

You do not have to separate them, and the holdout does not care: treated and held-out cohorts differ only in whether a per-customer decision was made, in the same period, with the same models and the same seasonality. That is why the verified range we quote is post-holdout rather than pre/post.

How is Markin different from the decisioning or AI already inside your data science team?

A decisioning engine ranks actions a human already defined, inside the campaign surface it was given. Markin forms the hypotheses itself, marketing, product, pricing or a technical anomaly holding growth back, sizes them, executes them inside your data science team and your product surfaces, and reads each one against a randomised holdout. It behaves like a data science and growth team, not like an optimiser.

Does Markin only test messages and offers?

No. Anything a human growth scientist would investigate is in scope: onboarding friction, feature adoption, pricing and packaging, dunning, and technical health issues such as a checkout error rate or a broken deeplink quietly killing conversion. Marketing is one of four hypothesis domains, not the boundary.

What is the business case for adding Markin on top of your data science team?

On the assumptions preloaded above, 4.0M customers at 19 a month, a small move in ARPU is a large number in absolute terms, because it applies to the whole installed base every month rather than to a campaign. Across Markin deployments the verified range on treated cohorts is +17% to +35% ARPU against a randomised holdout. The point is not more messages: it is finding the highest-value action per customer, launching it, and proving it against control before it scales.

How long before it pays for itself?

First sized opportunities are in test within six weeks and the first holdout-verified result lands inside 90 days. Payback depends on your base, margin and programme cost, the calculator on this page computes it from your own numbers, after applying the 20% to 40% haircut BCG finds when next-best-action programmes are incrementality-tested.