The Markin ARPU report for B2C enterprisesRead now
MARKIN

COMPARE/Markin and your stack

Markin + Optimizely: from winning experiments to deciding what is worth testing

+17–35% ARPU against holdoutObserved range across Markin deployments, measured on treated cohorts.

Optimizely is an experimentation and feature-flagging platform: it runs A/B tests and bandits on variations you define and reports which wins on a metric you chose. Markin sits above it and decides which commercial opportunity deserves a test in the first place, sizes the revenue behind it, and acts continuously on customers who never enter an experiment. The learning happens in Optimizely; the decision of what to learn and who to act on happens in Markin.

What is at stake

A decision layer is not a tool line item. It moves ARPU on the whole base, every month.

Installed base

3.0M

customers at $22 ARPU / month

Addressable revenue

$396.0M

per year, reachable base

Verified ARPU uplift

+17% to +35% ARPU

on treated cohorts, against holdout

What that is worth

$67.3M – $138.6M

incremental revenue per year

Measured on treated cohorts against a randomised holdout, read over a full measurement window rather than the first weeks. Anonymised range across Markin deployments in large B2C bases; your own holdout is the number that decides. The figures above apply that range to the reachable share of the base on this page's assumptions; they are arithmetic, not a forecast for your business.

Run it on your own numbers

What your stack does today.

01

Optimizely

An experimentation and feature-flagging platform that deploys code behind flags, runs A/B/n tests and multi-armed bandits, and reports which variation performs on a defined success metric with valid statistics.

02

Markin, the decision + execution layer

A layer that discovers and sizes revenue opportunities per customer, selects the treatment with the highest expected incremental value, and assigns a control group, deciding what is worth testing and who to act on.

Side by side

The differences that change outcomes.

DimensionOptimizelyMarkin, the decision + execution layer
Question it answersWhich variation performs better on this metric, and how confidently?Which opportunity deserves an experiment, and which action should this customer receive now?
Primary inputVariations, a success metric, traffic allocation, audience targeting.Customer context, outcomes, margins, costs, contact history, past experiment results.
Primary outputA statistical read per variation: lift, confidence interval, significance.A ranked, sized decision per customer, including hold, that can feed an Optimizely flag or act directly in other channels.
Usual ownerProduct, engineering and experimentation.Growth, data science and revenue leadership.
How it's measuredLift on the chosen metric with significance and confidence interval.Incremental revenue and ARPU against a holdout.

The unsolved part

What stays unsolved when experimentation is running well

A mature Optimizely setup tests whatever the team writes into it, at the cadence the team can ship. Experimentation throughput is bounded by the hypotheses a human decides to formalise, and the customers a flag is wired into, rather than by what the data is signalling across the base.

  • An experiment only covers customers who hit the flagged surface; the rest of the base receives no decision at all.
  • Stats Accelerator finds the best variation inside a test; it does not decide whether the test itself is the highest-value use of that contact.
  • Each experiment reports on one metric. Nothing sums the effect of every experiment a customer saw into one revenue number.
  • The backlog of hypotheses arrives at the speed the team can write specs, not at the speed the data changes.

The actual difference

Markin is not another decisioning engine.

Markin is not a decisioning engine. A decisioning engine ranks actions a human already defined. Markin works like a data science and growth team: it forms its own hypotheses about why ARPU is stuck, marketing, product, pricing or technical, sizes them, executes them inside the systems you already run, and reads each one against a holdout.

 A decisioning engineMarkin
Where the hypothesis comes fromA human authors it. The engine chooses between options someone already approved.Markin authors it. It reads the base, finds where revenue is leaking or unclaimed, and writes the hypothesis itself.
What it is allowed to questionMessage, offer, channel, timing, inside the campaign surface it was given.Anything that moves ARPU: onboarding friction, pricing and packaging, a feature nobody adopts, a payment failure spike, a broken deeplink.
Who does the analysisYour analysts, before and after. The engine optimises; it does not investigate.Markin does the analysis. Sizing, segment definition, experiment design and readout are automated end to end.
Where it stopsAt the recommendation. Someone still has to build and launch it.It launches. Markin executes inside your existing platforms and product surfaces, then closes the loop on the result.
ThroughputAs many hypotheses as your roadmap has room for, typically a handful per quarter.Hundreds in parallel, every one carrying a control group.
What happens when it is wrongThe programme keeps running until someone reviews it.It is retired automatically. Failing to beat control is a normal, cheap outcome.

A decisioning engine picks the best action from a list you wrote. Markin writes the list, and runs it in your stack.

Hypothesis space

Everything a human growth scientist would look at.

Most growth problems are not message problems. Markin is not restricted to the campaign surface: if something is holding ARPU back, it is in scope, and it gets tested the same way.

Marketing

The classic surface, but chosen per customer rather than per segment, and always against a holdout.

  • Which offer this specific customer is worth making
  • Channel and timing chosen per person, not per campaign
  • Contact pressure and fatigue arbitrated across every programme
  • Win-back economics: who is worth a discount and who is not

Product

Where the customer actually experiences the value, and where most silent revenue loss happens.

  • Onboarding steps that lose customers before first value
  • A feature with high retention correlation that half the base never discovers
  • Paywall and upgrade prompt placement
  • In-product surfaces used as a treatment arm, not just email and push

Commercial

Pricing, packaging and the shape of the offer itself, tested rather than argued about.

  • Plan and bundle structure by cohort
  • Discount depth against margin, not against conversion alone
  • Annual versus monthly framing per customer
  • Dunning and involuntary churn recovery sequences

Technical health

Anomalies nobody asked it to look for. This is the category no decisioning engine covers.

  • A checkout error rate that rose on one device and one region
  • Payment failures concentrated in a single issuer or method
  • A broken deeplink quietly killing a high-value journey
  • Latency or delivery degradation eating conversion before any message does

Think of Markin as a data science and growth team that never sleeps: it investigates, forms hypotheses, ships them into your own stack and proves each one against a control group, at a volume no human team can reach.

Their decisioning layer

What Optimizely decides — and where it stops.

Optimizely is an experimentation and feature-flagging platform. It deploys code behind flags, runs A/B tests and bandits on variations you define, and proves which variation wins on a success metric you chose. It does not decide which commercial opportunity is worth testing, size the revenue behind it, or act on customers who never enter an experiment.

Products referenced: Optimizely Feature Experimentation, Optimizely Web Experimentation, Optimizely Personalization, Stats Engine, Stats Accelerator

What it optimises

Documented boundaries

What the evidence actually says

Where Markin is different.

Decides what to test, not which variation wins

Optimizely answers 'which variation performs better on this metric'. Markin answers 'which opportunity is worth running an experiment on at all', and surfaces the ones no one thought to test.

Optimises revenue across the base, not a metric on one test

Stats Accelerator reallocates traffic inside an experiment. Markin reallocates contact and treatment across the whole customer estate, on expected incremental revenue net of margin and cost.

Holds across the estate, by default

An Optimizely holdout is a traffic split within one experiment. Markin attaches a control group to every decision, so the reported number is incremental ARPU against holdout rather than a significance read on a flag.

Architecture

How the two run together

Step 01

Context in

Markin reads customer context where it already lives, warehouse, CDP, product and billing systems, plus outcome history. Optimizely experiment results and flag exposure can be part of that context, so learning feeds the next decision.

Step 02

Decision

Markin generates and sizes revenue opportunities, ranks them per customer, chooses a treatment and assigns a control group. Some decisions become Optimizely experiments; most become direct actions in other channels.

Step 03

Activation back into Optimizely

Where a hypothesis needs a controlled test, the decision is written back as flag targeting or audience criteria so Optimizely runs it with its own stats engine. Flag logic, variation code and significance reporting remain in Optimizely.

The last step is execution, not a hand-off. Markin does not email a recommendation to someone who then has to build it: it launches the treatment inside Optimizely and your product surfaces directly, with the holdout attached, and reads the result itself.

The loop

Execution is a step in the loop, not a hand-off.

  1. 01

    Observe

    Markin reads the behavioural, transactional and product signal you already collect, continuously.

  2. 02

    Hypothesise

    It writes the hypothesis itself, marketing, product, commercial or technical, and states the expected direction.

  3. 03

    Size

    Each opportunity is ranked by expected value, so the queue is ordered by money rather than by opinion.

  4. 04

    Design

    Segment, treatment, guardrails and a randomised holdout are set before anything ships.

  5. 05

    Execute

    It launches inside the systems you already run, your engagement platform, your product surfaces, your APIs. Nothing waits on a build queue.

  6. 06

    Read

    Results are measured against the holdout over a full window, so novelty is not mistaken for effect.

  7. 07

    Scale or retire

    What beats control is scaled across the base. What does not is switched off automatically.

Where Markin fits

Not a replacement. A growth-science team on top.

Markin does not run experiments or replace flags. It decides which revenue opportunity is worth a test, and acts on the customers a flag never reaches. Experimentation keeps its stats engine, its bandits and its release safety; what changes is the input that decides what gets tested and who gets acted on.

The experiment queue becomes a decision, not a backlog

Instead of a product roadmap of tests, the highest-value hypotheses are surfaced and sized from the data, so experimentation capacity is spent where the revenue is.

Coverage extends beyond the flagged surface

Customers who never enter an experiment still receive a decision, a save, an attach, a hold, measured against control in the channels they do use.

Learning and acting share one record

Optimizely results feed Markin, and Markin's holdout feeds back, so the next round of hypotheses improves on proven numbers rather than drifting.

Operating model

The constraint is not ideas. It is how many you can test.

 Today, with OptimizelyWith Markin on top
Revenue hypotheses tested per quarter4 to 8, whatever the roadmap had room forHundreds, generated and run in parallel
What can be hypothesised aboutMessages, offers and audiences, the campaign surfaceMarketing, product, pricing and technical health alike
From decision to live in the channelA ticket, a build queue, a release windowMarkin launches it in your existing platforms itself
Time from idea to a result you trust6 to 10 weeks of analysis, build and readoutDays, because sizing and design are automated
Share of decisions with a control groupThe flagship programmes, when there is timeEvery decision, by default
Coverage of the baseTop segments and the customers a rule caughtOne decision per customer, across the whole base
Cost of testing the 500th hypothesisAnother analyst, another quarterEffectively zero
What the team spends its time onPulling data, building lists, reconciling reportsJudgement: constraints, economics, what to scale

Markin does not replace your data science team. It removes the ceiling on how much of the base that team can act on, and how fast it finds out whether it worked.

What Markin does not replace.

To be explicit about scope, because procurement will ask:

  • Markin does not run A/B tests, feature flags or bandits. Experimentation stays in Optimizely.
  • Markin does not replace Stats Engine, Stats Accelerator or Optimizely's statistical reporting.
  • Markin is not an analytics or release-management tool; it decides what is worth testing and who to act on.
  • Markin does not manage variation code, rollouts or flag environments.
  • Markin does not sit beside Optimizely making suggestions. It drives it, the action is launched there, in the system your team already knows, and the result comes back into the loop.

Evidence standard

Most of this category reports its own lift.

None of the major engagement, CDP or personalisation vendors publishes an independently verified uplift figure for its decisioning product. Where numbers exist, they come from vendor-commissioned studies or single-customer case studies with no disclosed holdout methodology. The most rigorous public research in the category is not flattering to anyone, including us, which is exactly why we build against it.

How Markin holds itself to it

  • Every decision Markin makes carries a control group. Uplift is reported against that holdout, not against the customers who did not qualify.
  • Results are read over a full measurement window rather than in the first weeks, so novelty is not mistaken for effect.
  • Programmes that fail to beat control are retired automatically. Killing decisions that do not pay is part of the loop, not an annual review.
  • The one figure we quote about ourselves is a range, not an average: +17% to +35% ARPU on treated cohorts against a randomised holdout, across Markin deployments in large B2C bases. We publish no industry benchmark, because we could not source one we would be willing to defend. Your holdout is the number that matters.

Size it yourself

Size the decision layer above your Optimizely programme

Preloaded for a consumer subscription business that runs Optimizely at scale: mature feature-flagging and A/B testing, a product team that ships many experiments. Optimizely proves which variation wins on the metric you chose. The figure below is the incremental margin available from deciding what is worth testing in the first place, and from acting on every customer continuously rather than only where a flag is wired in, measured against a holdout rather than a per-experiment significance read.

Your base

3.0M

Accounts that generated revenue in the last 30 days. Not registered users.

$22

Recurring plus non-recurring revenue divided by active customers.

60%

Margin on the next unit sold, not blended company margin.

Your programme today

50%

Consented, non-fatigued, reachable on at least one channel.

2.6%

Revenue lost to cancellations each month, as a share of the base.

The bet

$1.0M

Licences, data, incentives and the people running it.

3%

Before any incrementality haircut. 2–4% is a defensible planning assumption.

Verified annual impact

$4.0M

Net incremental gross margin in the central case, after the programme cost and after the share of decisioning programmes that independent research finds deliver no real lift.

Reported uplift

$11.9M

What a before/after dashboard would claim, with no control group.

Verified uplift

$8.3M

What survives a holdout in the central case.

Return on programme cost

5.0×

Payback

3 mo

If 20–40% of it does nothing

Best case · 20% no lift$4.7M
Central case · 30% no lift$4.0M
Worst case · 40% no lift$3.3M

What it takes to prove it

To detect a 3% lift on revenue per customer you need roughly 40K customers in the control arm, about 2.7% of your addressable base, read over at least 8 weeks, so novelty is not mistaken for effect.

Addressable base

1.5M

Revenue at risk from churn

$214.7M

Annualised, at the current monthly rate.

Open the full calculator, with the method behind it

Time to value

90 days to a number that survived a holdout.

No replatform, no data migration, no rebuild of the channels you already run. If the first cohorts do not beat control, nothing scales and you have lost a quarter, not a roadmap.

  1. Weeks 0–2

    Read the context you already have

    Markin connects to the data and the channels you run today, Optimizely included. No migration, no replatform, no new source of truth.

  2. Weeks 3–6

    First sized opportunities in test

    Opportunities are ranked by expected value, treatments are chosen per customer, and the first cohorts go live with a randomised holdout attached.

  3. Weeks 7–12

    First verified incremental revenue

    Results are read over a full measurement window. What beats control scales; what does not is retired. Nothing scales on a number that has not survived a holdout.

When you don’t need Markin.

  • You run a handful of experiments a quarter and the team can still reason about which to prioritise in a meeting.
  • Every customer decision already flows through a flagged surface, so there is no unaddressed base to act on.
  • Your success metric cannot be tied to revenue, so neither experiments nor decisions can be valued.

Questions buyers ask.

Doesn't Optimizely already decide things with bandits?

Stats Accelerator and multi-armed bandits reallocate traffic between variations inside one experiment to reach significance faster or maximise reward during the test. They optimise the test you already wrote; they do not decide which opportunity is worth testing, size the revenue behind it, or act on customers who never enter the experiment.

Do we have to replace Optimizely?

No. Optimizely remains the experimentation and feature-flagging layer. Markin decides which hypothesis deserves a test and passes that decision into Optimizely as flag targeting, while acting on the rest of the base directly in other channels.

How does the decision reach Optimizely?

As flag targeting rules or audience criteria on a flag, so an existing experiment picks it up. The integration surface is the same one your team already uses for any other upstream signal.

Isn't running experiments all the time the same as decisioning?

No. Experimentation tests a few hypotheses at human cadence on the surfaces you flagged. Continuous decisioning acts on every customer, every cycle, with a control group by default, including customers no experiment reaches. Experiments prove what works; decisioning decides who gets it and who gets held.

Who owns the output, product or growth?

Both. Product and engineering own the experiments and the flags; growth and data science own the revenue decision. The decision layer is the shared contract between them.

What does the first ninety days look like?

One revenue theme, one channel, a real holdout. The point of the first quarter is a defensible incremental number, not full coverage of the experiment backlog.

How is Markin different from the decisioning or AI already inside Optimizely?

A decisioning engine ranks actions a human already defined, inside the campaign surface it was given. Markin forms the hypotheses itself, marketing, product, pricing or a technical anomaly holding growth back, sizes them, executes them inside Optimizely and your product surfaces, and reads each one against a randomised holdout. It behaves like a data science and growth team, not like an optimiser.

Does Markin only test messages and offers?

No. Anything a human growth scientist would investigate is in scope: onboarding friction, feature adoption, pricing and packaging, dunning, and technical health issues such as a checkout error rate or a broken deeplink quietly killing conversion. Marketing is one of four hypothesis domains, not the boundary.

What is the business case for adding Markin on top of Optimizely?

On the assumptions preloaded above, 3.0M customers at 22 a month, a small move in ARPU is a large number in absolute terms, because it applies to the whole installed base every month rather than to a campaign. Across Markin deployments the verified range on treated cohorts is +17% to +35% ARPU against a randomised holdout. The point is not more messages: it is finding the highest-value action per customer, launching it, and proving it against control before it scales.

How long before it pays for itself?

First sized opportunities are in test within six weeks and the first holdout-verified result lands inside 90 days. Payback depends on your base, margin and programme cost, the calculator on this page computes it from your own numbers, after applying the 20% to 40% haircut BCG finds when next-best-action programmes are incrementality-tested.