COMPARE/Markin and your stack
Markin + Amplitude: from behavioural insight to revenue decisions
+17–35% ARPU against holdoutObserved range across Markin deployments, measured on treated cohorts.
Amplitude is product analytics with experimentation and recommendations: it observes behaviour, builds cohorts, runs feature experiments and can recommend content toward a predictive goal. Markin sits above it and converts that observation into sized revenue decisions per customer, which action, in which channel, at what cost, with a holdout attached. Amplitude shows what is happening; Markin decides what to do about it, in revenue terms.
What is at stake
A decision layer is not a tool line item. It moves ARPU on the whole base, every month.
Installed base
2.5M
customers at $16 ARPU / month
Addressable revenue
$192.0M
per year, reachable base
Verified ARPU uplift
+17% to +35% ARPU
on treated cohorts, against holdout
What that is worth
$32.6M – $67.2M
incremental revenue per year
Measured on treated cohorts against a randomised holdout, read over a full measurement window rather than the first weeks. Anonymised range across Markin deployments in large B2C bases; your own holdout is the number that decides. The figures above apply that range to the reachable share of the base on this page's assumptions; they are arithmetic, not a forecast for your business.
Run it on your own numbersWhat your stack does today.
01
Amplitude
A product analytics platform with feature experimentation and content recommendations: behavioural insight, cohorts, A/B tests and AutoML item recommendations toward a predictive goal.
02
Markin, the decision + execution layer
A layer that turns behavioural signal into sized revenue decisions per customer, selecting the treatment with the highest expected incremental value and assigning a control group.
Side by side
The differences that change outcomes.
| Dimension | Amplitude | Markin, the decision + execution layer |
|---|---|---|
| Question it answers | What is happening in the product, and which feature variation or item performs? | Given what we see, which revenue action is worth taking for this customer, and what is it worth? |
| Primary input | Instrumented events, user properties, experiment variants, predictive goals. | Behaviour and cohorts from analytics, outcomes, margins, costs, contact history, constraints. |
| Primary output | Behavioural analysis, cohorts, experiment significance, ranked recommendations. | A ranked, sized decision per customer, including hold, pushed into the channels that can act on it. |
| Usual owner | Product, analytics and experimentation. | Growth, data science and revenue leadership. |
| How it's measured | Significance on the experiment metric; recommendation take rate and relevance. | Incremental revenue and ARPU against a holdout. |
The unsolved part
What stays unsolved when analytics is running well
A mature Amplitude setup shows exactly what is happening and tests features well. Insight does not decide itself: turning 'users who do X churn more' into 'who to save, with what, at what cost, and who to leave alone' is a separate job, and it is normally done in a planning meeting and a spreadsheet.
- A cohort describes a population; it does not attach a revenue value, a margin or a cost per intervention to each customer in it.
- An experiment proves a feature works on a metric; it does not decide who should receive the rolled-out feature, or who should be held.
- Recommendations maximise a predicted engagement goal, not incremental revenue net of contact cost, so the highest-relevance item is not always the highest-value action.
- Holdout groups measure the program's combined lift, not the incremental revenue of each commercial decision as it is made.
The actual difference
Markin is not another decisioning engine.
Markin is not a decisioning engine. A decisioning engine ranks actions a human already defined. Markin works like a data science and growth team: it forms its own hypotheses about why ARPU is stuck, marketing, product, pricing or technical, sizes them, executes them inside the systems you already run, and reads each one against a holdout.
| A decisioning engine | Markin | |
|---|---|---|
| Where the hypothesis comes from | A human authors it. The engine chooses between options someone already approved. | Markin authors it. It reads the base, finds where revenue is leaking or unclaimed, and writes the hypothesis itself. |
| What it is allowed to question | Message, offer, channel, timing, inside the campaign surface it was given. | Anything that moves ARPU: onboarding friction, pricing and packaging, a feature nobody adopts, a payment failure spike, a broken deeplink. |
| Who does the analysis | Your analysts, before and after. The engine optimises; it does not investigate. | Markin does the analysis. Sizing, segment definition, experiment design and readout are automated end to end. |
| Where it stops | At the recommendation. Someone still has to build and launch it. | It launches. Markin executes inside your existing platforms and product surfaces, then closes the loop on the result. |
| Throughput | As many hypotheses as your roadmap has room for, typically a handful per quarter. | Hundreds in parallel, every one carrying a control group. |
| What happens when it is wrong | The programme keeps running until someone reviews it. | It is retired automatically. Failing to beat control is a normal, cheap outcome. |
A decisioning engine picks the best action from a list you wrote. Markin writes the list, and runs it in your stack.
Hypothesis space
Everything a human growth scientist would look at.
Most growth problems are not message problems. Markin is not restricted to the campaign surface: if something is holding ARPU back, it is in scope, and it gets tested the same way.
Marketing
The classic surface, but chosen per customer rather than per segment, and always against a holdout.
- Which offer this specific customer is worth making
- Channel and timing chosen per person, not per campaign
- Contact pressure and fatigue arbitrated across every programme
- Win-back economics: who is worth a discount and who is not
Product
Where the customer actually experiences the value, and where most silent revenue loss happens.
- Onboarding steps that lose customers before first value
- A feature with high retention correlation that half the base never discovers
- Paywall and upgrade prompt placement
- In-product surfaces used as a treatment arm, not just email and push
Commercial
Pricing, packaging and the shape of the offer itself, tested rather than argued about.
- Plan and bundle structure by cohort
- Discount depth against margin, not against conversion alone
- Annual versus monthly framing per customer
- Dunning and involuntary churn recovery sequences
Technical health
Anomalies nobody asked it to look for. This is the category no decisioning engine covers.
- A checkout error rate that rose on one device and one region
- Payment failures concentrated in a single issuer or method
- A broken deeplink quietly killing a high-value journey
- Latency or delivery degradation eating conversion before any message does
Think of Markin as a data science and growth team that never sleeps: it investigates, forms hypotheses, ships them into your own stack and proves each one against a control group, at a volume no human team can reach.
Their decisioning layer
What Amplitude decides — and where it stops.
Amplitude is product analytics with an experimentation and recommendations layer. It observes behaviour, builds cohorts, runs feature experiments and can recommend content toward a predictive goal. It does not size the revenue opportunity behind the behaviour or decide cross-channel actions in margin and cost terms.
Products referenced: Amplitude Analytics, Amplitude Experiment, Recommendations (Activation), Personalization, Holdout groups
What it optimises
Amplitude is a product analytics platform that turns instrumented events into behavioural insight, funnels, retention, cohorts and user journeys.
Vendor docsAmplitude docs, Recommendations (Activation)Amplitude Experiment runs feature experiments using a sequential testing method that keeps results valid whenever they are viewed and ends experiments early, on average, with fewer observations.
Vendor docsAmplitude docs, Sequential testing for statistical inferenceRecommendations (Activation) use AutoML to determine which items are most likely to maximise each user's predicted goal, for use in personalisation campaigns.
Vendor docsAmplitude docs, Recommendations: help users reach their goals
Documented boundaries
Amplitude's holdout groups measure the long-term, combined lift of an experimentation program as a whole, not the incremental revenue of a per-customer commercial decision across channels.
Vendor docsAmplitude docs, Holdout groupsThe optimisation unit is an experiment metric or a content recommendation toward a predictive goal. Revenue at stake, margin and contact cost per customer are not part of the recommendation decision.
Vendor docsAmplitude docs, Build a recommendation
What the evidence actually says
Amplitude's documented performance material concerns experiment significance and recommendation relevance. No independent benchmark of Amplitude's impact on incremental revenue is published.
Vendor docsAmplitude docs, Analyze your experiment data with the T-test
Where Markin is different.
Observes and tests; Markin decides the commercial action
Amplitude answers 'what is happening, and which feature variation wins'. Markin answers 'which revenue action is worth taking for this customer, and what is it worth', turning observation into a decision.
From a metric to a revenue decision
A recommendation maximises a predicted goal; a Markin decision maximises expected incremental revenue net of margin and contact cost, which is why hold is a legitimate output.
Per-customer decision across the estate
Amplitude cohorts and experiments group users. Markin produces one sized decision per customer, across every channel, and reports it against a holdout, not an experiment metric.
Architecture
How the two run together
Context in
Markin reads behavioural context where it already lives, Amplitude events and cohorts, warehouse, CDP, billing, plus outcome history. Observation becomes one of the inputs to the decision, not the output.
Decision
Markin generates and sizes revenue opportunities, ranks them per customer, chooses a treatment and assigns a control group. The output is one decision per customer, with an expected value attached.
Activation into channels
The chosen decision is written back as user attributes or triggered events, so engagement and lifecycle channels act on it. Where a feature needs validation, the decision can feed an Amplitude Experiment as targeting.
The last step is execution, not a hand-off. Markin does not email a recommendation to someone who then has to build it: it launches the treatment inside Amplitude and your product surfaces directly, with the holdout attached, and reads the result itself.
The loop
Execution is a step in the loop, not a hand-off.
- 01
Observe
Markin reads the behavioural, transactional and product signal you already collect, continuously.
- 02
Hypothesise
It writes the hypothesis itself, marketing, product, commercial or technical, and states the expected direction.
- 03
Size
Each opportunity is ranked by expected value, so the queue is ordered by money rather than by opinion.
- 04
Design
Segment, treatment, guardrails and a randomised holdout are set before anything ships.
- 05
Execute
It launches inside the systems you already run, your engagement platform, your product surfaces, your APIs. Nothing waits on a build queue.
- 06
Read
Results are measured against the holdout over a full window, so novelty is not mistaken for effect.
- 07
Scale or retire
What beats control is scaled across the base. What does not is switched off automatically.
Where Markin fits
Not a replacement. A growth-science team on top.
Markin does not replace analytics or experiments. It turns what Amplitude observes into per-customer revenue decisions with a holdout attached, and acts in the channels that can move revenue, including customers a feature experiment never reaches.
Insight becomes a decision, not a deck
A behavioural finding is converted into a sized opportunity and a chosen treatment per customer, so the insight leaves the meeting and reaches the customer.
Recommendations are ranked by revenue, not relevance
Treatments are selected on expected incremental revenue net of margin and cost, so the highest-value action wins the contact even when it is not the most relevant item.
Measured against holdout, by default
Every decision carries a control group and reports incremental ARPU, rather than a significance read on a feature metric or a program-level lift.
Operating model
The constraint is not ideas. It is how many you can test.
| Today, with Amplitude | With Markin on top | |
|---|---|---|
| Revenue hypotheses tested per quarter | 4 to 8, whatever the roadmap had room for | Hundreds, generated and run in parallel |
| What can be hypothesised about | Messages, offers and audiences, the campaign surface | Marketing, product, pricing and technical health alike |
| From decision to live in the channel | A ticket, a build queue, a release window | Markin launches it in your existing platforms itself |
| Time from idea to a result you trust | 6 to 10 weeks of analysis, build and readout | Days, because sizing and design are automated |
| Share of decisions with a control group | The flagship programmes, when there is time | Every decision, by default |
| Coverage of the base | Top segments and the customers a rule caught | One decision per customer, across the whole base |
| Cost of testing the 500th hypothesis | Another analyst, another quarter | Effectively zero |
| What the team spends its time on | Pulling data, building lists, reconciling reports | Judgement: constraints, economics, what to scale |
Markin does not replace your data science team. It removes the ceiling on how much of the base that team can act on, and how fast it finds out whether it worked.
What Markin does not replace.
To be explicit about scope, because procurement will ask:
- Markin does not replace Amplitude Analytics, dashboards, funnels or cohort building.
- Markin does not run feature experiments or replace Amplitude Experiment's stats engine or holdout groups.
- Markin is not a content recommendation engine; it decides commercial actions in revenue terms.
- Markin does not manage event instrumentation or data governance.
- Markin does not sit beside Amplitude making suggestions. It drives it, the action is launched there, in the system your team already knows, and the result comes back into the loop.
Evidence standard
Most of this category reports its own lift.
None of the major engagement, CDP or personalisation vendors publishes an independently verified uplift figure for its decisioning product. Where numbers exist, they come from vendor-commissioned studies or single-customer case studies with no disclosed holdout methodology. The most rigorous public research in the category is not flattering to anyone, including us, which is exactly why we build against it.
BCG reports that when organisations adopt rigorous incrementality testing, they typically find 20% to 40% of their active next-best-action programmes deliver marginal to negative lift.
Independent researchBCG, How Measurement Is Evolving in Next-Best Action (2026)The same research flags novelty effects, new programmes show inflated early results, and recommends 8 to 12 weeks before drawing conclusions.
Independent researchBCG, How Measurement Is Evolving in Next-Best Action (2026)Global-holdout, programme-level ROI measurements often overstate impact through halo effects, pull-forward effects and experiment contamination.
Independent researchBCG, How Measurement Is Evolving in Next-Best Action (2026)
How Markin holds itself to it
- Every decision Markin makes carries a control group. Uplift is reported against that holdout, not against the customers who did not qualify.
- Results are read over a full measurement window rather than in the first weeks, so novelty is not mistaken for effect.
- Programmes that fail to beat control are retired automatically. Killing decisions that do not pay is part of the loop, not an annual review.
- The one figure we quote about ourselves is a range, not an average: +17% to +35% ARPU on treated cohorts against a randomised holdout, across Markin deployments in large B2C bases. We publish no industry benchmark, because we could not source one we would be willing to defend. Your holdout is the number that matters.
Size it yourself
Size the decision layer above your Amplitude programme
Preloaded for a product-led consumer business running Amplitude at scale: rich behavioural analytics, a mature experimentation team, recommendations in personalisation. Amplitude observes behaviour and tests features; it does not size the commercial opportunity behind the behaviour or decide cross-channel actions in revenue terms. The figure below is the incremental margin available from turning those observations into per-customer revenue decisions, measured against a holdout, not a per-experiment significance read.
Your base
Accounts that generated revenue in the last 30 days. Not registered users.
Recurring plus non-recurring revenue divided by active customers.
Margin on the next unit sold, not blended company margin.
Your programme today
Consented, non-fatigued, reachable on at least one channel.
Revenue lost to cancellations each month, as a share of the base.
The bet
Licences, data, incentives and the people running it.
Before any incrementality haircut. 2–4% is a defensible planning assumption.
Verified annual impact
$1.4M
Net incremental gross margin in the central case, after the programme cost and after the share of decisioning programmes that independent research finds deliver no real lift.
Reported uplift
$5.8M
What a before/after dashboard would claim, with no control group.
Verified uplift
$4.0M
What survives a holdout in the central case.
Return on programme cost
2.6×
Payback
5 mo
If 20–40% of it does nothing
What it takes to prove it
To detect a 3% lift on revenue per customer you need roughly 40K customers in the control arm, about 4.0% of your addressable base, read over at least 8 weeks, so novelty is not mistaken for effect.
Addressable base
1.0M
Revenue at risk from churn
$147.0M
Annualised, at the current monthly rate.
Time to value
90 days to a number that survived a holdout.
No replatform, no data migration, no rebuild of the channels you already run. If the first cohorts do not beat control, nothing scales and you have lost a quarter, not a roadmap.
Weeks 0–2
Read the context you already have
Markin connects to the data and the channels you run today, Amplitude included. No migration, no replatform, no new source of truth.
Weeks 3–6
First sized opportunities in test
Opportunities are ranked by expected value, treatments are chosen per customer, and the first cohorts go live with a randomised holdout attached.
Weeks 7–12
First verified incremental revenue
Results are read over a full measurement window. What beats control scales; what does not is retired. Nothing scales on a number that has not survived a holdout.
When you don’t need Markin.
- You cannot connect a behavioural signal to a revenue outcome, so neither analytics nor decisions can be valued.
- Your only actions are in-product feature rollouts with no commercial treatment or channel to act in.
- Your base is small enough that the team can reason about every cohort by hand.
Questions buyers ask.
Doesn't Amplitude already have recommendations and experiments?
Yes. Amplitude Experiment runs feature experiments with sequential testing, and Recommendations uses AutoML to suggest items that maximise a predicted goal. Both optimise a metric or an engagement goal inside Amplitude. They do not size the revenue at stake per customer, choose a cross-channel treatment in margin-and-cost terms, or decide who to hold.
Do we have to replace Amplitude?
No. Amplitude remains the analytics and experimentation layer. Markin reads the signals Amplitude produces, turns them into sized revenue decisions, and hands them to the channels that can act on them.
How does the decision leave Amplitude?
As user attributes or triggered events on the profile, so engagement and lifecycle channels can key off them. Where a feature needs validation, the decision can also feed an Amplitude Experiment as targeting.
What is the difference between an Amplitude holdout and a Markin holdout?
An Amplitude holdout group measures the long-term combined lift of your experimentation program as a whole. A Markin holdout is attached to each commercial decision, so the reported number is the incremental revenue of that specific action, not a program-level average.
Who owns the output, product or growth?
Both. Product and analytics own the insight and the experiments; growth and data science own the revenue decision. The decision layer is the shared contract between them.
What does the first ninety days look like?
One revenue theme, one channel, a real holdout. The point of the first quarter is a defensible incremental number, not full coverage of every cohort.
How is Markin different from the decisioning or AI already inside Amplitude?
A decisioning engine ranks actions a human already defined, inside the campaign surface it was given. Markin forms the hypotheses itself, marketing, product, pricing or a technical anomaly holding growth back, sizes them, executes them inside Amplitude and your product surfaces, and reads each one against a randomised holdout. It behaves like a data science and growth team, not like an optimiser.
Does Markin only test messages and offers?
No. Anything a human growth scientist would investigate is in scope: onboarding friction, feature adoption, pricing and packaging, dunning, and technical health issues such as a checkout error rate or a broken deeplink quietly killing conversion. Marketing is one of four hypothesis domains, not the boundary.
What is the business case for adding Markin on top of Amplitude?
On the assumptions preloaded above, 2.5M customers at 16 a month, a small move in ARPU is a large number in absolute terms, because it applies to the whole installed base every month rather than to a campaign. Across Markin deployments the verified range on treated cohorts is +17% to +35% ARPU against a randomised holdout. The point is not more messages: it is finding the highest-value action per customer, launching it, and proving it against control before it scales.
How long before it pays for itself?
First sized opportunities are in test within six weeks and the first holdout-verified result lands inside 90 days. Payback depends on your base, margin and programme cost, the calculator on this page computes it from your own numbers, after applying the 20% to 40% haircut BCG finds when next-best-action programmes are incrementality-tested.