The Markin ARPU report for B2C enterprisesRead now
MARKIN

COMPARE/Markin vs Claude with MCP

MarkinvsClaude logoClaude

vs Claude with MCP: strong reasoning, no control group

+17–35% ARPU against holdoutObserved range across Markin deployments, measured on treated cohorts.

Claude with MCP is the strongest general option for long-context analysis over your warehouse, and MCP is Anthropic's own protocol. Markin is a different layer: versioned growth skills that size hypotheses in money, run randomised holdouts, and route work across cheap models and classical ML so per-customer decisions stay affordable.

The short answerLast updated: August 2026

If we had to pick one assistant for an analyst to reason over messy data with tools, it would be a close call and Claude would be in the final two. That is still analysis. Markin's job starts when someone has to decide what each customer gets this week and prove afterwards that it was worth doing.

01

Claude with MCP

Anthropic's assistant with MCP servers over your warehouse, code and documents. Long context, careful tool use, strong at multi-step reasoning on unfamiliar data.

Choose it when the work is deep, exploratory analysis that a careful analyst would otherwise do by hand.

  • Long-context reasoning
  • MCP is Anthropic's protocol
  • Per-token pricing

02

Markin

An autonomous growth-science team: hypotheses generated, sized, launched into your stack and read against a control group, continuously.

Choose it when the constraint is how many hypotheses get proven, not how well one gets analysed.

  • Continuous loop
  • Holdout on every action
  • Routed model stack

Line by line

The same ten questions, answered for both.

Markin compared with Claude with MCP across ten dimensions
DimensionMarkinClaude with MCP
What it isAn autonomous growth-science team: it investigates why revenue per customer is stuck and acts on what it finds.A frontier assistant with first-class MCP tool use over your systems.
What it decidesWhich commercial opportunity deserves to exist for each customer this week, what it is worth, and when the right answer is to do nothing.What to recommend to the person in the session. The human owns every downstream decision.
Where hypotheses come fromGenerated by Markin from customer, product, pricing and technical-health data, then sized before anyone builds anything.Reasoned out in the session, often very well, but without sizing or memory across sessions.
How a hypothesis is evaluatedSized in money on the eligible population, filtered by statistical power, then killed or kept by a randomised holdout.Argued rather than tested. No eligible population, no expected value, no power threshold.
Scope of actionMarketing, product, pricing and technical-health hypotheses, arbitrated against each other in one queue.Analysis, code and document work across whatever tools are connected.
Which models do the workA routed mix: frontier language models to write and explain, small cheap models for volume classification, and classical ML and deep learning, uplift, survival, time series, embeddings, for the numbers.One frontier model doing every step, including the ones a gradient-boosted tree does better.
What the cost scales withDecisions taken and revenue proven, not tokens burned. Per-customer reasoning is handled by the cheap layers by design.Tokens, and long context is expensive context. Cost grows with depth of analysis, not with revenue.
How the work reaches the customerWritten back into the systems you already run, as attributes, events or API calls. Markin does not add a new customer-facing surface.A human takes the conclusion into the execution system.
How impact is provenA randomised holdout on every decision. The reported number is incremental revenue and ARPU, not attributed conversions.None built in. Whatever the team sets up afterwards.
Where the data sitsReads context where it already lives, warehouse, CDP, product and billing systems. No new system of record.Live reads through MCP servers you host.
Governance and controlEvery action carries its hypothesis, its expected value, its guardrails and its control group, reviewable before launch.Tool scoping and workspace policy. No experiment record.
Time to a verified numberOne revenue theme, one channel, one holdout: a defensible incremental number inside 90 days.An excellent analysis in an afternoon. A verified number, not yet.
Best fitLarge B2C bases where the constraint is how many good hypotheses get tested, not how many messages get sent.Analysts doing deep, one-off investigations.

What is at stake

A decision layer is not a tool line item. It moves ARPU on the whole base, every month.

Installed base

2.0M

customers at $24 ARPU / month

Addressable revenue

$259.2M

per year, reachable base

Verified ARPU uplift

+17% to +35% ARPU

on treated cohorts, against holdout

What that is worth

$44.1M – $90.7M

incremental revenue per year

Measured on treated cohorts against a randomised holdout, read over a full measurement window rather than the first weeks. Anonymised range across Markin deployments in large B2C bases; your own holdout is the number that decides. The figures above apply that range to the reachable share of the base on this page's assumptions; they are arithmetic, not a forecast for your business.

Run it on your own numbers

The unsolved part

Reasoning quality is not the binding constraint

Most growth programmes do not fail because the analysis was shallow. They fail because good analysis arrives once a quarter, is never sized against alternatives, and is validated by the same attribution model that produced the problem.

  • One deep answer per investigation, versus hundreds of sized candidates per quarter.
  • Nothing carries forward: what was tried and failed is not in the context next time.
  • Depth costs context, and context costs money, so the analysis stays occasional.
  • No holdout, so a confident conclusion and a proven one look identical.

Hypothesis space

Everything a human growth scientist would look at.

Most growth problems are not message problems. Markin is not restricted to the campaign surface: if something is holding ARPU back, it is in scope, and it gets tested the same way.

Marketing

The classic surface, but chosen per customer rather than per segment, and always against a holdout.

  • Which offer this specific customer is worth making
  • Channel and timing chosen per person, not per campaign
  • Contact pressure and fatigue arbitrated across every programme
  • Win-back economics: who is worth a discount and who is not

Product

Where the customer actually experiences the value, and where most silent revenue loss happens.

  • Onboarding steps that lose customers before first value
  • A feature with high retention correlation that half the base never discovers
  • Paywall and upgrade prompt placement
  • In-product surfaces used as a treatment arm, not just email and push

Commercial

Pricing, packaging and the shape of the offer itself, tested rather than argued about.

  • Plan and bundle structure by cohort
  • Discount depth against margin, not against conversion alone
  • Annual versus monthly framing per customer
  • Dunning and involuntary churn recovery sequences

Technical health

Anomalies nobody asked it to look for. This is the category no decisioning engine covers.

  • A checkout error rate that rose on one device and one region
  • Payment failures concentrated in a single issuer or method
  • A broken deeplink quietly killing a high-value journey
  • Latency or delivery degradation eating conversion before any message does

Think of Markin as a data science and growth team that never sleeps: it investigates, forms hypotheses, ships them into your own stack and proves each one against a control group, at a volume no human team can reach.

Job to be done

The same work, at a different throughput.

Nothing below needs a tool that does not exist. It needs the work to happen continuously instead of once a quarter, and to be proven against a holdout instead of argued about.

Job to be done, compared between With claude with mcp alone and With Markin
Job to be doneWith claude with mcp aloneWith Markin
Notice that revenue per customer is drifting in a segmentSomeone spots it in a dashboard review, weeks after it started.Detected as a signal the day the drift clears noise, with the segment already sized.
Explain why it is happeningAn analyst is pulled off the roadmap for a two-week investigation.An investigation runs automatically and returns the drivers with their evidence.
Come up with hypotheses worth testingA workshop produces the handful of ideas the room happened to think of.Hypotheses are written continuously across marketing, product, pricing and technical health.
Decide which hypotheses deserve budgetPrioritised by seniority and gut feel, with no size attached.Each one is sized in revenue and ranked before anything is built.
Choose the next best action for one customerSegment rules and campaign calendars decide, refreshed when someone has time.Chosen per customer, per moment, against everything else competing for that customer.
Actually launch itA ticket to the lifecycle team, then a slot in next month's calendar.Executed inside the systems you already run, with no new channel to adopt.
Prove it caused the revenueReported against non-qualifiers or a global holdout, if at all.Every decision carries a randomised control group; uplift is read against it.
Kill what does not workProgrammes survive because nobody owns retiring them.Failing to beat control retires the programme automatically.
Do all of it again next weekCapacity-bound: four to eight tests a quarter.Hundreds of hypotheses in flight in parallel, continuously.

Honest take

What Claude with MCP does better.

A comparison that only flatters one side is not worth reading. These are the cases where we would tell you to stay where you are.

  • Claude reasons better over unfamiliar data than any skill library

    Give it a schema it has never seen and a vague question, and it will do genuinely good work. Markin is narrow by design and would simply not have a skill for that.

  • MCP is Anthropic's protocol and their tool use shows it

    Multi-step tool orchestration is more reliable there than in most alternatives. We would rather say that plainly than pretend otherwise.

  • For one hard question, it is cheaper than any platform

    A single deep investigation costs a few dollars in tokens. Nothing we sell competes with that, and nothing should.

Where Markin fits

Not a replacement. A growth-science team on top.

Use Claude for the hard questions. Markin industrialises the ordinary ones: it runs the loop continuously and reserves expensive reasoning for the steps where language genuinely changes the outcome.

Versioned skills

Detection, sizing, design, reading and arbitration, each with its own evaluations.

Model routing by task

Cheap models for volume, frontier models for language, trained models for prediction.

Every action has a control

Incremental ARPU, measured, not argued.

Evidence standard

Most of this category reports its own lift.

None of the major engagement, CDP or personalisation vendors publishes an independently verified uplift figure for its decisioning product. Where numbers exist, they come from vendor-commissioned studies or single-customer case studies with no disclosed holdout methodology. The most rigorous public research in the category is not flattering to anyone, including us, which is exactly why we build against it.

How Markin holds itself to it

  • Every decision Markin makes carries a control group. Uplift is reported against that holdout, not against the customers who did not qualify.
  • Results are read over a full measurement window rather than in the first weeks, so novelty is not mistaken for effect.
  • Programmes that fail to beat control are retired automatically. Killing decisions that do not pay is part of the loop, not an annual review.
  • The one figure we quote about ourselves is a range, not an average: +17% to +35% ARPU on treated cohorts against a randomised holdout, across Markin deployments in large B2C bases. We publish no industry benchmark, because we could not source one we would be willing to defend. Your holdout is the number that matters.

Time to value

90 days to a number that survived a holdout.

No replatform, no data migration, no rebuild of the channels you already run. If the first cohorts do not beat control, nothing scales and you have lost a quarter, not a roadmap.

  1. Weeks 0–2

    Read the context you already have

    Markin connects to the data and the channels you run today, claude with mcp included. No migration, no replatform, no new source of truth.

  2. Weeks 3–6

    First sized opportunities in test

    Opportunities are ranked by expected value, treatments are chosen per customer, and the first cohorts go live with a randomised holdout attached.

  3. Weeks 7–12

    First verified incremental revenue

    Results are read over a full measurement window. What beats control scales; what does not is retired. Nothing scales on a number that has not survived a holdout.

Which one you should pick.

Choose Markin if

  • The output has to be a decision per customer, not a document.
  • Every claimed win needs a randomised control group.
  • Hypotheses must be ranked by expected value across marketing, product and pricing.
  • Per-customer reasoning has to run weekly on millions of records.
  • The method must be reproducible across quarters and teams.

Choose Claude with MCP if

  • The work is deep, one-off analysis.
  • Your analysts want leverage, not autonomy.
  • The data model is unfamiliar and needs interpretation before anything else.
  • You are not running controlled experiments yet.

When you don’t need Markin.

  • You need an analyst's assistant, not an autonomous loop.
  • The base is too small for controlled measurement.
  • Every treatment requires individual legal review before launch.

Questions buyers ask.

Could we replicate Markin with Claude and a few MCP servers?

You can replicate the demo, not the economics or the discipline. The gaps show up as cost per customer, absence of sizing and power checks, and the missing control group that makes a result defensible.

What does Claude do better?

Open-ended reasoning over unfamiliar data, long-context work and reliable multi-step tool use. For a single hard investigation it is the better and cheaper choice.

Does Markin use Claude?

Where a frontier model is the right instrument, we route to whichever performs best on that step. Most of the per-customer work never reaches a frontier model, because propensity, uplift and anomaly detection are done by trained models.

How do you evaluate hypotheses differently?

Every candidate gets an eligible population, an expected value net of margin and contact cost, and a power check. What clears the threshold runs with a randomised holdout, and what loses is retired rather than quietly kept.

How is Markin different from the decisioning or AI already inside claude with mcp?

A decisioning engine ranks actions a human already defined, inside the campaign surface it was given. Markin forms the hypotheses itself, marketing, product, pricing or a technical anomaly holding growth back, sizes them, executes them inside claude with mcp and your product surfaces, and reads each one against a randomised holdout. It behaves like a data science and growth team, not like an optimiser.

Does Markin only test messages and offers?

No. Anything a human growth scientist would investigate is in scope: onboarding friction, feature adoption, pricing and packaging, dunning, and technical health issues such as a checkout error rate or a broken deeplink quietly killing conversion. Marketing is one of four hypothesis domains, not the boundary.

What is the business case for adding Markin on top of claude with mcp?

On a large B2C base, a small move in ARPU is a large number in absolute terms, because it applies to the whole installed base every month rather than to a campaign. Across Markin deployments the verified range on treated cohorts is +17% to +35% ARPU against a randomised holdout. The point is not more messages: it is finding the highest-value action per customer, launching it, and proving it against control before it scales.

How long before it pays for itself?

First sized opportunities are in test within six weeks and the first holdout-verified result lands inside 90 days. Payback depends on your base, margin and programme cost, the calculator on this page computes it from your own numbers, after applying the 20% to 40% haircut BCG finds when next-best-action programmes are incrementality-tested.