RESOURCES/Guide
How to evaluate a decisioning platform
Evaluating a decisioning platform comes down to four things: who authors the candidate actions, what objective the ranking optimises, whether the system can execute without a human build step, and whether every decision carries a control group. A product that fails any of the four is a rules engine with a model attached.
Why demos are a bad evaluation instrument
Every decisioning demo looks the same: a customer profile, a ranked list, a confident number next to each row. The differences that matter are not visible on that screen. They show up three months in, when someone asks where the candidate actions came from, why the winner was the one with the highest click probability, and who is holding the control group.
- A ranked list proves nothing about how the list was assembled.
- Engagement objectives look excellent in a demo and over-contact people who would have converted anyway.
- 'Real time' usually describes the API, not the decision.
- Uplift on the slide is almost always attributed, not measured against a holdout.
Twelve questions, and the answers that should worry you
Ask these in the order below. The first four are disqualifying; the rest are trade-offs you can live with once you know about them.
| Question | The answer that should worry you | What good looks like |
|---|---|---|
| Who writes the candidate actions? | "Your team configures them in the UI." | The system proposes actions you did not think of, and can explain the evidence behind each one. |
| What does the ranking optimise? | Propensity to open, click or convert. | Expected incremental revenue, net of margin, contact cost and fatigue. |
| Can it decide to do nothing? | Hold is a suppression rule you configure. | Hold is a first-class candidate that wins on expected value, routinely. |
| Who launches the winning action? | It exports a recommendation; your team builds it. | The platform executes inside the systems you already run, without a build queue. |
| Where does the control group live? | "You can set one up per campaign." | Randomised holdout attached to every decision by default, not per campaign. |
| What is in scope beyond messaging? | Message, offer, channel, timing. | Pricing, packaging, onboarding, in-product surfaces and technical anomalies too. |
| How many hypotheses can run in parallel? | A number bounded by seats or by campaign slots. | Bounded by the base and by statistical power, not by human capacity. |
| What happens to a losing programme? | It appears in a quarterly review. | It is retired automatically when it fails to beat control. |
| How is uplift reported? | Conversions attributed to the journey. | Incremental revenue against a randomised holdout, over a full window. |
| What data does it need to own? | A full migration into their profile store. | It reads your warehouse or CDP and owns only the decision log. |
| How are guardrails expressed? | Frequency caps only. | Margin floors, contact economics, brand and legal constraints, per-segment eligibility. |
| What does the pilot prove? | Engagement metrics on a hand-picked segment. | A holdout-verified revenue number on a segment you chose. |
Two things to verify independently
Ask for the incrementality-tested number rather than the reported one. BCG finds 20% to 40% of measured next-best-action uplift disappears once a randomised control is applied, so the two figures are not interchangeable.
Independent researchBCG, incrementality in personalisation programmes (2026)Check who funded any performance study you are shown. The quantified economic-impact research published for BrazeAI Decisioning Studio, for instance, is a Total Economic Impact study conducted by Forrester Consulting and commissioned by Braze, which is a different evidence class from independent analyst research.
Vendor-commissionedForrester TEI, commissioned by Braze (May 2026)
What to demand in the pilot contract
A pilot that cannot fail is not a pilot. Write the failure condition down first.
- You choose the segment, not the vendor.
- A randomised holdout of a size you agree in advance, held for the full measurement window.
- One primary metric, stated before the pilot starts, expressed in revenue rather than engagement.
- A stated stopping rule and a stated failure condition, both written down before launch.
- Access to the decision log: for any customer, why that action, what it was worth, what happened.
- No migration as a precondition. If the pilot requires you to move your data first, it is not a pilot.
Reasons to walk away from the category entirely
- Your outcome data is unreliable or heavily delayed. Fix measurement before buying decisioning.
- Your constraint is delivery capacity, not decision quality. Buy execution capacity instead.
- You cannot obtain a randomised holdout for organisational reasons. Without one you will never know whether it worked.
- The commercial case rests on a small base. Below roughly 100,000 active customers, the arithmetic rarely justifies the programme.
Markin is an autonomous growth-science team for large B2C businesses. It investigates why revenue per customer is stuck, forms its own hypotheses across marketing, product, pricing and technical health, chooses the next best action for each customer, launches it through the systems the business already runs, and proves every one against a randomised holdout.
Decisioning tools choose between the actions your team already built. Markin decides what to build.
Questions people ask
- What is the single most important question to ask a decisioning vendor?
- Who writes the candidate actions. If the honest answer is that your team configures them and the system only ranks them, the product's ceiling is the imagination and capacity of your team. Everything else, models, real-time APIs, channel coverage, is downstream of that limit.
- How much should a decisioning pilot cost?
- The number that matters is not the fee but the ratio between the fee and the incremental margin the pilot has to produce to justify itself. Write that threshold down before the pilot starts and express it in incremental revenue against a holdout, not in engagement metrics.
- Should we build this in-house?
- Building the ranking is the easy part and most good data science teams can do it. The parts that consume years are activation into every channel, guardrail management, automated experiment design and readout, and keeping it all running when the schema changes. Evaluate the build against those, not against the model.
- Do we need a CDP before we buy decisioning?
- No. A decision layer needs access to reliable customer context, which can come from a warehouse, a CDP, or a mixture. Requiring a CDP first is a sequencing preference, not a technical dependency, and it delays the revenue case by a year in most organisations.