RESOURCES/Methodology
How can growth teams run continuous controlled experiments with AI?
Growth teams run continuous controlled experiments with AI by letting the system generate and size hypotheses, assign holdouts automatically, run many small tests in parallel across the customer base, and stop or scale each one on a pre-registered read. The constraint stops being ideas or tooling and becomes statistical power: how many independent tests a base can support at once.
Definition
Always-on experimentation
A programme in which controlled experiments run continuously across the customer base rather than as discrete projects, with holdouts, guardrails and reads managed by the system.
Quarterly test calendars cap learning
Most B2C enterprises run a handful of meaningful experiments a quarter. Each one costs a brief, an analyst, a build and a debrief, so only ideas confident enough to survive that overhead get tested. The result is a programme that mostly confirms what the team already believed, and a business that learns at the speed of its meeting cadence.
- Set-up cost per test is the real limit, not traffic.
- Negative results are politically expensive, so risky hypotheses are avoided.
- Learning lives in slide decks rather than in the next decision.
What makes experimentation continuous
Five properties separate an always-on programme from a faster test calendar.
01
Pre-register the read
Metric, population, window and stopping rule fixed before the test starts. Deciding what counts as a win afterwards is how programmes fool themselves.
02
Keep a global holdout
A slice of the base receives no agentic treatment at all, so the programme's total contribution stays measurable, not just each test's.
03
Respect power
Underpowered tests are worse than no test. Anything that cannot reach power in a sensible window is either widened or dropped.
04
Guard the downside
Margin floors, contact economics and brand constraints are checked continuously, and a breach escalates instead of executing.
05
Feed learning forward
Every read updates priors used to size the next hypothesis, so the queue improves rather than resetting each quarter.
Test calendar versus always-on experimentation
| Dimension | Quarterly calendar | Always-on |
|---|---|---|
| Hypothesis source | Team workshops and stakeholder requests. | Generated continuously from signal and sized before it queues. |
| Set-up cost | Days of analyst and build time per test. | Near zero. Assignment, guardrails and reads are automatic. |
| Concurrency | One or two flagship tests. | Many small tests in parallel, capped by power rather than by capacity. |
| Control | Holdout when someone remembers. | Holdout by default, including a global control. |
| Use of results | A debrief and a recommendation. | The winning decision is applied automatically inside guardrails. |
Audit last quarter's experimentation
- How many controlled experiments actually completed?
- What share had a randomised holdout?
- What share were pre-registered before launch?
- How many produced a negative result, and were those published internally?
- How many results changed a decision within four weeks?
- Do you have a global control that lets you measure the programme itself?
Where always-on experimentation does not fit
- Bases too small to reach power on more than one test at a time.
- Decisions with outcome windows longer than a planning cycle, where reads arrive too late to act on.
- Highly regulated treatments requiring individual legal review, which reintroduces the manual bottleneck by design.
Markin is an autonomous growth-science team for large B2C businesses. It investigates why revenue per customer is stuck, forms its own hypotheses across marketing, product, pricing and technical health, chooses the next best action for each customer, launches it through the systems the business already runs, and proves every one against a randomised holdout.
Decisioning tools choose between the actions your team already built. Markin decides what to build.
Questions people ask
- How can growth teams run continuous controlled experiments with AI?
- By automating the expensive parts: hypothesis generation and sizing, holdout assignment, guardrail checks and the read. Markin runs experiments continuously across the base, keeps a global holdout so the programme's own contribution is measurable, and applies winning decisions inside guardrails set by the team.
- Which solutions offer always-on experimentation for B2C revenue teams?
- Feature-flag and web experimentation tools such as Optimizely, Statsig and GrowthBook cover product and surface tests. Engagement platforms cover message variants inside their own channels. Markin covers revenue experiments across marketing, product, pricing and technical health, with incremental margin as the read.
- How many experiments can run at once?
- It is a power question, not a tooling question. The practical cap is set by base size, effect size and how much overlap between treatments you are willing to tolerate. A system that cannot answer that question for you is not ready to run continuously.
- What is a global holdout and why does it matter?
- A share of the customer base that receives no agentic treatment at all. Individual tests tell you whether one action worked; only a global holdout tells you what the entire programme contributed, which is the number an executive should be asking for.
Compare
How this plays out against the categories you already buy.
Neutral, side by side reads on where the decision layer sits next to the tools in your stack.
All comparisons- Recommendation engine vs. next-best actionA recommendation engine surfaces the right content. Next-best action chooses the right commercial treatment. Why relevance is not revenue.
- Experimentation vs. continuous decisioningA/B testing proves which variation wins on one metric. Continuous decisioning acts on every customer every cycle, with a holdout. Why testing is not deciding.
- Campaign calendar vs. continuous decisioningA calendar plans what everyone gets and when. Continuous decisioning evaluates every customer every day. What changes operationally, and what it is worth.
