Reinforcement learning for marketing
Also called: Contextual bandits
Reinforcement learning treats customer decisions as a sequence, optimising cumulative long-term reward rather than the next click. In marketing it usually appears as contextual bandits, which balance exploiting the current best action against exploring alternatives that might be better.
Why it matters for ARPU
Sequences matter for ARPU: the action that maximises this month can suppress next quarter. Optimising the trajectory is what separates lifetime value from short-term conversion.
Related terms
Multi-armed banditA multi-armed bandit allocates traffic adaptively toward variants that are performing well while continuing to sample the others.Next best actionNext best action is the single intervention that maximises expected value for a specific customer at a specific moment, chosen across every available option including doing nothing.Uplift modelAn uplift model estimates the change in outcome caused by treating a customer, rather than the outcome itself.Lifetime valueLifetime value is the discounted margin a business expects from a customer over the whole relationship.