Next steps

Knowing it worked

Targeting is the easy half. This is the other half: which segments are large enough to prove anything about at all, whether customers actually moved over a quarter, and what gets tested before a single message goes out.

Where it sits: Measure and Interpret: the half of the loop that decides what the next revolution changes.

The gap

A campaign report cannot answer the question you are being asked

It tells you what happened after a send: who opened, who clicked, who converted in the window. What it cannot tell you is whether the customer is in a better place than they were, because nothing wrote down where they started.

It measures a send, not a person

A campaign report is one send to one list. A programme is a customer over a year, touched by forty of them. Summing the sends does not reconstruct the person.

It has nothing to compare against

A base moves on its own. Without someone deliberately held back, a rising line is a claim, not a result.

It has no memory

A baseline cannot be retrofitted. If nobody snapshotted the state last quarter, that quarter is gone, and the comparison you want in six months cannot be made.

Before you measure anything

What can this model actually prove?

Segment sizes decide what is knowable. A cohort of a few hundred will never produce a weekly significant read, however good the targeting is, and presenting it beside a cohort of sixty thousand at the same visual weight is how someone ends up acting on a swing that means nothing. So the first question is not what moved. It is which segments could show a move at all.

Your programme

The model sets each segment's share of the base. The size of the base is yours, so start with it.

Base size150,000
The addressable base, before any segmentation

Every segment below is this number multiplied by the share the model gives it.

Lift you need to detect8%
A segment is steerable when it can see a lift this small

Bigger effects are easier to prove. A campaign hoping for 3 per cent needs far more people than one expecting 25.

Test design

The design of the test itself. These trade reach against certainty.

Baseline conversion4%
A held-back customer converts in 4 per cent of weeks

Rarer events are harder to read. A 1 per cent action needs roughly four times the audience of a 4 per cent one.

Holdout10%
15,000 customers held back across the base

A bigger control reads tighter but reaches fewer people. Below about 10 per cent the control itself becomes the limit.

Weeks running12
12 weeks of evidence

Letting a test run longer is the cheapest way to make a small segment readable, and the thing most often cut short.

What each segment could prove

smallest detectable lift

3 of 13 segments can carry a read on their own, covering 63 per cent of the base. The rest are real segments and worth targeting. They are simply not worth steering on weekly, and saying so is what stops a quiet fortnight being read as a failing campaign.

Casual / Low-engagement61,770 customers, 41.2% of the base3.8%Steerable
Onboarding17,550 customers, 11.7% of the base7.2%Steerable
Rising Customer15,870 customers, 10.6% of the base7.6%Steerable
Prospect12,240 customers, 8.2% of the base8.7%Read it rolled up
Subscriber10,320 customers, 6.9% of the base9.5%Read it rolled up
At-Risk7,290 customers, 4.9% of the base11.4%Read it rolled up
Champion4,770 customers, 3.2% of the base14.3%Read it rolled up
One-and-done4,620 customers, 3.1% of the base14.6%Read it rolled up
Loyal Habitual4,050 customers, 2.7% of the base15.6%Read it rolled up
Single-Category Loyalist3,270 customers, 2.2% of the base17.5%Monitor only
Reactivated2,970 customers, 2.0% of the base18.5%Monitor only
Deal-Driven2,700 customers, 1.8% of the base19.5%Monitor only
Hibernating2,580 customers, 1.7% of the base19.9%Monitor only

Read it at the level that has the power

When a segment cannot carry a read, the answer is not a longer squint at the same number. It is to roll up a level, and to say which level you are reading.

The whole base
150,000
2.4%
The largest tier
60,690
3.8%
The largest segment
61,770
3.8%

Illustrative and deterministic. The shares come from a modelled base of 5,000 scored by the same engine as the tuner; the headcounts are those shares at the base size you set. The threshold is worked out from the same maths the holdout simulator uses, so the two tools can never disagree about what counts as proven.

The baseline

Did anyone actually move?

One quarter of the whole base, with nothing done to it: where every customer started, and where they ended up. No campaign report can produce this view, and it is the reason the model writes down where everyone stood each time it runs instead of only working out where they stand today.

rows: where they startedcolumns: where they ended, one quarter later
Platinum
Gold
Silver
Bronze
Dormant
Left
started
Platinum
73%
27%
131
Gold
11%
70%
19%
1%
371
Silver
9%
69%
21%
1%
880
Bronze
8%
76%
14%
2%
2,023
Dormant
15%
80%
5%
1,595
Joined
2%
70%
28%
350
25%
of the base changed tier or left it during the quarter
9%
is the largest move any tier TOTAL made, which is what a trend chart would have shown you
136
customers left the base entirely, and appear in no tier at the end

Read a row across: of everyone who started in that tier, where they were a quarter later. The hatched column is the one a migration report usually omits. The Joined row has its own denominator, a share of all arrivals, because a new customer has no starting tier to be a share of. Illustrative and deterministic: the same modelled base as the tuner, advanced one quarter and rescored by the same engine, with no campaign applied.

Why the totals lie

The tier totals barely move, while a quarter of the base changes tier underneath them. A trend chart of tier counts would have called this a flat quarter. It was not a flat quarter, it was a busy one that happened to net out, and the difference matters because the people flowing down are not the people flowing up.

This is the whole base with nothing done to it, which is what makes it a baseline. The same question asked of a single treated cohort is below.

The other grain

And what one treated cohort actually bought

The matrix above is the whole base with nothing done to it. This is one cohort that WAS worked, set against the slice of itself deliberately held back. The gap between the columns is what the plan earned, as opposed to what would have happened anyway.

Quarter startSingle-Category Loyalist at SilverQuarter end
Single-Category LoyalistStill in the same cell
46%
62%-16
Loyal HabitualGained breadth, the move this plan is buying
27%
12%+15
Single-Category Loyalist at GoldSame journey, dialled up a tier
14%
11%+3
At-RiskSlipped, and picked up by the retention plan
13%
15%-2

The same cohort three months later, treated against the held-out control. Illustrative and deterministic: a worked example, not a measurement.

Breadth gained is the move this plan is buying, and it is the one the control does not produce on its own.

The plan that produced it is on In Braze. Note the grains differ on purpose: this moves between behavioural segments, the matrix above moves between value tiers.

What counts as working

Working is five different claims

They get collapsed into one, and which system can even answer them is the argument. Most programmes report level two and imply level four.

1

It sent

Braze

The message left the building and reached an inbox or a handset. Delivery, not effect.

2

They engaged

Braze

Opened, clicked, tapped. Real, useful, and still entirely inside the channel: it says the creative worked, not that the customer did anything.

3

They did the thing

Warehouse

They purchased. The first level the activation tool cannot answer on its own, because the outcome it is being asked about happens somewhere it cannot see.

4

They did it because of us

Warehouse plus a holdout

The held-back control did not do it at the same rate. Without the control this is the tide being read as the boat.

5

Their state changed, and stayed changed

Snapshot table, over time

A quarter later they are in a different segment, and still there. This is the only level that describes a programme rather than a campaign, and it is the one almost nobody produces.

Levels four and five are the two that need something Braze cannot hold: a stable control and a history. Both have to be decided before the first send, not after the first question.

The trap

A reactivation that produces one purchase and no change in cadence is a level three success and a level five failure. The campaign report says yes. The programme says no. Both are reading their own numbers correctly.

That is not an argument for ignoring campaign reporting. It is an argument for not letting it answer a question about the programme.

How the holdout itself works →
The release loop

What ships before it sends

Learning that cannot be shipped is a slide, and a pipeline with nothing measuring the result is faster guessing. The mechanism that turns one into the other is ordinary release discipline, and how much of it you get depends entirely on where the logic lives.

01Fully automatable

Definitions and models

The rules get tested the way software does, automatically, on every change: every state still has somewhere to go, every KPI still has a lever, and each example customer still lands where they are meant to. The most useful check of the lot is the simplest, though. If a change moves an audience by more than a few per cent, the build stops and asks. That is what catches the change nobody thought was a change.

02Fully automatable

The sync into Braze

Before anything reaches Braze, the handover is checked against what was agreed: only the named attributes cross, no personal data goes with them, and no customer arrives in a state that no journey is built to handle. After the sync, the counts on both sides are compared. A silent mismatch here is how an audience quietly empties.

03Partly, and be honest about the rest

Journeys inside Braze

Content, templates and product feeds can be managed properly and rolled back. Test customers walked through every branch in a sandbox workspace is the real safety net, and the only thing that catches two journeys contradicting each other before a live customer sees both. What cannot be automated is the canvas itself: it is drawn by hand, so there is no way to review a change to it line by line. The written plan is the source of truth, and the canvas is built from it.

04Fully automatable

Measurement

The scorecard is a build artefact that regenerates on the model's schedule. If answering "is the win-back working" takes an analyst three days, it gets asked once a quarter and the loop never closes. Standing, not ad hoc, is the requirement.

The gradient is the argument

Everything upstream of Braze can be tested properly. Inside it, some can. So the more logic that moves into a canvas condition, the less of the programme can be tested at all, because every rule that moves there leaves version control, the test suite, the diff and the audit trail behind it.

The question worth asking of any programme: point at the rule that decides who gets the win-back, then show the test that proves it does what you think, and the commit where it changed.

And the ceiling on all of it

The model's refresh cadence sets the maximum learning rate of the programme. A monthly recompute gives twelve observations of movement a year, and no amount of release tooling speeds that up. If the weekly job also persists a snapshot row carrying value tier and behavioural state, that becomes fifty-two. The full recompute can stay monthly; only the resolved state is written more often. One line in the specification, four times the learning rate.

How we start

Source events, a Braze workspace, agreement on the size of the universal holdout, and the first two or three segments to target. Foundation first, then a quick win live, then the proof. The holdout and the snapshot are decided at the start, because those are the two things that cannot be added retrospectively.

The build, what you are responsible for and what it takes to run are set out on the service page. The Growth scorecard mapping is on the scorecard.

If you want to design a holdout and watch the confidence interval close, the incrementality simulator does exactly that. It is a separate model, so it opens in its own set of pages.

That is RFM working on sample data. The service page has the rest: what it needs from you, how long it takes to build, what it pairs with, and how the lift gets proven.