Knowing it worked
Targeting is the easy half. This is the other half: which segments are large enough to prove anything about at all, whether customers actually moved over a quarter, and what gets tested before a single message goes out.
Where it sits: Measure and Interpret: the half of the loop that decides what the next revolution changes.
A campaign report cannot answer the question you are being asked
It tells you what happened after a send: who opened, who clicked, who converted in the window. What it cannot tell you is whether the customer is in a better place than they were, because nothing wrote down where they started.
It measures a send, not a person
A campaign report is one send to one list. A programme is a customer over a year, touched by forty of them. Summing the sends does not reconstruct the person.
It has nothing to compare against
A base moves on its own. Without someone deliberately held back, a rising line is a claim, not a result.
It has no memory
A baseline cannot be retrofitted. If nobody snapshotted the state last quarter, that quarter is gone, and the comparison you want in six months cannot be made.
What can this model actually prove?
Segment sizes decide what is knowable. A cohort of a few hundred will never produce a weekly significant read, however good the targeting is, and presenting it beside a cohort of sixty thousand at the same visual weight is how someone ends up acting on a swing that means nothing. So the first question is not what moved. It is which segments could show a move at all.
Your programme
The model sets each segment's share of the base. The size of the base is yours, so start with it.
Every segment below is this number multiplied by the share the model gives it.
Bigger effects are easier to prove. A campaign hoping for 3 per cent needs far more people than one expecting 25.
Test design
The design of the test itself. These trade reach against certainty.
Rarer events are harder to read. A 1 per cent action needs roughly four times the audience of a 4 per cent one.
A bigger control reads tighter but reaches fewer people. Below about 10 per cent the control itself becomes the limit.
Letting a test run longer is the cheapest way to make a small segment readable, and the thing most often cut short.
What each segment could prove
smallest detectable lift3 of 13 segments can carry a read on their own, covering 63 per cent of the base. The rest are real segments and worth targeting. They are simply not worth steering on weekly, and saying so is what stops a quiet fortnight being read as a failing campaign.
Read it at the level that has the power
When a segment cannot carry a read, the answer is not a longer squint at the same number. It is to roll up a level, and to say which level you are reading.
Illustrative and deterministic. The shares come from a modelled base of 5,000 scored by the same engine as the tuner; the headcounts are those shares at the base size you set. The threshold is worked out from the same maths the holdout simulator uses, so the two tools can never disagree about what counts as proven.
Did anyone actually move?
One quarter of the whole base, with nothing done to it: where every customer started, and where they ended up. No campaign report can produce this view, and it is the reason the model writes down where everyone stood each time it runs instead of only working out where they stand today.
Read a row across: of everyone who started in that tier, where they were a quarter later. The hatched column is the one a migration report usually omits. The Joined row has its own denominator, a share of all arrivals, because a new customer has no starting tier to be a share of. Illustrative and deterministic: the same modelled base as the tuner, advanced one quarter and rescored by the same engine, with no campaign applied.
Why the totals lie
The tier totals barely move, while a quarter of the base changes tier underneath them. A trend chart of tier counts would have called this a flat quarter. It was not a flat quarter, it was a busy one that happened to net out, and the difference matters because the people flowing down are not the people flowing up.
This is the whole base with nothing done to it, which is what makes it a baseline. The same question asked of a single treated cohort is below.
And what one treated cohort actually bought
The matrix above is the whole base with nothing done to it. This is one cohort that WAS worked, set against the slice of itself deliberately held back. The gap between the columns is what the plan earned, as opposed to what would have happened anyway.
The same cohort three months later, treated against the held-out control. Illustrative and deterministic: a worked example, not a measurement.
Breadth gained is the move this plan is buying, and it is the one the control does not produce on its own.
The plan that produced it is on In Braze. Note the grains differ on purpose: this moves between behavioural segments, the matrix above moves between value tiers.
Working is five different claims
They get collapsed into one, and which system can even answer them is the argument. Most programmes report level two and imply level four.
It sent
BrazeThe message left the building and reached an inbox or a handset. Delivery, not effect.
They engaged
BrazeOpened, clicked, tapped. Real, useful, and still entirely inside the channel: it says the creative worked, not that the customer did anything.
They did the thing
WarehouseThey purchased. The first level the activation tool cannot answer on its own, because the outcome it is being asked about happens somewhere it cannot see.
They did it because of us
Warehouse plus a holdoutThe held-back control did not do it at the same rate. Without the control this is the tide being read as the boat.
Their state changed, and stayed changed
Snapshot table, over timeA quarter later they are in a different segment, and still there. This is the only level that describes a programme rather than a campaign, and it is the one almost nobody produces.
Levels four and five are the two that need something Braze cannot hold: a stable control and a history. Both have to be decided before the first send, not after the first question.
The trap
A reactivation that produces one purchase and no change in cadence is a level three success and a level five failure. The campaign report says yes. The programme says no. Both are reading their own numbers correctly.
That is not an argument for ignoring campaign reporting. It is an argument for not letting it answer a question about the programme.
How the holdout itself works →What ships before it sends
Learning that cannot be shipped is a slide, and a pipeline with nothing measuring the result is faster guessing. The mechanism that turns one into the other is ordinary release discipline, and how much of it you get depends entirely on where the logic lives.
Definitions and models
The rules get tested the way software does, automatically, on every change: every state still has somewhere to go, every KPI still has a lever, and each example customer still lands where they are meant to. The most useful check of the lot is the simplest, though. If a change moves an audience by more than a few per cent, the build stops and asks. That is what catches the change nobody thought was a change.
The sync into Braze
Before anything reaches Braze, the handover is checked against what was agreed: only the named attributes cross, no personal data goes with them, and no customer arrives in a state that no journey is built to handle. After the sync, the counts on both sides are compared. A silent mismatch here is how an audience quietly empties.
Journeys inside Braze
Content, templates and product feeds can be managed properly and rolled back. Test customers walked through every branch in a sandbox workspace is the real safety net, and the only thing that catches two journeys contradicting each other before a live customer sees both. What cannot be automated is the canvas itself: it is drawn by hand, so there is no way to review a change to it line by line. The written plan is the source of truth, and the canvas is built from it.
Measurement
The scorecard is a build artefact that regenerates on the model's schedule. If answering "is the win-back working" takes an analyst three days, it gets asked once a quarter and the loop never closes. Standing, not ad hoc, is the requirement.
The gradient is the argument
Everything upstream of Braze can be tested properly. Inside it, some can. So the more logic that moves into a canvas condition, the less of the programme can be tested at all, because every rule that moves there leaves version control, the test suite, the diff and the audit trail behind it.
The question worth asking of any programme: point at the rule that decides who gets the win-back, then show the test that proves it does what you think, and the commit where it changed.
And the ceiling on all of it
The model's refresh cadence sets the maximum learning rate of the programme. A monthly recompute gives twelve observations of movement a year, and no amount of release tooling speeds that up. If the weekly job also persists a snapshot row carrying value tier and behavioural state, that becomes fifty-two. The full recompute can stay monthly; only the resolved state is written more often. One line in the specification, four times the learning rate.
How we start
Source events, a Braze workspace, agreement on the size of the universal holdout, and the first two or three segments to target. Foundation first, then a quick win live, then the proof. The holdout and the snapshot are decided at the start, because those are the two things that cannot be added retrospectively.
The build, what you are responsible for and what it takes to run are set out on the service page. The Growth scorecard mapping is on the scorecard.
If you want to design a holdout and watch the confidence interval close, the incrementality simulator does exactly that. It is a separate model, so it opens in its own set of pages.
That is RFM working on sample data. The service page has the rest: what it needs from you, how long it takes to build, what it pairs with, and how the lift gets proven.
