Insights·Testing·7 min read

Incrementality Testing: The Only Measurement That Actually Proves Causation

Every other measurement tool tells you what correlated with a conversion. Incrementality testing tells you what caused it. That distinction is worth a lot of money.

Byline
James Murray
March 15, 2024
IncrementalityTestingExperimentationROI
Pivotal Consulting Group2024

The question attribution can't answer

A channel appears in every winning conversion path. Spend on it rises, conversions rise. Spend drops, conversions drop. The platform calls this proof of performance. It isn't. It's correlation, and it lives inside a model that was built by the channel doing the claiming.

Incrementality testing answers the question those models can't: if this channel went dark tomorrow, what would actually be lost? The answer requires a real experiment, not a retrospective analysis.

What the test actually measures

The setup is a control group and a test group, as similar as possible on every dimension that matters, with one difference: one group sees the marketing and one doesn't. The lift between them, after accounting for the natural difference in their baselines, is the incremental effect.

The design varies by what the program will support. Geo holdouts split similar markets. Audience holdouts randomly withhold ads from a matched user segment. Ghost-ad studies serve a non-clickable placeholder to control exposure while withholding the real ad. Platform-native lift studies use the same logic inside a single channel's measurement infrastructure. The right choice depends on spend levels, addressability, and how clean you need the result to be.

None of these are cheap to run well. All of them are cheaper than funding a channel that isn't working.

The design decisions that determine whether you learn anything

Statistical power is the one most teams underinvest in. A test that can't detect a 5% lift difference will produce a null result that looks like evidence the channel doesn't work, when it was actually just underpowered. Power analysis before launch is not optional.

Test duration matters for the same reason. Running a geo holdout for ten days on a channel with a 30-day consideration cycle measures nothing. The test needs to run long enough to capture the full conversion window, including the lagged effects that most platforms don't report.

Contamination is the silent killer of geo tests. If treatment and control markets share media overlap, share distribution, or share customer populations that interact with each other, the holdout isn't clean. In our experience, most B2B geo holdouts require more market separation than the initial design accounts for.

Where tests fit in the measurement stack

Incrementality is the strongest signal in the toolkit. It's also slow and relatively expensive, which means you can't test everything simultaneously. The sequencing matters: test the channels carrying the most spend first, because that's where a wrong assumption costs the most. Test structural questions (does this channel work at all?) before tactical ones (which creative drives more lift?).

The results don't just answer the channel question. A well-run incrementality test gives you the number you need to calibrate your attribution model, which then makes the day-to-day tactical decisions more reliable even when you're not running a formal experiment.

What changes when you do this

Teams that run incrementality regularly stop arguing about attribution and start running experiments. The channel that looked like a top performer on last-click either proves it or doesn't. Budget follows evidence instead of the vendor with the best reporting interface. That shift is harder to achieve than it sounds, and it's exactly what it looks like when measurement is working.

Let's talk

What can we
do for you?