Most incrementality tests are autopsies. By the time the lift number lands, the outcome was mostly decided already, not by the test, but by where and how the ad ran. Marketers treat incrementality as a grade they receive at the end of the quarter. It's closer to a decision they made at the start, when they chose where the ad would run.

That reframe matters because incrementality has become the KPI everyone wants and almost no one trusts. Across the industry's major measurement surveys, most advertisers say their stack, attribution, MMM, and incrementality together, still falls short on the rigor, speed, and trust they need (IAB, State of Data 2026). We'd argue the gap isn't a measurement failure. It's a media one.

What a lift test actually measures

Every lift test asks one question: did this ad reach someone who wouldn't have converted anyway? That "anyway" is the whole game. Incrementality is what happened minus what would have happened without the ad, the counterfactual, and the size of that gap is set less by the test than by the placement you chose to run.

Which leads somewhere uncomfortable: the channels marketers trust most are often their least incremental. Branded search and retargeting convert beautifully and look great on the dashboard, precisely because they collect demand that was already headed for the cart. Across 225 DTC geo tests, branded search posted a median 0.70x incremental ROAS, while the top performer, Tatari CTV, reached 3.30x (Stella). A high conversion rate, in other words, can be a symptom of low incrementality, not proof of high.

Call it the difference between demand harvesting and demand introduction. Harvesting collects intent that already exists. Introduction creates a sale that wouldn't have happened otherwise. Both belong in a media plan. Only one is incremental.

And the read is only as trustworthy as the control group beneath it. Platform ad experiments can quietly confound the ad's effect with the algorithm's targeting, because the platform delivers each variant to its own optimized mix of users, a bias documented in a 2025 Journal of Marketing study (Braun & Schwartz) that only deepens as the algorithms get smarter. The fix is a real, randomized counterfactual, the same population, minus the ad. That's what separates a causal read from a flattering correlation.

Where incrementality fits in the stack

No single model is the source of truth. Attribution is fast and granular for daily optimization. MMM sets the strategic picture. Incrementality is the causal check that calibrates both, tightening MMM's priors and validating what attribution claims. The strongest programs run all three and treat the disagreements as signal, not noise.

But designing for incrementality is only half the job. The test's real purpose isn't to manufacture lift; it's to reveal which placements earned it, so you can move budget behind them. That's where most programs stall. 71% of advertisers call incrementality their top retail media KPI, yet only 15% rate themselves very or extremely effective at measuring it (Skai, 2026). The gap usually isn't methodology. It's the discipline to move money once the numbers challenge a line item that looked good on paper. Platform-reported ROAS can overstate true incremental impact by 30 to 60% on harvesting channels like paid social and retargeting (Measured), which means the placements that look most efficient are often the ones quietly overfunded.

Why Rokt is built for introduction

Rokt operates at the Transaction Moment, the window between checkout and confirmation, when someone has just decided to buy. Structurally, that's demand introduction by design. The shopper has committed to one purchase, so a relevant offer from a different advertiser reaches someone who wasn't shopping for that brand at all. The shopper's intent was for a different product entirely, so relative to the advertiser we're introducing, there's no existing demand to harvest. We're creating a moment of consideration that wouldn't have existed.

Engagement there is high, a 4.03% average CTR (10x Google Display, 4x Facebook) and a 6.32% CVR. But high conversion is exactly the trap above, so we treat it as a symptom to test, not proof to trust. Structural position earns favorable odds; it doesn't hand out a free pass, which is why every placement still gets measured. And we measure with more than one rigorous method. Ghost ads give a user-level holdout, where the control group is made of people who would have seen the offer, matched the targeting, cleared suppression, won the auction, but were randomly held back at the last step, so treatment and control come from the same eligible population. Geo-based lift testing, including the matched-market and synthetic-control designs partners like Haus run, compares whole geographies and captures the offline and omnichannel effects a user-level test can miss. Different designs, different strengths. What they share is a real, randomized control, and we use both.

Incrementality was never a grade waiting at the end of the quarter. It's a choice you make every time you decide where the next dollar runs.

Sources

ANA retail media KPI survey (2024); IAB x BWG Global State of Data 2026; Braun & Schwartz, "Where A/B Testing Goes Wrong," Journal of Marketing (2025); Stella 2025 DTC Incrementality Benchmarks (225 geo tests); Skai 2026 State of Retail Media Measurement; Measured (iROAS overstatement); Rokt by the Numbers (2026).