Incrementality testing, explained: how a causal lift is measured when the screen is outdoors
Incrementality testing measures the outcomes a campaign caused by comparing what happened with a counterfactual: what would have happened without it. Brand lift asks what people think and attribution assigns credit; incrementality asks whether the outcome would have occurred anyway. Out-of-home cannot randomise exposure, so the design starts with geography or with exposure inferred from device location.
Every incrementality test compares a number that was observed with a number that was not, and the discipline is in how the second number gets made. Digital channels can randomise people; a screen on a street shows the same thing to everyone who passes. Below: the designs that survive that constraint, the arithmetic that decides whether a test can see anything, and what a brief has to settle before the flight.
What is incrementality testing?
Incrementality testing measures the part of an outcome that a campaign caused, by comparing the observed result with a counterfactual. Meta’s GeoLift documentation puts it in one sentence: a campaign’s incremental effect is “the difference between what we observed as the results of that campaign and what would have happened in a world where it didn’t take place”. The outcome is a behaviour, a store visit, a sale, an install. The counterfactual is a control that did not get the campaign.
That separates it from the two measures it is confused with. A brand lift study measures a change in what people think, through a survey. Attribution assigns credit for an outcome that did happen to the touchpoints that preceded it. The OAAA’s 2020 measurement guide keeps measurement (plays, impressions, reach, frequency) and attribution (analytics that “demonstrate causality or correlation”) as separate disciplines; incrementality is the causal end of the second. Attribution can credit a channel with visits that would have happened regardless. An incrementality test is built so that it cannot.
How is incrementality calculated?
Incremental outcome equals observed outcome minus counterfactual outcome. In a two-group design incremental lift is exposed rate minus control rate, and relative lift divides that gap by the control rate. The OAAA guide defines lift as the “percent difference in visitation rates between exposed audience and unexposed audience” and adds a variant, incremental lift, in which “a known visiting quantity is removed from the calculated set” so only visits above baseline remain.
A published example: Clear Channel Outdoor’s RADARProof study of an Atlanta fast casual campaign (Cuebiq data, August 2021) reports that over 15% of exposed adults visited, 30.6% more than non-exposed adults. Read backwards, the control rate is about 11.5% and the absolute lift about 3.5 points; its second figure, 28% of exposed visits classed as incremental, is the OAAA’s variant. A result that cannot be converted both ways has left out a base rate.
Which test designs work for out-of-home?
Four designs cover the field. They differ in where the counterfactual comes from, and that decides which of them out-of-home can use.
| Design | Where the counterfactual comes from | When it fits out-of-home | Where it breaks |
|---|---|---|---|
| Geo holdout / matched market | Whole regions assigned to treatment or control: randomised (Google, 2011) or matched on shared traits (IAB, 2025). | Multi-market campaigns with a clean pre-period. Needs no exposure log. | One market, or a campaign that must run everywhere. Measured: 10 to 15 matched markets, 95%+ historical correlation. |
| Synthetic control | A weighted combination of untreated units built to track the treated unit’s pre-period (GeoLift), or propensity-matched devices (Reveal Mobile). | A single exposed market, or a device panel where nobody was randomised. One of the IAB’s two primary DOOH methods. | Only as good as the matching variables. A control of devices never logged near a screen still contains people who passed it. |
| Ghost ads / PSA control | Individuals randomised at the ad server; the control sees a public-service message, or the platform logs what they would have seen. | Digital channels where a platform decides, person by person, who sees what. | Out-of-home. A screen is one-to-many. Withholding at screen level is a small geo holdout, not a ghost-ad test. |
| Pre/post with seasonality control | The same market before the flight, adjusted for what normally happens in that period. | A first read when nothing can be held out, or a sanity check beside a proper control. | Everything else that changed in the period is inside the number. |
Geo holdout and matched market tests
Vaver and Koehler’s 2011 Google paper describes the geo experiment: “non-overlapping geographic regions are randomly assigned to a control or treatment condition, and each region realizes its assigned condition through the use of geo-targeted advertising”. The matched market test is the same idea without randomisation; the IAB’s July 2025 DOOH Measurement Guide describes dividing markets into test and control on shared traits, prioritised by campaign goal. Neither needs an exposure log. What they need is markets: Measured’s July 2026 guidance is 10 to 15 matched markets with 95% or higher historical correlation before the test begins, most tests running 4 to 8 weeks with a holdout of 10 to 20% of the test region’s budget. A campaign in one city has nothing to hold out.
Synthetic control
With one treated market, or no randomisation at all, the counterfactual is modelled. GeoLift builds it from untreated locations with synthetic control methods (SCMs): “using historical information prior to the treatment SCMs find the combination of untreated units that most closely replicate the treated”. Haus describes the out-of-home version, a New York flight against a synthetic control built from the remaining DMAs, because out-of-home audiences “operate without individual-level tracking, pixels, or PII”. The device-level analogue is Reveal Mobile’s net lift study, which propensity-matches control users so they are “identical to the exposed group, except for the fact that they were not exposed”, reported at 95% confidence. The IAB lists synthetic control beside matched markets as the two primary DOOH methods, with a caveat: “since exposure records may be incomplete, blending synthetic models with geo-based holdouts can help reduce contamination of the control group”.
Ghost ads and PSA controls, and why they barely work outdoors
In a PSA test the ad server randomises individuals and shows the control group a public-service message in the brand ad’s slot. Johnson, Lewis and Nubbemeyer’s ghost ads paper (2016; Journal of Marketing Research, 2017) lists the problems: PSAs “are expensive and prone to errors because they require coordination among advertisers, publishers, and third-party charities”, and performance-optimised delivery breaks them by sending the PSA and the brand ad to different people. Ghost ads instead log, at the platform, the impressions a control user would have received, at “at least an order of magnitude less” cost. Both rest on a system that assigns and records exposure person by person. A screen has none: it cannot show the PSA to half the pavement, and a PSA on a public screen costs the full slot while producing a control group of everyone who passed. Holding out some screens is a geo holdout at the scale of a street corner, and the same person passes both.
Pre/post with a seasonality control
The OAAA guide permits lift “versus the control group or before the ad buy occurred”; the IAB describes comparing sales during the campaign with historical data “while accounting for seasonality or external factors”. That clause is the work. Seasonally adjusted, a pre/post read is a first estimate and a sanity check beside a proper control. On its own, everything else that changed in the period is inside the number.
Why can’t out-of-home randomise exposure?
Because nobody is assigned to a screen. A display faces a street and whoever walks along it is exposed, so an out-of-home study constructs its exposed group after the fact, from device location. The OAAA defines exposure as presence in a screen’s exposure zone, or viewshed, while content is viewable, and this “does not require that the content be viewed”. OUTFRONT’s May 2024 explainer gives the mechanics: measurement partners geofence the viewshed, and a phone passing through is marked present by its mobile advertising identifier. Reveal Mobile computes device-level exposure from media plans, ad rolls, experiential specs and moving waypoints, joined to a panel of 60 million US daily active users. Panel devices logged inside a viewshed are exposed; the control is drawn from the rest. Exposure is inferred, not served, and the inference is only as good as the location signal behind the panel.
The control is defined by absence of evidence: a device is in it because it was not logged in the viewshed, and a panel of 60 million is a sample. Some of the control passed the screen unrecorded, which drags measured lift toward zero. A Geopath impression, the number of people “likely to notice” an ad, is a modelled audience with no roster of devices; it sizes the audience and cannot populate a test cell. How that estimate is built, and why its inputs change for a vehicle, is on how out-of-home impressions are measured.
A moving screen sharpens both problems. A fixed viewshed is one polygon drawn once; a vehicle’s viewshed is a polygon per timestamp, and the exposed set is the union of every phone logged inside it along the route. The exposure record for mobile inventory is a route log joined to a panel, not a site list, which is why Reveal lists moving waypoints as an input. The IAB’s guide says exposure for exterior moving formats “is highly variable depending on vehicle speed, obstructions, and route patterns” and that fixed screens are “more compatible with geofencing”. For a test on vehicle-based inventory the route data is part of the exposure definition. Firefly, the operator behind the Foursquare-measured study in the results table below, lists foot traffic, web conversion and app conversion studies on its measurement page and describes its digital tops as GPS-equipped, which is where the route log comes from. What a moving fence does to the exposed panel is on geofencing advertising.
How large does an incrementality test need to be?
Large enough that the smallest lift worth acting on is bigger than the noise. That lift is the minimum detectable effect (MDE), fixed by the baseline rate, the confidence level and the statistical power. For a two-group visitation study the sample per group is
n = (zα/2 + zβ)2 × [p0(1 − p0) + p1(1 − p1)] ÷ (p1 − p0)2
where p0 is the control visit rate, p1 the exposed rate to detect, zα/2 is 1.96 for a 95% two-sided test and zβ is 0.84 for 80% power. Worked once: a 5.0% control rate and a 10% relative lift, which is 5.5%. The bracket is 0.0475 + 0.0520 = 0.0995; the denominator is 0.005 squared, 0.000025; the z-terms sum to 2.80 and square to 7.85. So 7.85 × 0.0995 ÷ 0.000025: about 31,200 devices per group before a 10% lift on a 5% base is reliably visible.
| Devices per group (approx.) | Lift detectable at 95% confidence, 80% power, 5% baseline |
|---|---|
| 3,800 | 5.0% to 6.5% (30% relative) |
| 8,200 | 5.0% to 6.0% (20% relative) |
| 31,200 | 5.0% to 5.5% (10% relative) |
| 100,000 | 5.0% to 5.27% (5.5% relative), read the other way: the smallest lift 100,000 per group can see |
Small samples can only see large lifts, one reason single-campaign case studies publish large percentages; 90% power pushes the 10% case to about 41,800 per group. Geo tests do not use this closed form, because their unit is a market and the noise is a time series. GeoLift ships power calculators to decide “which are the best test locations, how many you should include, investment, and even how long you should run the test”; Measured says a well-powered geo test targets an MDE of 2 to 5%; Haus’s rule of thumb is enough conversions per cell to detect a 10 to 20% lift at 80% power. Vendors state confidence differently, Clear Channel at 90% on its Atlanta study, Reveal at 95% on every study. A lift reported without one is a number, not a finding.
What do published out-of-home incrementality results look like?
Few, and rarely with base rates. The table lists the named, dated results this desk could verify on the publisher’s own page, with the method each states: datapoints on what the genre reports, not benchmarks to plan against.
| Study | Published by | Stated method | Reported result |
|---|---|---|---|
| Fast casual restaurant, Atlanta, digital billboards | Clear Channel Outdoor (RADARProof) | Exposed versus non-exposed devices, Cuebiq location data; lift confidence level stated as 90%. | Over 15% of exposed adults visited, 30.6% more than non-exposed; 28% of exposed visits classed as incremental; 27% on the day of exposure. Dated August 2021. |
| Cook Out restaurants, 24 panels, May 8 to June 4, 2023 | Lamar Advertising | 149,290 devices tracked; exposed versus non-exposed; 7-day lookback from exposure to visit. | Published as “more consumers exposed to the OOH campaign visited”; the page text states the design and the lookback, not a lift percentage. |
| Convenience store chain, digital taxi and rideshare tops | Published by the operator (Firefly); measured by Foursquare Attribution | Foursquare Attribution, which ingests proof-of-play logs to compare exposed and control visitation. | 9.31% behavioural lift, 227,000 incremental visits; highest lifts mid-week, 38% Wednesday and 19% Thursday. Published May 5, 2023. No base rates or sample per cell. |
All three are matched-device designs on a location panel, none is a geo experiment, and each publishes a relative figure. The Clear Channel row states a confidence level and separates total lift from the incremental share. The Lamar row matters for its lookback, 7 days, which caps how long after exposure a visit can count; the IAB says attribution windows “vary widely by category”, and two results with different windows are not the same measurement.
What should an incrementality test brief ask for?
The design, the exposure definition, the sample and the window, in writing, before the flight. A test designed afterwards has no pre-period and nothing left to hold out.
- The outcome and its window. Visits, sales, installs or sessions, and the lookback in days inside which an outcome counts.
- The design, and why. Geo holdout, matched market, synthetic control or pre/post. If markets are held out: how many, chosen how, with what pre-period correlation.
- The exposure definition. Device observed in a viewshed, or market-level assignment. For vehicle-based inventory, the route log that defines the viewshed and who supplies it.
- The control construction. The matching variables, and whether a synthetic control is blended with a geographic holdout as the IAB suggests.
- MDE, power and confidence, and the sample per cell they imply. Ask for the arithmetic.
- The pre-period and contamination. Weeks of baseline before the flight; how people who commute between control and test markets, or pass both a held-out and a live screen, are treated.
- Base rates in the report, for every metric, so absolute and relative lift can both be computed; and who runs and pays for the study, on the record.
The OAAA’s 2020 checklist already asks most of this: how baseline and lift are determined, whether a control group is used, the seasonality considerations, how time of day and day of week are factored in, whether there is panel bias or normalisation to census. Six years on, still the right list.
Where to go next
The category these tests measure is laid out in the guide to out-of-home advertising. Programmatic buys report delivery by screen and hour, the exposure log a device-panel study needs; direct buys usually reconstruct one, covered on programmatic DOOH. Why a vehicle’s exposure record is a route rather than a site is on location-based advertising.
Frequently asked questions
- What is incrementality testing?
- A test that measures how much of an outcome, such as store visits, sales or app installs, a campaign caused, by comparing what happened with a counterfactual: what would have happened without it. The counterfactual comes from a control that did not get the campaign, whether a held-out region, a synthetic model of one, or matched devices that were not exposed.
- How do you calculate incrementality?
- Incremental outcome equals observed outcome minus the counterfactual. When the test compares groups, incremental lift is the exposed group’s rate minus the control group’s rate, and relative lift divides that difference by the control rate. The OAAA’s 2020 guide defines lift as the percent difference in visitation rates between exposed and unexposed audiences, and incremental lift as that figure with the known baseline visiting quantity removed.
- How do you prove incrementality?
- With a control the campaign could not have influenced, set before the flight starts. The strongest proof is random assignment: regions randomised to treatment and control in a geo experiment, or individuals randomised at an ad server. Where randomisation is impossible, as in out-of-home, a synthetic control built from pre-period data is the accepted substitute, and its matching variables should be stated.
- What is the difference between incrementality and lift?
- Lift is the gap between an exposed group and a control group on any metric. Incrementality is lift on a behaviour, measured against a counterfactual designed to isolate cause. Brand lift is lift on attitudes measured by survey. A visitation lift from a matched-device study is an incrementality estimate; a lift from a pre/post comparison with no control is only a change.
- What is a geo holdout test?
- A test in which some geographic regions receive the campaign and comparable regions do not, so the difference in outcomes between the two sets estimates the campaign’s effect. Google’s 2011 geo experiment method randomises non-overlapping regions to treatment or control. Measured’s 2026 guidance is 10 to 15 matched markets, 95% or higher historical correlation, and a 4 to 8 week run.
- What is a matched market test?
- A geo test in which test and control markets are chosen for similarity on the traits that matter to the outcome, rather than randomised. The IAB’s 2025 DOOH guide describes dividing markets into test and control groups based on shared traits and prioritising traits by campaign goal, and lists it beside synthetic control as one of two primary methods for DOOH incrementality.
- Can incrementality be measured for out-of-home advertising?
- Yes, with two constraints. Exposure cannot be randomised per person, so it is either assigned by geography or inferred from device location inside a screen’s viewshed. And the exposure record is only as complete as the location panel behind it, which is why the IAB suggests blending synthetic controls with geo holdouts to limit contamination of the control group.
- How is incrementality testing different from A/B testing?
- An A/B test randomises individuals between two versions and needs a system that can assign and log each person. Incrementality testing asks a different question, whether the campaign caused an outcome at all, and in offline media usually cannot randomise individuals, so it randomises regions or builds a statistical counterfactual instead.
- How long should an incrementality test run?
- Long enough to cover the purchase cycle and to reach the sample the minimum detectable effect needs. Measured reports most geo tests run 4 to 8 weeks with 4 often sufficient, plus about 7 days for conversions to settle; Haus suggests extending the window 4 to 8 weeks beyond exposure for considered purchases. A device-panel visitation study is bounded by its lookback window, 7 days in Lamar’s Cook Out study.
Sources
- Measuring Ad Effectiveness Using Geo Experiments (Vaver and Koehler) · Google Research, 2011
- GeoLift Methodology: incrementality, quasi-experiments, synthetic control methods, power calculations · Meta Open Source, Accessed September 14, 2026
- OOH Performance Measurement: The Net Lift Study, exposure computation, synthetic control, 95% confidence · Reveal Mobile, Accessed September 14, 2026
- Out of Home Advertising: Measurement and Analytics Guide for Agencies and Advertisers (lift definition, attribution solutions matrix, measurement checklist) · Out of Home Advertising Association of America (OAAA), Data Use and Analytics Committee, March 2020
- Digital Out-Of-Home (DOOH) Measurement Guide, Chapter 6 (attribution windows, foot traffic lift) and Chapter 7 (synthetic control, matched market and matched store tests) · Interactive Advertising Bureau (IAB), July 2025
- Glossary: impressions, circulation, Likelihood to See, Visibility Adjustment Index · Geopath, Accessed September 14, 2026
- Ghost Ads: Improving the Economics of Measuring Ad Effectiveness (Johnson, Lewis and Nubbemeyer; working paper, later Journal of Marketing Research 54(6), 2017) · NBER conference paper, February 18, 2016
- When Incrementality Testing Is Not the Right Method (and When You Are Actually Ready for It): MDE of 2 to 5%, 10 to 15 matched markets, holdout size, test length · Measured, July 15, 2026
- Can You Measure the Incrementality of Out-of-Home (OOH) Marketing? (synthetic control from remaining DMAs) · Haus, September 26, 2025
- Incrementality testing frameworks for health brands: detect a 10 to 20% lift at 80% power, 4 to 8 week test window · Haus, Accessed September 14, 2026
- OOH FAQ: Billboard and Transit Media Measurement and Attribution (geofenced viewsheds, MAIDs) · OUTFRONT Media, May 21, 2024
- OOH delivers incremental visits to fast casual restaurant chain (RADARProof; Cuebiq, August 2021; lift confidence level 90%) · Clear Channel Outdoor, Accessed September 14, 2026
- RADARProof: comparing the actions of consumers exposed to OOH campaigns with those not exposed · Clear Channel Outdoor, Accessed September 14, 2026
- Cook Out Case Study: 149,290 devices, 7-day lookback, exposed versus non-exposed · Lamar Advertising, Campaign May 8 to June 4, 2023; accessed September 14, 2026
- Omnichannel Attribution and Store Visit Measurement (OOH attribution from proof-of-play logs) · Foursquare, Accessed September 14, 2026
- Firefly and Foursquare Drive 9.31% Behavioral Lift for Major Convenience Store with DOOH Campaign · Firefly, May 5, 2023
- Measurement & Attribution: brand study, foot traffic, web conversions, app conversion, retargeting · Firefly, Accessed September 14, 2026
About the author
Ercan Bozkurt writes the methodology pages on this site: how out-of-home inventory is counted, measured and bought, with a particular interest in screens that move. He works from the primary document in each case, the transit authority's own rulebook, the measurement body's own definitions, the company's own filing, and writes the page so that a media planner can check every figure against it.
Ercan Bozkurt on LinkedIn · Reviewed by Melih Koray, who publishes this site · How this site is edited and what it refuses to publish: about this site.