Holdout test measurement

How to Measure Real Advertising Incrementality with Holdout and Geo Experiments

Advertising reports can show thousands of conversions without answering the question that matters most to a business: how many of those sales, leads or subscriptions happened because of the advertising? A customer may click a paid search advert shortly before buying, yet that customer might already have known the brand and intended to make the purchase. Attribution will usually give advertising some credit, while incrementality measurement asks what would have happened if the advertising had not been shown. Holdout and geographic experiments are among the clearest ways to answer this question because they create a realistic comparison between advertising activity and a credible alternative without it. In 2026, this distinction is particularly important as marketers combine automated bidding, several advertising channels, first-party data and offline sales. Instead of relying only on reported conversions or return on ad spend, businesses can use controlled experiments to estimate additional sales, revenue or profit genuinely caused by their advertising.

Why Attributed Conversions Are Not the Same as Incremental Growth

Standard advertising reports are useful for monitoring campaigns, but attribution and incrementality measure different things. Attribution assigns credit for an observed conversion to one or more marketing interactions. If someone clicks an advert, returns later and completes an order, the advertising system may record a conversion according to its attribution rules. Incrementality asks a stricter question: would that order still have happened without the advertising exposure? This difference is important for well-known brands, existing customers and high-intent search traffic. People who are already planning to buy can be relatively easy for advertising algorithms to identify. A campaign may therefore produce an impressive attributed conversion total while part of that activity represents demand that already existed.

Consider an online retailer that records £500,000 of revenue attributed to a paid advertising campaign. Treating the entire £500,000 as advertising-generated revenue would assume that none of those customers would have purchased otherwise. In practice, some may have reached the retailer through an organic search, a bookmark, an email, a physical shop or a direct visit if the advertising had disappeared. An experiment creates a counterfactual estimate — a practical approximation of what would have happened without the campaign. If comparable customers or regions without the advertising would have generated £420,000 during the same period, the estimated incremental revenue is closer to £80,000 than £500,000. That difference can completely change the assessment of advertising efficiency.

This does not make attribution useless. Day-to-day campaign management still needs fast signals such as conversions, cost per acquisition and attributed revenue. The problem begins when those metrics are treated as proof of causality. A strong measurement process uses attribution for operational decisions and controlled experiments for larger questions about budget value. For example, attribution can help identify which creative or keyword is producing more recorded conversions this week, while an incrementality study can test whether the wider campaign creates additional business at all. The two approaches therefore serve different purposes. The strongest decisions come from understanding where they agree, where they disagree and why.

What a Holdout Experiment Actually Measures

A holdout experiment creates a group that is deliberately withheld from the advertising being evaluated. Ideally, eligible people are assigned to a treatment group and a control group before advertising exposure. The treatment group can receive the campaign as normal, while the control group is prevented from receiving the relevant advertising wherever the test design allows. At the end of the study, the business compares outcomes such as purchases, subscriptions, qualified leads or revenue. Because assignment happens before the advertising can influence behaviour, differences between the groups provide much stronger evidence of causality than simply comparing people who happened to see an advert with people who did not.

The business question must be defined carefully before the test begins. A company testing the incremental value of one advertising channel does not necessarily need to stop every other marketing activity for the control group. Instead, both groups should continue to experience normal business conditions, with the tested activity being the main intentional difference. If the question is whether a particular paid social campaign adds sales beyond search, email and other existing marketing, those other activities should normally remain comparable across the groups. If the question concerns the incremental value of an entire media programme, the required holdout is broader. The design should match the decision that will be made after the results arrive.

Holdouts also involve a genuine business cost because some potential customers are intentionally excluded from advertising. That is why the control group should not simply be made as large as possible. A larger holdout can improve measurement precision, but it can also mean giving up more potential revenue during the test. A very small group reduces this commercial cost but may provide too little information to detect a modest advertising effect. Modern lift tools therefore use historical conversion volume, expected effect size, campaign budget and study duration to assess whether an experiment has enough statistical power. The practical objective is to create a control group large enough to answer the question reliably without withholding more advertising than necessary.

How to Design a Holdout Test That Answers a Business Question

A useful holdout study begins with a business outcome rather than a convenient advertising metric. Click-through rate, impressions and video views can explain how people interact with advertising, but they rarely provide the strongest measure of commercial incrementality. For an ecommerce business, completed orders, net revenue or contribution margin may be more meaningful. A subscription service might examine paid subscriptions rather than account registrations, while a lead-generation business may prefer qualified leads or completed contracts to raw form submissions. The closer the selected KPI is to the financial result that management actually cares about, the easier it becomes to use the experiment when making budget decisions.

Measurement should also be agreed before the campaign starts. Changing the primary KPI after seeing preliminary results creates a risk of selecting whichever number looks most favourable. The same applies to the audience, conversion window, test period and treatment definition. Teams should record what is being tested, which outcome determines success and what action will follow different possible results. For example, the plan could state that a clearly positive incremental return supports continued investment, a near-zero result triggers a review of targeting and creative, and an uncertain result leads to another test rather than an immediate budget cut. This makes the experiment a decision-making process rather than a search for a desirable statistic.

Business-as-usual conditions should remain as stable as reasonably possible during the experiment. Large price changes, major promotions, stock shortages, a redesigned checkout or an unexpected change in another advertising channel can affect conversions independently of the campaign being tested. Not every real-world change can be avoided, but important changes should be documented so that the result can be interpreted correctly. It is particularly risky to redesign targeting, replace most creative assets or substantially alter budgets halfway through a study unless those changes are part of the planned treatment. A test that continually changes what it is testing may still produce numbers, but those numbers become much harder to turn into a clear business lesson.

Reading Lift, Incremental Revenue and iROAS Without Overclaiming

The first result to understand is absolute incremental outcome. If the treatment group produces more conversions than would reasonably be expected without the advertising, the difference represents estimated incremental conversions. The result can also be expressed as percentage lift, which shows the improvement relative to the control expectation. Neither figure should be considered in isolation. A 10% lift could represent substantial value for a high-volume retailer but relatively few extra transactions for a small campaign. Conversely, a modest percentage lift applied to a large revenue base can justify significant investment. Reporting both the percentage effect and the number of additional business outcomes gives decision-makers a much clearer picture.

Incremental return on advertising spend, often shortened to iROAS, is especially useful when the objective is revenue. Unlike conventional ROAS, which divides attributed revenue by advertising spend, iROAS uses revenue estimated to have been caused by the advertising. Suppose a campaign costs £50,000, advertising reports attribute £200,000 of revenue to it, but a holdout study estimates that only £75,000 was genuinely incremental. Conventional reported ROAS would be 4.0, while incremental ROAS would be 1.5. Neither number automatically shows whether the campaign is profitable because product margins and other costs still matter, but the incremental figure provides a more realistic starting point for that calculation.

A credible experiment should also communicate uncertainty. The result is an estimate based on observed data, not a guarantee that the exact same lift will occur every month. A study may indicate a positive effect but still leave a wide range of plausible values when conversion volume is low or the underlying behaviour is volatile. This is why marketers should avoid turning a weak directional result into a precise claim such as “advertising generated exactly £87,450 of extra revenue”. A better interpretation reports the estimated effect together with the degree of certainty and considers whether the possible range would change the business decision. If several plausible values lead to the same budget decision, perfect precision is less important. If the decision changes depending on small variations in the estimate, further evidence may be worthwhile.

Holdout test measurement

When Geo Experiments Are the Better Choice

User-level holdouts are not always practical. Some advertising channels do not provide a reliable way to prevent selected individuals from being exposed, and businesses may need to measure outcomes that cannot easily be linked to individual advertising exposure. Retail sales, call-centre orders, dealer transactions and other offline activity can make geographic testing particularly useful. Instead of separating individual customers, a geo experiment separates locations. Advertising can be reduced, stopped or increased in selected test regions while comparable regions provide the control. The resulting change in sales or another KPI is then evaluated against what would reasonably have happened in the absence of the advertising change.

A simple comparison between two arbitrary cities is rarely sufficient. Birmingham and Manchester, for example, may have different customer bases, competitive conditions, sales trends and seasonal patterns. A credible geo experiment uses historical data to identify locations or combinations of locations that behaved similarly before the intervention. The aim is not to find two areas that look identical on a map, but to establish a control that provides a believable estimate of the test region’s normal trajectory. Modern geo-testing methods can combine information from several control regions when no single location provides a good match. This makes geographic experiments useful even when markets differ considerably in size.

Geo tests can take several forms depending on the business question. A holdback or “go dark” design reduces or removes the tested advertising in selected areas while maintaining normal activity elsewhere. A heavy-up test moves in the opposite direction by increasing spend in chosen regions and comparing the additional outcome with control areas. The first approach is useful when the business wants to determine how much existing activity would be lost without advertising. The second can answer whether additional budget is capable of producing additional growth. These are related but not identical questions. A channel that protects existing sales effectively at the current budget may still have limited room to scale, while another channel may show stronger incremental returns when spend is increased.

Practical Rules for Reliable Geo Tests in 2026

Historical stability is one of the most useful indicators when selecting test and control regions. Before making any advertising change, analysts should examine whether the areas follow sufficiently similar movements in the chosen KPI over time. One unusually good week is not enough. The comparison should cover a representative pre-test period that includes ordinary variations in demand. Businesses should also check whether planned events will affect regions differently. Local holidays, shop openings, sporting events, extreme weather, competitor activity or regional promotions can create changes that have nothing to do with advertising. These factors do not automatically invalidate a study, but known differences should influence market selection and the interpretation of results.

Geographic leakage deserves particular attention. Consumers do not remain inside advertising boundaries simply because an experiment uses them. Someone may live in a control region, work in a treatment region and see advertising there. Another customer may see an advert in one city and complete a purchase in another. National television, influencers, organic social activity and other broad-reach marketing can also affect both groups. Advertising systems that support geo experiments increasingly account for this type of contamination when selecting experimental areas, but the business still needs to consider real customer behaviour. Areas with heavy cross-border travel or closely integrated local economies may be poor choices for a clean comparison.

The most valuable practice is to treat incrementality testing as a repeated measurement programme rather than a one-off attempt to certify that advertising “works”. Advertising effects can change as budgets grow, audiences become saturated, competitors react, products mature and consumer demand shifts. A holdout study may show that a channel creates incremental sales at one spending level without proving that twice the spend will create twice the result. Geo experiments can then test the effect of higher or lower investment, while later studies can examine different campaigns or periods. In 2026, this combination of controlled experimentation, attribution and broader marketing measurement gives businesses a more defensible basis for allocating budgets. The goal is not to replace every advertising metric with one lift number, but to determine which spending genuinely changes customer behaviour and by how much.