Chuyển đến nội dung
Chandler Nguyen
AI7 phút đọc

Measuring marketing incrementality with AI

Incrementality measures what would not have happened without the marketing. AI helps design tests, estimate effects from more data, and read who the effect is concentrated in. It does not remove the need for a control. Without a counter- factual, a model just makes a correlation sound causal.

Incrementality is the measure of what would not have happened without the marketing. AI helps in three places: designing better tests, estimating effects from more and messier data, and reading which parts of the audience or the mix the effect is concentrated in. It does not supply the counterfactual. Without a proper control, a model will make a correlation sound causal with great confidence.

That last sentence is the whole risk. AI makes it easier to produce a number that looks like incrementality, and hardest part of the job is still deciding what to compare against. Teams that skip that step and jump to the model end up with elegant estimates they cannot defend when someone asks the obvious question: compared to what?

What incrementality actually means

Incrementality answers a counterfactual question: how many conversions, sales, or brand outcomes occurred because of the marketing, rather than happening anyway. The people who would have bought without seeing the ad are not incremental, however much the ad touched them. Attribution counts credit; incrementality measures causation, and they are different questions that often disagree.

This is why attribution data alone cannot establish incrementality. A last-click report will give credit to the ad someone clicked, but it cannot tell you whether they would have converted regardless. To know that, you need a group that did not receive the marketing and a way to compare. That group is the counterfactual, and no amount of model sophistication replaces it.

Why attribution is not incrementality

Attribution and incrementality answer different questions, and confusing them produces confident, wrong budget decisions. Attribution allocates credit across touchpoints for conversions that happened; incrementality estimates how many of those conversions depended on the marketing. A channel can look strong in attribution and be nearly zero incremental, or look weak and be carrying the campaign.

The practical implication is that incrementality should carry more weight than attribution for budget decisions, and attribution is useful mainly for in-flight optimisation. When the two disagree, trust the experiment. That is the discipline set out in the wider measurement frame in how to measure AI marketing impact, and it is even more important when AI is generating the analysis.

AI's three jobs in incrementality

AI helps at three points, and it is worth being precise about each, because the value and the risk are different at every one. Design: generating test structures, power calculations, and candidate holdout designs faster than a human, then leaving the choice to a person. Estimation: fitting models that pull signal from large or sparse data where simple comparisons fail. And segmentation: finding where the effect is concentrated, so a blended average does not hide a channel that works only for one audience.

In each case the model is a tool inside a causal design, not a substitute for one. The design, in particular, is where judgement still dominates, because a badly designed test gives a model nothing to estimate. AI can propose designs; a human who understands the business has to choose one that can actually run.

The main methods and where AI helps

Different incrementality methods suit different situations, and AI's role varies by method. The table below maps the common approaches.

MethodWhen it fitsAI's role
Geo holdoutChannel-level, region-splittableDesign, effect estimation, heterogeneity
SwitchbackTime-based, auction-drivenDesign, sequential analysis
Ghost bids / PSASearch and auction channelsEstimation, spend-response curves
Matched controlWhen a holdout is impossibleMatching, bias checks
MMMPortfolio-level, always-onFaster fitting, scenario testing

The more a method depends on a designed counterfactual, the less AI should be trusted to invent the design alone. The more it depends on estimating from data, the more AI adds. Use the table to pick the method for the situation rather than defaulting to the one the model handles most easily.

The danger of confident causation

The failure mode is a model that finds a relationship in observational data and presents it as incremental. Modern methods can sound authoritative while resting on assumptions that do not hold — parallel trends that are not parallel, controls that are contaminated, spend that responds to demand rather than causing it. The output looks like a causal estimate and is really a sophisticated correlation.

The defence is to keep the counterfactual explicit and to test the assumptions. If the design says two groups are comparable, check it. If a holdout is meant to be clean, verify that no other change hit one side. And be willing to report that the data cannot answer the question, which is a more useful answer than a number you cannot defend.

Combining experiment and model

The strongest measurement systems pair a designed experiment with a model that extends its finding. A geo holdout gives you a clean, credible estimate for one channel in a defined window; a model uses that estimate to calibrate a broader view across channels and periods. The experiment anchors causation, and the model provides coverage where you cannot run a test.

This pairing is also how you handle the channels where a holdout is impractical. You trust the model more where it has been calibrated against real experiments and less where it has not. Over time, a library of past experiments becomes the evidence base that keeps the always-on model honest, and it compounds the same way a client-memory layer does.

A practical starting point

Start by fixing the counterfactual question for each major channel: what would you compare against, and can you actually run that comparison? Where you can, run a holdout and treat it as the anchor. Where you cannot, use a matched control or a model, and label the estimate as weaker. Then use AI for design, estimation, and segmentation inside that frame, and report every estimate with its method and its confidence.

Then build the habit of comparing the model's increments to the next experiment. When they diverge, the experiment wins and the model gets recalibrated. That loop is what turns a pile of estimates into a trustworthy measurement system, and it is the same discipline whether or not AI is involved.

Attribution decay and why it matters

The gap between a report and a decision is not new; I wrote about it years ago in don't waste time on reporting, and AI makes the number easier to produce without making it any more causal. Attribution data decays for reasons that have nothing to do with the campaign: privacy changes, identifier loss, measurement gaps, and platform-side modelling that fills the holes with estimates. As decay increases, attribution drifts further from the truth, which is exactly when teams lean on it more because the cleaner signals are gone. Incrementality becomes more valuable as attribution gets weaker, not less.

This is the practical reason to invest in tests now. The counterfactual is the one measurement that does not depend on tracking a person across the internet; it depends on comparing groups. In a world of decaying attribution, the test-based estimate is the durable one, and the model's role is to extend it further into the plan than a single experiment could reach on its own. The corollary is to date your estimates. An incrementality number from two years ago is not a fact about today; the auction, the audience, and the creative have all changed. Re-test on a rolling basis and treat old estimates as priors rather than truths. A measurement system that never expires is a measurement system that is quietly out of date.

FAQ

Can AI tell us incrementality without a test?

No. It can estimate from observational data, and the estimate rests on assumptions that a test would verify. Without a counterfactual, you are reading a correlation and calling it causation, however sophisticated the method. Use AI to support a designed test, not to replace one.

Should we stop using attribution?

No, but demote it. Attribution is useful for in-flight optimisation and for understanding the path to conversion. For budget decisions, incrementality should carry more weight, and when the two disagree, the experiment is the one to trust.

How often should we run incrementality tests?

Enough to anchor the model, typically per major channel on a rolling basis rather than all at once. Each clean experiment recalibrates the always-on view and raises confidence everywhere. A few good tests a year, well designed, beat a constant stream of weak ones.

What if we cannot run a holdout?

Use a matched control, a switchback, or a model, and report the estimate as weaker. Label the method and the confidence clearly, and treat any decision that depends on a weak estimate as provisional. Being explicit about the limitation is more credible than presenting every estimate as equal.

The short version

Incrementality measures what would not have happened without the marketing, and AI helps design tests, estimate effects, and find where they concentrate. It never supplies the counterfactual, so keep a control and label every estimate with its method. The how to measure AI marketing impact piece covers the operating metrics, and the Execution guide places measurement in the wider workflow. The for team leads track works it through the operating model.

If you run incrementality, I would like to hear which method you trust least — that is usually the honest answer.

Cheers, Chandler