本文へ移動
Chandler Nguyen
AI6分で読めます

How to measure AI marketing impact

Measure decisions changed and cycle time, not output volume. Output is a vanity metric once the model makes it cheap. The signals that matter are how fast work moves, how much review it takes, how often the AI-assisted work is actually better, and what business outcome followed.

Measure decisions changed and cycle time, not output volume. Output volume is a vanity metric the moment the model makes production cheap: a team can triple its asset count and learn nothing, or ship the same number of assets and learn a great deal. The signals that actually indicate AI is working are how fast a decision moves, how much review it needs, how often the AI-assisted version beats the old one, and what business outcome followed.

I have watched teams celebrate a jump in content output as proof that their AI programme was working, when the output was the least interesting part of the change. Measurement is where an AI programme either gets honest or gets lost, so this is worth getting right before you scale anything.

The trap: measuring what got easy

The default instinct is to measure throughput. How many assets did we produce, how many markets did we launch, how many reports did we ship. These numbers go up under AI, and they are satisfying in a deck, but they measure the thing AI made cheap rather than the thing you are trying to improve.

Throughput can rise for reasons that have nothing to do with value. More variants can mean more noise if nothing is tested. More reports can mean more meetings nobody needed. More markets can mean more places you are mediocre. The question a good measurement frame answers is not "did we produce more?" but "did more production turn into better decisions and better outcomes?" Those are different questions, and only the second one matters.

What to measure instead

Measure the operating model, not the production line. Four families of signals cover most of it. Speed: cycle time from brief to live, and time to a decision. Effort: senior review time per output, which is the real cost of an AI-native workflow. Quality: error rate, revision rate, and how often the AI-assisted work needs rework. And outcome: cost per result, and the business metric the work was supposed to move.

None of these is exotic, and none requires new tooling. What makes them different from the old metrics is that they measure judgement and flow rather than volume. A team that cuts cycle time, holds error rate flat, and keeps review time sane has genuinely improved its operating model, even if the asset count did not change.

Leading and lagging signals

Split the signals by timing. The lagging signal is the business outcome — revenue, pipeline, brand lift, whatever the campaign was for — and it is the one that matters, but it arrives late and is influenced by many things. The leading signals are the operating ones: cycle time, review time, revision rate, decisions per week. They move first and they are under your control.

You need both, and you need to be honest about attribution. It is tempting to claim the lagging number for AI when the market moved. The honest frame is to use leading signals to prove the operating change, and to use the lagging signal with proper attribution — incrementality, holdouts, a test design — to argue the business effect. The reporting discipline is what keeps those two from being conflated.

The quality counter-metric

Every speed metric needs a quality counter-metric, or the team will optimise speed by shipping worse work. If cycle time is the headline, error rate and revision rate are the guardrails. If review time is the efficiency story, the counter is whether escaped defects — problems the client or the customer found — went up or down.

This is the single most important discipline in measurement for AI work, because AI makes speed easy and quality hard. A team that reports faster cycles without a quality counter-metric is not measuring its operating model; it is measuring its ability to move work out the door, which is exactly the behaviour that produces review debt.

A simple measurement frame

Pick one primary operating metric — I would start with cycle time on the workflow you redesigned — and two counter-metrics, error rate and review time. Add one business outcome the workflow is meant to influence, measured with a design that can attribute it honestly. Report all four together, every month, and treat any move in the primary that comes with a worse counter-metric as a failed improvement, not a win.

That frame is deliberately thin, because thin frames get used. A dashboard with twenty metrics gets ignored. Four numbers that force the speed-versus-quality trade-off to the surface are enough to run the operating model, and they make the argument for expansion with evidence rather than enthusiasm.

Instrumentation before expansion

The hardest part of measurement is having the numbers before you need them. Most teams expand an AI workflow first and try to reconstruct its impact later, which is how a credible programme ends up with an unprovable claim. The discipline is to instrument the workflow at the moment you build it: log the cycle time, record the review step, and capture the error rate from the start.

That sounds like overhead, and early on it is. But the alternative is a year in with no baseline and no way to tell whether the workflow improved or just changed. A simple log — when a brief entered, when it went live, how many review passes, what was wrong — is enough. You do not need a platform to start; you need the habit of recording the four numbers before you scale the thing that produces them.

A worked example

Imagine a paid social creative loop. The old cycle time was measured in days: brief, design, review, revisions, launch. After moving to a model-first workflow, you expect that to fall. But if review time doubled and revisions did not, you have not improved the operating model; you have moved the bottleneck.

The honest read comes from looking at all four numbers together. Cycle time down, review time flat or slightly up, error rate flat, and the test-readout quality up — that is a genuine improvement. Cycle time down, review time up sharply, and revisions up — that is a team shipping faster and worse, and the counter-metrics caught it. Without them, the cycle-time chart alone would have looked like a triumph.

Common measurement mistakes

Three mistakes recur. Measuring volume and calling it impact, which flatters the programme and teaches nothing. Dropping the quality counter-metric, which lets speed hide a decline in the work. And attributing a market movement to AI without a design that supports the claim, which works until someone asks how you know.

The fix for all three is the same discipline: pick the operating metric, keep the counter-metric attached, and be explicit about what you can and cannot attribute. A modest, defensible claim beats a large, unfalsifiable one, because the defensible claim survives the next quarter and the large one does not.

What to do on Monday

Take the workflow you have moved to AI and write down its cycle time, its review time, and its error rate for the last month. If you cannot, that is the first finding: you are running an AI-native workflow without the instrumentation to tell whether it is better. Then pick one business outcome and decide how you would credibly attribute a change in it. Do that before the next campaign, so the measurement exists by design rather than as a reconstruction.

FAQ

Should we count hours saved?

Hours saved is a useful internal signal but a dangerous external one, because it depends on assumptions about what people did with the time. Count cycle time and review time, which are observable, and let hours fall out of those rather than leading with them.

How do we prove AI caused the outcome?

With a design set up before the work: a holdout, a matched control, or a test where one condition is AI-assisted and one is not. Without that, you can show correlation and improvement, but you cannot claim causation. Set the design up front or accept the weaker claim.

What if the business outcome moves for unrelated reasons?

Say so. Use the operating metrics to prove the workflow change and be candid that the business movement is shared with the market. Overclaiming is how an AI programme loses credibility when the next quarter is flat.

How often should we review these metrics?

Monthly is enough for a stable workflow, and weekly when you are actively tuning it. The point is the cadence of the conversation, not the dashboard. A standing monthly review with the four numbers is the whole system.

The short version

Measure decisions changed and cycle time, not output volume. Use speed, effort, quality, and outcome signals, pair every speed metric with a quality counter-metric, and attribute business outcomes honestly. The Execution guide covers the rest of the lane, and the for team leads track works measurement through the operating model.

If you are measuring your AI programme, I would like to hear which number you found hardest to defend — that is usually the honest one.

Cheers, Chandler