본문으로 건너뛰기
Chandler Nguyen
AI7분 읽기

How to pilot AI in a marketing org

A pilot is a time-boxed test of one workflow with a named owner, a baseline, and a decision rule agreed before it starts. The output is a decision — expand, change, or stop — not a demo. Pilots fail when they are built to impress rather than to decide.

How to pilot AI in a marketing org: pick one well-defined workflow, name an owner, record a baseline, and agree the decision rule before anything runs. A pilot is not a showcase; it is a test designed to produce a decision. If you cannot say in advance what result would make you stop, you are not running a pilot, you are running a demo with extra steps.

Most AI programmes that stall do so because they skipped this step. They went from interest to rollout, or from interest to a pilot that was really a launch in disguise. The pilot is where an organisation earns the right to scale, and it only works if it is designed to be able to fail.

Why most pilots fail

The common failure is a pilot built to succeed. The workflow is chosen because it is easy, the scope creeps until it touches everything, and success is defined after the fact as "the team liked it." Nobody wrote down what would have counted as a failure, so nothing did. The pilot ends with applause and no transferable conclusion.

The second failure is the opposite: a pilot so hedged that it tests nothing. It uses synthetic data, runs on a side project nobody cares about, and is reviewed by people with no stake in the outcome. It cannot fail, and it cannot teach you anything, which makes it a waste of the one thing a pilot is for.

Pick one workflow

Choose a workflow that is narrow, repetitive, and measurable, with a clear start and end and a team that already feels the pain. Good candidates are things like first-draft creative production, reporting assembly, or campaign localization. Bad candidates are anything that depends on ten other teams, has no baseline, or is really a strategy question dressed as a task.

The test is whether you can describe the workflow as an input, a process, and an output. If you can, it can be piloted. If you cannot, it needs designing first, not piloting. Starting with the hardest, most ambiguous workflow is the fastest way to conclude that AI does not work, when the real issue was the target.

Record a baseline

A pilot without a baseline cannot prove anything, because there is nothing to compare against. Before the AI workflow runs, capture the current numbers for the few things that matter: cycle time, senior review time, error or revision rate, and the business outcome the workflow feeds. This is the discipline set out in how to measure AI marketing impact, and it has to happen before the pilot, not after.

The baseline does not need to be perfect; it needs to exist. A rough number recorded last week beats a precise one reconstructed next quarter from memory. If you have no baseline and no time to build one, run the old workflow for two more cycles and record it. That delay is cheaper than a pilot whose result nobody can defend.

Decide the rule before you start

Write the decision rule down before the first run: what result expands the pilot, what result changes it, and what result stops it. For example, expand if cycle time falls by a fifth with no rise in error rate; change if speed improves but quality drops; stop if neither moves. The specific numbers matter less than committing to them while you are still unbiased.

This is the step teams skip because it feels bureaucratic. It is the step that makes the pilot honest. Once you have run the workflow and seen a result, every interpretation becomes self-serving; the only defence is a rule written when you did not yet know the answer.

Keep the human in the loop

A pilot that removes the human entirely tests the wrong thing. What you want to learn is whether an AI-assisted workflow can produce better output per unit of senior time, and that requires a review step you can measure. Track how often the human changes the output and how much time the review takes, because the review step is where the real cost of an AI-native workflow lives.

The pilot should also leave room for the operator's judgement. If the workflow forces a rigid path, you will learn whether that path works, not whether the team can do more with better tools. Give the pilot room to be adapted, and record the adaptations — they are the seeds of the operating model you scale.

Design the pilot on one page

Keep the pilot design short enough that everyone can hold it. The table below is the minimum set of decisions to make before you run anything.

DecisionWhat to write down
WorkflowThe one input-output job under test
OwnerA named person, not a committee
ScopeWho is in, who is out, what is excluded
BaselineCycle time, review time, error rate, outcome
Decision ruleThe result that expands, changes, or stops
DurationA fixed end date for the test

If you cannot fill this table, the pilot is not ready. Filling it takes an hour and saves a quarter.

Run it small, then decide

Run the pilot with a small team on real work, for a fixed period — four to eight weeks is usually enough to see whether the workflow moves. Keep the decision rule visible, review the numbers weekly, and resist the urge to expand scope mid-pilot. The point is to learn one clean thing, not to fix everything at once.

At the end, make the decision you wrote down. If it expands, the pilot becomes a rollout with a real baseline behind it. If it changes, you now know which part failed, which is more valuable than a vague success. If it stops, you have saved the organisation a year and earned the credibility to try the next pilot — the honest failure is the cheapest outcome available.

Governance is part of the pilot

Even a small pilot needs a data boundary and a review step, so agree those before the first run rather than discovering them halfway through. This is the lighter end of the policy and governance work, and a pilot is a good place to test which controls you actually need. Controls that prove necessary here are the ones worth scaling; controls that add friction without catching anything can be dropped.

Making the pilot's result transferable

A pilot that teaches one team something nobody else can use has half-failed. From the start, capture what you learned in a form that travels: the workflow as it ended up, the prompts and templates that worked, the failure modes you found, and the decision you made. Treat the pilot's output as two things — a decision and a reusable kit.

That kit is what turns a pilot into a rollout. The next team does not start from zero; it starts from a working design and adapts it to its own work. The kit only works if the next team can find and read it, which is why a memory layer matters as much as the workflow itself — the point I make in AI without memory is just an expensive chatbot. This is also where a centre of excellence, if you have one, earns its place, by holding the kit and the standards that make it usable. A pilot without a transferable output is an experiment that has to be repeated, which is the most expensive kind.

A counter-metric for the pilot

Speed alone will flatter almost any pilot, so pair it with a counter-metric that catches the hidden cost. Track how often the human changes the output, how long the review takes, and whether the error or rework rate moves at all. A workflow that halves drafting time while doubling review time has not saved anything; it has moved the work to a more expensive person. The honest measure is output per unit of senior time, with quality held, and the only way to see it is to record the review step as carefully as the production step. If the pilot improves the headline number and worsens the counter-metric, you have found the real question the rollout has to answer.

FAQ

How long should a pilot run?

Four to eight weeks on a narrow workflow is usually enough to see whether cycle time and quality move. Shorter and the numbers are noise; longer and you are running a rollout without having decided to. Pick the end date at the start.

What if we have no baseline?

Run the current workflow for two more cycles and record the numbers before you start. It feels like a delay, but a pilot without a baseline produces a result you cannot defend, which is worse than waiting two weeks.

Who should be in the pilot?

A small team that already does the work and feels the pain, plus the owner who will make the decision. Keep leaders informed rather than embedded; their presence changes behaviour and muddies the result. The pilot should reflect the real work, not a performance for management.

What is the best first workflow?

The most repetitive, well-defined, painful one. Reporting assembly, first-draft production, and localization are common good starts. Avoid anything that needs ten teams to align, because you will spend the pilot on coordination rather than learning.

The short version

Pilot AI by testing one workflow with a named owner, a baseline, and a decision rule agreed in advance. Run it small on real work, keep a human in the loop, and make the decision you wrote down. The Execution guide shows where the pilot fits in the wider workflow, and the for team leads track works through the pilot and the operating model it leads into.

If you have run one of these, I would like to hear what your decision rule was — and whether you actually followed it.

Cheers, Chandler