Zum Inhalt springen
Chandler Nguyen
KI7 Min. Lesezeit

How to govern marketing data for AI

Governing marketing data for AI comes down to four decisions you make before the model sees anything: what it may see, where it may run, who owns the output, and how long you keep it. Teams that build the memory layer first and govern later end up rebuilding it under a lawyer's supervision.

Governing marketing data for AI comes down to four decisions you make before the model sees anything: what it may see, where it may run, who owns the output, and how long you keep it. Make those decisions before you build a memory layer and they shape the architecture. Skip them and you will rebuild the layer later, usually under a lawyer's supervision, after something has already gone wrong.

I have built memory into systems and I have watched teams skip this step in the name of speed. The pattern is consistent: the model sees more than anyone intended, a question gets asked in a review, and a rushed retrofit follows. Governance is not the tax on the interesting work; it is the precondition for doing the interesting work safely.

The four decisions

What may the model see? Not every piece of marketing data is equal. Public research, internal performance data, customer PII, and regulated categories all carry different obligations. Decide the classification before the model touches anything, because once data is in a context window, "unseeing" it is not a real option.

Where may it run? A model that runs inside your environment is a different risk from one you send prompts to over an API, and both are different from a public tool a team member pastes into. Decide which classes of data may go to which kind of model, and make the answer enforceable rather than aspirational.

Who owns the output? When the model produces a plan, a piece of copy, or an audience definition, someone has to own it. Governance means naming that person, and being clear that accountability does not transfer to the tool. This is the same principle as the review standard in the execution lane.

How long do you keep it? Retention is the decision teams defer and later regret. If context improves the model, you want to keep it. If it includes personal or regulated data, you may be obliged to delete it. Decide retention by data class, and make deletion real.

Classify the data first

The classification is the foundation for the other three decisions, so do it first and keep it simple. A workable scheme is four buckets: public, internal, confidential, and regulated. Public data can go almost anywhere. Internal data — performance, spend, strategy — stays with approved models. Confidential data — client-specific context, unreleased plans — stays inside controls. Regulated data — anything touching personal or legally protected information — gets the strictest handling and often never enters a general model at all.

The value of the classification is that it makes the rules legible to everyone. A marketer who knows the bucket knows what they can paste into which tool, without needing a policy memo for every case. Simple beats comprehensive here, because a rule nobody remembers is not governance.

Where the model runs

The hosting decision is where most of the real risk sits. A model deployed inside your own environment with your own controls is the most contained option and the most work to run. A model accessed through a vendor API under a data-processing agreement is the common middle path, and the terms matter enormously — whether inputs train the vendor's models, where data is stored, and what the deletion guarantees are.

And then there is the public-tool case, which is the one to close explicitly. If team members can open a browser and paste client data into any tool, none of the rest of this matters. Name the approved tools, block or discourage the rest, and explain why. The ungoverned path is almost always the convenient one.

Who owns the output

Every AI-assisted deliverable needs a named owner, and the owner is a person, not a workflow. That is not a technicality; it is what makes the output defensible. If a model drafts a media plan and no one signs off, the agency or team has effectively published a machine's guess. Governance assigns the signature.

This connects directly to purchasing. When you buy an agency's AI-native service, the contract should say who owns the output and who is accountable when it is wrong. Ambiguity there is how a quality incident becomes a legal one. The governance answer and the review standard are the same discipline applied to different risks.

Retention and deletion

Retention is where the compounding benefit of memory and the obligation to delete collide, and there is no way around thinking it through. Context makes the model useful, so you want to keep the briefs, results, and decisions. Some of that same data may be subject to deletion requirements or contractual limits. The resolution is retention by class: keep the derived, de-identified context indefinitely; keep raw personal data only as long as you are allowed to; and make the boundary between them explicit.

The reason to do this before you build is that retrofitting deletion into a memory layer is painful. If raw data and derived context are mixed together, separating them later is a migration. Design the separation in from the start. This is the practical heart of why AI without memory is just an expensive chatbot — the memory is the value, and it is also the governance surface.

A worked classification

Take a typical campaign and sort the inputs. The category research and the competitor landscape are public — they can go anywhere. The client's spend, performance history, and internal strategy are internal — approved tools only. The unreleased creative concepts and the client's customer list are confidential — keep them inside controls. And anything that identifies a person, or falls into a regulated category, gets the strictest handling and often stays out of general models entirely.

Once that sort exists, the rules write themselves. Public data can go into public tools. Internal and confidential data go only into contracted, approved models. Regulated data goes nowhere without a specific review. The classification is what turns a vague policy into a decision a marketer can make in ten seconds, which is the only kind of governance that survives contact with a busy week.

Vendor terms that matter

When you evaluate a model provider, four terms decide most of the risk. Does the provider train on your inputs, or are your prompts excluded from training? Where is the data stored and processed, which matters if you have data-residency obligations? What are the deletion guarantees and how do you enforce them? And what happens to your data if the contract ends? Ask for the answers in writing, because the marketing website will not tell you.

Most enterprise vendors now offer a no-training option and reasonable deletion terms; the point is to confirm it rather than assume it. For a team that pastes client data into a chat window, none of these terms apply, which is exactly why the approved-tool list has to be explicit and short.

The order of operations

Classify the data, decide where each class may run, name the owner of the output, and set retention by class. Do this in that order, before the first prompt that touches real data. Then write it down where the team can find it, and make the approved tools the easy path. Governance that relies on people remembering a rule is not governance; it is hope.

FAQ

Do we need a formal policy or is a team agreement enough?

Start with a one-page decision record covering the four questions, and formalise it later if your scale or regulator requires it. A short, enforced agreement beats a long, ignored policy. The point is to decide, not to produce a document.

Can we use public AI tools for internal marketing data?

Only if the internal data is not confidential or regulated and the tool's terms allow it. For most teams the honest answer is to route internal and client data through approved, contracted tools and keep public tools for public or synthetic data.

How do we set the data boundary itself?

By classification, not by tool. Sort every input into one of four buckets — public, internal, confidential, regulated — and let the bucket decide where it may go: public data to public tools, internal and confidential to approved contracted models, regulated data nowhere without a specific review. The boundary is the classification, and once it exists a marketer can make the call in ten seconds without a policy memo.

What if we cannot tell which bucket something belongs in?

Default to the stricter bucket and put it on a short review list rather than guessing. Ambiguous data is usually personal, client-specific, or regulated, which is exactly the data you do not want to discover was mishandled later. A two-line list of the cases people actually flag, reviewed monthly, resolves the grey areas without slowing the work down.

The short version

Govern marketing data for AI with four decisions made before the model sees anything: what it may see, where it may run, who owns the output, and how long you keep it. Classify first, host deliberately, name an owner, and set retention by class — then build the memory layer on top of that. The Execution guide covers the rest of the operating questions, and the for in-house teams track works governance through to measurement.

If you have set this up on your own team, I would like to hear which of the four decisions was hardest to get agreement on.

Cheers, Chandler