Saltar al contenido
Chandler Nguyen
IA7 min de lectura

The AI marketing stack for agencies

An agency AI stack is not a list of tools. It is five layers — client memory, orchestration, production, review, and measurement — assembled around how the agency already works. Buy the commodity layers, build the one that holds your client knowledge, and never let the stack outgrow the operating model.

An agency AI stack is the five layers an agency runs its AI work on: client memory, orchestration, production, review, and measurement. It is not a folder full of subscriptions. The stack exists to make the agency faster without making it less accountable, and the only layer worth building rather than buying is the one that holds what the agency knows about its clients.

I have watched agencies buy their way into an AI story and end up with a tool list that impresses in a pitch and confuses in delivery. The stack question is really an operating-model question wearing a technology costume, so it is worth separating the two before you spend.

Why the tool count is the wrong frame

The default instinct is to count tools. Twelve platforms, five models, a design tool, a scraping tool, a reporting tool. Tool count is easy to present and mostly meaningless, because two agencies with the same subscriptions can have completely different results. One has wired them into a workflow with a single source of client truth; the other has twelve tabs and no memory.

The stack that matters is defined by what flows between the layers, not by what sits in each one. If a brief enters once and produces a plan, an activation, a report, and a learning that feeds the next brief, you have a stack. If each tool starts from zero, you have a pile. That distinction is the whole game, and it is invisible in a procurement spreadsheet.

The five layers

Every working agency stack I have seen resolves into the same five layers. They are not products; they are jobs to be done, and you can fill each one with a different mix of build, buy, and process depending on size.

LayerJobBuy or build
Client memoryHold briefs, past work, brand rules, resultsBuild — this is the moat
OrchestrationRoute work, hold context, run the loopBuy, then configure
ProductionDraft copy, creative, plans, codeBuy — fastest-moving layer
ReviewCatch errors, enforce brand, check factsBuild the rubric, buy the tooling
MeasurementTie output to outcomes honestlyBuy for data, build the frame

The ordering matters. Memory sits underneath everything, because every other layer is more useful when it can read what the agency already knows. Review sits above production, because production without review is how agencies ship confident mistakes at speed.

Layer one: client memory

This is the layer I would build first and defend hardest. It is where the agency stores the accumulated understanding of each client: the brief history, the brand rules, the winning and losing creative, the measurement readouts, the decisions and why they were made. None of it is exotic. All of it is hard to copy, because it is specific to the work you have done.

A generic model knows the category. Your stack should know your client. That difference is the reason an agency can do in an afternoon what a client's in-house tooling cannot do at all, and it compounds every month you keep it. If you want the long version of why this layer is the one that lasts, I wrote about the client-memory data moat separately.

Layer two: orchestration

Orchestration is the plumbing that moves a job through the stack while carrying context. It answers questions like: when a brief lands, what runs, in what order, with which memory attached, and where does a human look before anything is sent. In small agencies this is often a person and a checklist, and that is a legitimate answer. In larger ones it is a workflow tool or a thin internal app.

The trap is buying an orchestration platform before you know your workflow. The software will happily encode a process you have not designed, which is how agencies end up automating their existing mess. Design the loop on paper first, run it by hand for a month, and only then look for something to carry it.

Layer three: production

Production is the commoditising layer, and the one where buying almost always beats building. The models and tools that draft copy, generate creative, assemble plans, and write code are improving faster than any agency can keep up with internally. Your job here is not to build the engine; it is to choose a small number of engines and learn them deeply.

Resist the urge to add a new generator every week. Depth beats breadth in production, because the quality difference between a tool used casually and a tool used with a tuned prompt and a known failure mode is larger than the difference between the tools themselves. Pick two or three, write down what they are good and bad at, and revise that list quarterly.

The trap to watch is the tool that feels fast and quietly weakens the team's judgement — a failure I wrote about in the fast-but-judgment-weakening tool trap. That kind of tool costs more than its subscription, because the saving in drafting time comes back as errors nobody caught.

Layer four: review and QA

Review is where an agency distinguishes itself, because it is the layer clients are paying for and the layer most stacks neglect. AI makes production cheap and review expensive, which means the bottleneck moves to the review step the moment you scale. A stack with no review layer does not scale; it accumulates review debt until something embarrassing ships.

Build the rubric even if you buy the tooling. What counts as on-brand, what counts as a factual claim that needs a source, what counts as a hard stop before anything goes near a client. The rubric is the agency's judgement made explicit, and it is the thing that lets a junior operator run a senior-quality check without a senior sitting over their shoulder.

Layer five: measurement

Measurement closes the loop, and its job is honesty rather than volume. The stack should be able to say what a piece of work was supposed to change, what actually moved, and how confident you are that the work caused it. That is harder than counting outputs, and it is what separates an agency that can defend its fees from one that can only show activity.

Buy the data plumbing, because attribution data lives in platforms you do not control. Build the frame, because the frame is where judgement lives. A thin, disciplined frame — one business outcome, a few leading signals, an honest note about what you can and cannot attribute — beats a twenty-metric dashboard nobody reads.

How to sequence the build

Start with memory and a hand-run loop. Get the client knowledge into one place, write down the steps a job takes, and run it manually until the steps stop changing. Then add production tools to the steps that hurt most. Add review once production is fast enough to create review pressure. Add measurement when you have something worth measuring.

That sequence feels slow for the first month and saves a year of rework, because each layer is added when the layer below it is stable. The agencies that struggle are usually the ones that bought production first, then tried to bolt memory and review on the side. The stack has to grow upward from what the agency knows.

A worked stack review

Once a year, audit the stack by walking one real job through it end to end. Start where a brief enters, follow it through orchestration and production, name where a human reviews, and finish at the measurement readout. At each handoff, ask two questions: did the context carry across, and could a new joiner follow this without asking? If context is lost between two layers, that is the layer to fix first, because memory leaking between stages is what turns five tools into five unrelated tasks. If no one can trace a result back to the brief that produced it, the measurement layer is decorative. The review takes an afternoon and tells you more than a subscription audit, because it tests the connections rather than the list.

FAQ

Do we need our own models?

Almost never. The value of an agency stack is in memory, orchestration, and judgement, not in training. Fine-tuning is worth it only when you have a large, stable, proprietary dataset and a task the general models handle poorly. Most agencies have neither, and should spend the effort on the memory layer instead.

How much should we build versus buy?

Buy the layers that are improving fast and that other firms also use, because you cannot out-invest the market on those. Build the layers that hold your specific knowledge and encode your specific judgement, because those are the only ones a competitor cannot buy. Memory and review are build layers; production and data are buy layers.

What is the biggest stack mistake?

Adding tools before designing the workflow that connects them. It looks like progress, it produces a renewal list, and it leaves the agency with the same bottlenecks plus more subscriptions. The stack should follow the operating model, never the other way around.

How do we keep client data safe across all this?

Treat the memory layer as the boundary. Client data lives in your controlled store, tools read only what a task needs, and you keep a record of what went in and what came out. That is a governance conversation as much as a technical one, and it should be settled before the stack scales.

The short version

An agency AI stack is five layers — client memory, orchestration, production, review, and measurement — assembled around the operating model. Buy the commoditising layers, build the one that holds what you know, and grow the stack upward from memory. The Execution guide covers the workflow this stack serves, and the for agencies track works it through the agency operating model.

If your agency has built its stack the other way around — tools first — I would like to hear what forced the redesign.

Cheers, Chandler