Six months in, the work is faster. Everyone on the team feels it. Then the budget review arrives, someone asks what the system returned, and the honest answer is a shrug.
That is not a technology failure. It is a measurement failure, and it happens for a specific reason. The gain lands where nobody was counting. Fewer hours lost to a third revision. A brief that did not have to be rewritten from scratch. A launch that shipped in nine days instead of twenty-two. None of those had a number attached to them the quarter before, so there is nothing to compare them against now.
The window to fix this closes on install day. After that, the old baseline is gone and the argument becomes anecdote.
Take four numbers before anything is connected.
You do not need a measurement program. You need four figures, each answerable in an afternoon from calendars, project files, and email.
Time to first usable draft. Not first draft. First one a senior person would put in front of a client. Revisions to approval, counted per piece of work. Rework rate, meaning the share of work restarted rather than edited. And calendar days from request to publish, measured door to door, including the waiting.
Those four describe the expensive part of brand work: the loop between someone asking and someone approving. Every credible gain from an agentic system shows up in one of them. Take them cold, before install. A number captured after the fact is a guess wearing a decimal point.
Cost per approved output is the only ratio that survives a budget review.
The denominator matters more than the numerator. Count outputs that cleared your approval gate, not outputs produced. Volume is not a result. Approved volume is.
The numerator is total cost: model spend, operator time, and review time. Most buyers assume the model line dominates. It does not. Our own system runs 13 agents that logged more than 44,000 runs in 63 days on under $50 of model spend. The compute is a rounding error. The real cost is the review capacity in front of it, which is exactly why the system is designed around a human approval gate rather than around raw generation.
Divide honestly and the number is defensible in a room full of skeptics. Divide by everything produced and you are quoting a vanity metric back to your own CFO.
Instrument the gate, not the agent.
Agent-level telemetry tells you a run happened. It does not tell you whether the business moved. The approval gate does, because everything of value passes through it once and gets a verdict.
Log every run so it can be replayed, then read the gate: what got approved on the first pass, what came back, what was killed. First-pass approval rate is the single best proxy for whether a system has actually learned your standards, and it is the number that should climb month over month while throughput rises. Ours rises because the brief, the voice rules, and the prior decisions stay in the system rather than in someone's head. That is the same discipline behind our posture on read-only connectors and audit-grade logs.
Set the baseline, then run the ratio quarterly. A 3 to 5 times throughput gain in 90 days is a claim. Four before-numbers and one after-number make it a finding.
If you want the baseline taken properly before anything is installed, book the strategy blueprint call and we will walk your current numbers first.