Every agentic AI vendor will show you a demo. The demo is not the diligence.
A CMO evaluating three vendors sees three polished walkthroughs. Same energy, same claims, same watch-it-work moment. What the demo cannot show you is what happens in week six, when the system hits an edge case, a brand exception, a client the founder still personally handles differently. That is the part you are actually buying, and it is the part no sales deck is built to demonstrate.
The question most RFPs skip.
Most evaluation frameworks ask what the system can do. The better question is what happens when it is wrong. Every agentic system makes a bad call eventually. The system worth buying has a designed answer for that moment, not an apology after the fact.
That answer has a name: the approval gate. Work that touches brand risk, client-facing commitments, or spend routes through a human before it ships. Work that doesn't — drafting, research, first-pass structuring — runs without waiting on anyone. A vendor who cannot draw that line for you in the first meeting has not built the line. See /posture for how that gate gets set for a given engagement.
Four things worth actual diligence.
Ownership of the routing. Ask whether the system is locked to one model provider or routes model-agnostically across providers by task. A locked-in system inherits that provider's pricing, outages, and roadmap changes on someone else's timeline. Yours shouldn't have to.
Read access versus write access. Ask what the system can see versus what it can touch. Read-only connectors into your CRM, analytics, and content systems mean the system learns your business without becoming a liability inside it.
Replayability. Ask whether every output traces back to the input, the prompt, and the approval that cleared it. Audit-grade logs turn "the AI did something strange" from a mystery into a two-minute lookup, which matters the first time legal or finance asks.
Exit cost. Ask what happens the day you leave. A system built with no lock-in hands your data and workflows back in a form you can use elsewhere. A system built to trap you will make that question expensive to ask twice, and most buyers only find out which kind they bought after the contract is signed.
What the contract should say, not just the pitch.
None of this shows up in a demo. It shows up in the architecture, and in what a vendor is willing to put in writing. RDLB runs 13 specialized agents against a 12-operator roster, publishing read-only into client systems, routing model choice by task rather than by contract, and logging every run at audit grade. In 63 days that stack produced more than 44,000 runs for under $50 in total model spend — a ratio that only holds when the system is built for throughput, not for lock-in. See /system for the architecture and /agents for the roster doing the work.
The brands that win the next five years will not be the ones who adopted an agent first. They will be the ones who asked the harder procurement questions before they signed anything.
If you are evaluating an agentic partner right now, book the strategy blueprint call and bring your actual RFP. We'll tell you where it's testing the wrong thing.