rdlb · insights August 19, 2026 · 2 min read

The demo is not the test.

An agent's competence is measured against the work that actually arrives, not the examples you wrote for it.

Achilles, the RDLB Agentic mascot, focused, holding one small card beside a tray overflowing with unsorted cards, on a soft ivory background.

The pattern is common enough to set your watch by. An agent demos beautifully on Tuesday. It goes live on Monday. By Thursday somebody is quietly fixing its output before anyone downstream sees it, and by the following month the team has decided the technology is not ready.

Nobody lied in that demo. The agent did the work. The problem is what the demo was evidence of.

Here is the mechanism. An agent's reliability is not a property of the model. It is a property of the model against a distribution of inputs. Change the distribution and the same weights produce a different competence. So when you evaluate on examples you wrote yourself, you have measured performance on a sample of your own construction, and then deployed against a sample the world constructs.

Those two samples are different in a predictable direction. You write test cases from the cases you already understand. The requests that are ambiguous, half-specified, internally contradictory, or attached to a file called final_v3_ACTUAL — those do not occur to you, because if they had occurred to you they would already be handled. Your blind spots are structurally absent from any eval set you author. That is not carelessness. It is the definition of a blind spot.

Which gives the fix its shape. Build the eval set from arrivals, not from imagination.

Take twenty real requests from last month. Not cleaned up, not rewritten, not the good ones. The actual inbound, including the three that annoyed everyone. Run them cold, with no human steering mid-run. Then count how many produced output a person had to rescue before it could move.

That count is a rate, and a rate is a thing you can move. It is also the only honest answer to the question every founder is actually asking, which is not "is the model good" but "what fraction of this work can leave my hands." Twelve rescues out of twenty is a forty percent autonomy rate and a clear brief for what to fix. A flawless demo is a number you do not have.

Then keep the rescues. Every case a human saved is the most valuable data your system will generate this quarter, because it marks the exact boundary of what the agent can be trusted with. A failure that gets patched quietly and thrown away is a failure you will pay for twice — once now, once when it returns at scale and nobody remembers seeing it before. The rescue file is how the eval set stops being a snapshot and starts being a system.

I argued two weeks ago that a standard has to be testable by someone who did not write it. Close the loop from the other side: a test needs cases, and the cases have to be real ones, or the standard is only testing your imagination.

So the Monday move is small and slightly uncomfortable. Pull twenty real requests, run them cold, count the rescues. You will not enjoy the number. You will finally have one.

Takeaway: A demo measures the agent against the work you imagined. An eval measures it against the work that arrives. Only one of those numbers survives contact with Monday. ✱

Achilles Angle · evaluation · agentic systems · reliability

A 30-minute strategy blueprint call maps where a system takes over your highest-cost work.

Book the strategy blueprint call