The agent that worked in the demo and stalled in the operation
We arrived at a retailer that already had a service agent running. The digital director was proud of the demo and frustrated with the operation. Both were true at the same time.
In the pitch, the agent answered everything, recommended well, closed the order. In the store, it promised sold-out product, gave wrong delivery dates, and the customer complained to a human afterward. Same agent, two behaviors.
What we found
The demo was honest and dishonest at once. Honest because the agent really did converse well. Dishonest because it ran on a frozen catalog, a snapshot of stock from weeks earlier.
In the operation, stock changes every minute, and the agent did not see that change. It answered with the demo’s confidence about a world that had moved. The conversation was right; the world behind it was out of date.
What we changed
First, wire the agent to real stock, in near real time, instead of the frozen snapshot. That alone cut most of the broken promises.
Then, give it a path for “I do not know, I will hand you to a human”. An agent that admits its limit is worth more than one that guesses with confidence. And a written limit on what it can promise, so it never offers what the system does not confirm.
The number that moved
In about nine weeks, the resolution rate without a human rose, and false availability promises fell to near zero. The customer stopped complaining about the agent to the human.
The honest part: part of the gain came from cutting scope. We pulled the agent off questions it had no way to answer well and focused on reorder and availability. Less ambition, more trust.
What we got wrong
At first, we trusted the agent’s own confidence score to decide when to hand off to a human. The score was miscalibrated: the agent was most confident exactly on the cases it got wrong. We used the wrong signal for two weeks.
The fix was to stop trusting self-assessment and use a business rule: certain kinds of question go to a human, period. It is the kind of detail that does not show up in an optimistic timeline.
What we left behind
A handoff rule based on case type, not on the model’s self-confidence. A dashboard the internal team checks every week, with false-promise rate first. And the scope criterion: the agent only takes on what it can confirm.
Think about the agent or pilot that dazzled in your meeting room. Does it run on real stock, or on a snapshot? If nobody can answer, the demo is fooling you with your own confidence.
If this shape sounds familiar, between the demo that dazzles and the operation that stalls, describe in one paragraph what is happening. We step inside the operation for two weeks and leave with a one-page document your team keeps: what we saw, what the demo hid, and the three changes that close the gap. That is how we work.