The demo went well, the applause was real, and nothing changed. The incentives that keep the show running.
The demo went well. It nearly always does. The room was a good one, the executive team in the front row, the vendor's best presenter on the stage, and the AI did the thing it was built to do with a fluency that drew an actual round of applause. I have sat in that room many times, on both sides of the stage, and the applause is real. So is the feeling on the way out that the business has just moved.
Six months later, nothing has changed on the floor. Not one process, not one roster, not one decision that a store manager makes differently on a Tuesday. The pilot report says the technology performed. The dashboard exists. And the people who applauded have moved on to the next demo, because there is always a next demo.
I have started calling this AI innovation theatre, and I want to be careful with the phrase, because it sounds like an accusation and I do not mean it as one. Nobody in that room was pretending. The theatre is not a lie anyone tells. It is a set of incentives that make the performance of innovation more rewarding, for everyone involved, than the slower and less visible work of making innovation stick. Understanding the incentives is the only way I know to stop doing it.
The demo rewards the wrong moment. A demo proves that the technology can do something. That is the easiest question in the whole journey and, since the arrival of generative AI, almost always answered yes. The hard questions come later: whether anybody will use it on a busy Saturday, whether the data underneath it is good enough in the third-worst store rather than the best one, whether the process it replaces was actually the problem. Those questions have no stage. Nobody applauds a process map.
The pilot is designed to succeed. It runs in the stores most likely to say yes, with the manager who volunteered, supported by the project team and the vendor's people on the ground. Of course it works. It has been given every condition the rollout will not have. A pilot that is designed to prove the technology will prove the technology. It will not tell you whether the organisation can carry it, because that was never the question it was asked.
Announcing is safer than embedding. A retailer that announces an AI initiative gets a headline, an analyst note, and a slide for the results presentation. A retailer that simply embeds one gets a slightly better number eighteen months later that nobody can attribute. If you are a leader with a three-year horizon, the incentive is not subtle. And vendors, who are not villains either, are paid on the contract rather than the adoption.
The frontline is measured, not asked. The people who would tell you in the first week that the tool does not fit the day are the last people consulted and the first people scored on usage. So they use it when someone is watching and work around it when no one is. The usage number looks fine for a quarter. Then it does not.
Nobody set a parachute point. The moment to decide whether to continue, change course or stop was never fixed in advance. So the initiative drifts, technically live, functionally invisible, too sponsored to kill and too unloved to grow. I have written elsewhere about innovation accounting, and this is the failure it exists to prevent.
The obvious cost is the money, and MIT's finding that ninety-five per cent of enterprise generative AI pilots return nothing measurable puts a number on it. The less obvious cost is trust. Every cycle of theatre teaches the organisation something. It teaches the frontline that the new thing will pass. It teaches middle managers that the safe response to an announcement is polite compliance and quiet delay. It teaches the executive team, eventually, that AI does not work here, which is the wrong lesson drawn from the right evidence. By the third or fourth cycle the readiness gap has widened, not because the technology got harder but because the organisation has learned not to believe.
I do not think the answer is fewer demos. Demos are useful, and they are fun, and retail could use more fun. The answer is to change what happens before and after them, and it comes down to the three quiet questions we ask of every organisation we work with.
Before the demo: what problem is this for? Not a use case, a problem, one that someone in a store would recognise. If the purpose is not clear enough to aim the technology at something that matters, it will be aimed at whatever demos well, and the demo will be the high point.
Around the pilot: are the people ready to trust it, and ready to be trusted with it? Run the pilot in an ordinary store with an ordinary manager, ask before you measure, and treat the first week's workarounds as the most valuable data you will collect.
After the pilot: is there a repeatable way from here to everyday use? Someone owns the transition. The old process is switched off, not left running beside the new one. There is a parachute point, agreed in advance, where the data says continue or stop and stopping is a normal outcome. In ReFRAME we call the last stage Embed, and it is the one that never gets a standing ovation.
The retailers who get real returns from AI will not, I suspect, be the ones with the best demos. They will be the ones who found the theatre a little boring and went back to the store.
Every article ends with the assessment it pairs with. Free, ten minutes or less, results on screen.