Diego PérezAnalytics and AI leadership
EN ES
Method · 02

How to prove an AI project worked

2026 Note 02 of 03

Saying an AI project worked is easy. Sales went up after we launched, so the project worked. That sentence has ended more careers than bad models have, because the first person to check it finds out that sales went up everywhere.

The claim problem

Almost every reported AI result I have been handed to review compares the same places before and after. That comparison cannot separate the project from the season, the price change, the new competitor, the weather or the fact that the team paid unusual attention to those locations for three months.

The fix is not statistical sophistication. It is having something to compare against: a set of comparable places that deliberately did not get the change, watched over the same period, with the same reporting.

The cost of a control group is real. You knowingly leave part of the operation without an improvement you believe in. That is the price of being able to answer the question a board will eventually ask, and it is cheaper than not being able to answer it.

Choosing the comparison

The comparison group has to be similar in the things that drive the outcome, not similar in the things that are easy to match on. In retail that means demand shape, size, format and catchment, not the country a store happens to sit in.

The test I use is simple and it happens before the intervention starts: plot both groups for the previous year. If the two lines do not already move together, they are not comparable, and no amount of later adjustment will fix it. Choose again.

Two rules that save arguments later. Fix the groups in writing before the change goes live, with the list of locations recorded and dated. And never let the operation choose which sites get the treatment, because they will choose the ones they expect to do well.

Reading the gap

Once both groups move together historically, the method is difference in differences, and it is far less intimidating than the name suggests. Take the change in the treated group from before to after. Take the change in the untouched group over the same window. Subtract the second from the first. What is left is the part you can attribute to the intervention.

Everything that hit both groups equally, the season, the economy, the pricing decision taken centrally, cancels out. That cancellation is the whole point, and it is why the untouched line is worth more than the treated one.

The chart on the home page is this and nothing more: two lines from a common origin, a dashed line on the day the change started, and the vertical distance between them afterwards. If a result cannot be drawn that way, I do not present it as a result.

The four ways it breaks

Spillover

The treated sites take business from the untouched ones. Now the gap overstates the gain, because the control group got worse rather than the treatment getting better. This is common when the sites are close together, and it is the reason geography belongs in the group design.

Selection

Someone chose the treated group for a reason correlated with the outcome. The worst version is well-intentioned: the team picks the stores with the most upside, which guarantees the treated group would have improved anyway.

Drift in what is being measured

The definition of the metric changes mid-experiment. A category is reclassified, a store is remodelled, returns start being counted differently. This is why the measure has to be locked and documented at the start, and why I reconcile against the warehouse rather than against a working file.

Reporting the number you like

If you compute the gap on twelve metrics and present the one that came out best, you have not measured anything. Declare the primary metric before you look. The others are diagnostics, and they should be labelled as diagnostics when you present them.

What a board actually asks

In my experience it is three questions, in this order. How do you know it was the project? What did the comparison group do? Does this number match the one in the monthly report?

The third question is the one that catches people out. A result that does not reconcile with the figures the company already reports will be treated as an advocacy number, no matter how good the method is. So the reconciliation happens before the result leaves the team, not after someone challenges it.

Getting to the point where you have something worth measuring is note 01. The foundation that makes any of it possible is note 03.