Most AI business cases fail before anyone evaluates the technology. They fail because the numbers inside them cannot survive a finance review.
Know the odds before you write anything, because your CFO already does. RAND Corporation research from 2024 found more than 80% of enterprise AI projects fail to deliver on their original objectives. That is roughly twice the failure rate of non-AI IT projects. Gartner projected at least 30% of generative AI projects would be abandoned after proof of concept by end of 2025. MIT’s Project NANDA reviewed over 300 public enterprise AI initiatives and found 95% produced no measurable return at P&L level.
DEFINITION:
AI business case
An AI business case is the financial argument for an AI investment, setting out expected costs, quantified benefits, the payback period and the method used to verify results. In retail it is assessed against the same capital as new store openings and refits, so it competes on evidence rather than novelty.
Meanwhile S&P Global Market Intelligence tracked the same trend moving quickly. The share of enterprises abandoning most of their AI initiatives rose from 17% in 2024 to 42% in 2025.
So your case is not competing against enthusiasm. It is competing against a finance function that has watched this pattern and wants to know why yours is different. These six steps, in this order, are how you show it.
95%
The share of enterprise generative AI pilots that produced no measurable return at P&L level, across more than 300 public initiatives reviewed.
MIT Project NANDA, 2025.
Step 1: Start with a problem that already costs money
Lead with the operational problem, rather than the technology. Cases that open with AI capability invite the question of why now. Cases that open with a quantified operational cost invite the question of how to fix it.
So pick something specific and already bleeding. Promotional displays that never make it to the floor. Store managers spending their shift in the back office. Compliance checks that take longer than they should. Name the process, name who does it today, and name what it costs.
That framing also narrows scope, which matters for step 2. A case built around one problem is far easier to fund than a case built around a category.
Step 2: Check the payback window before you model anything
This is where most retail AI cases die, although it is avoidable.
Retail investment committees apply tighter horizons to software than to physical assets, since the two behave differently. A new store is modeled over seven to ten years. Software typically gets three to five, because the technology dates faster and contracts renew. Within that window, therefore, payback ceilings are firm. Finance expects point applications to pay back in six to twelve months. Enterprise platforms get twelve to twenty-four.
Store operations AI often needs longer than a single-purpose tool, because deployment, integration and adoption all take time. So scope your first phase to pay back inside the window. Proposing the full vision and asking finance to wait rarely works.
There is also a P&L point worth understanding here. Under FASB standards ASC 350-40 and ASU 2018-15, subscription and hosting fees cannot be capitalized. They are expensed as incurred, so they hit operating profit directly. They do not sit below the line as depreciation. Consequently the accounts treat a store refit more kindly than a software subscription.
INSIGHT
Physical assets get a seven to ten year horizon. Software gets three to five, and the cost hits operating profit immediately. That asymmetry is why an AI case competing against a store refit needs a much shorter route to payback.
Step 3: Pick value levers you can actually price
In practice, four levers carry most retail store operations cases. Each has a defensible external benchmark, which matters because finance will check.
First, labor time. Value reclaimed hours at fully loaded rates, not base pay. The US Bureau of Labor Statistics puts median hourly pay at $17.03 for retail salespersons. First-line supervisors sit at $23.33. Employer costs data shows benefits and payroll taxes add roughly 29% to 31%. That gives fully loaded rates near $24.50 and $33.50.
Second, on-shelf availability and promotional compliance. Corsten and Gruen set the global out-of-stock benchmark at 8.3%. That costs around 3.9% to 4.0% of store sales. Chintagunta, Chu and Cebollada studied 5,000 stores. They found 29% of planned promotional displays were never placed at all. Displays that were placed ran for only 62% of the intended campaign. Getting a display onto the floor during a campaign week lifted sales of those products by 9.6%, which is why visual merchandising execution carries real money.
Third, employee turnover. Boushey and Glynn reviewed 30 studies for the Center for American Progress. They put replacement cost for roles under $30,000 a year at 16.1% of annual pay. That is around $5,700 per associate at current wages. You will see 50% to 150% quoted widely. Those figures come from managerial and technical roles and do not transfer to hourly retail, so finance will reject them.
Fourth, store manager time. Decarolis and colleagues, in NBER working paper 31192, studied store-level data. Individual managers explain 25% to 35% of the variance in store productivity. Fisher and Raman went further. Redirecting manager attention toward sales floor execution produced profit improvement equal to 4.2% of store sales. That is the lever Store Manager Copilot is built to pull.
Step 4: Separate cashable savings from capacity
This distinction decides more business cases than any single number, and most proposals ignore it.
Hard savings clear payroll, because the schedule actually shrinks. A reduced shift, a removed schedule block, an avoided hire. Finance recognizes these as direct P&L reductions, because the money genuinely stops leaving the business.
Soft savings, by contrast, are minutes. Ten minutes back per associate per shift is real. It does not remove a schedule block, so it is capacity rather than cash. Finance treats it as non-cashable unless you say what the capacity gets used for.
Consequently, soft savings still count if you attach a redeployment plan. Name the margin-generating activity those hours move into, whether that is peak-hour floor coverage, order fulfillment or coaching. Pret A Manger reclaimed 154,000 hours a year across 525 shops, worth around $3.5 million. They reinvested that time into coaching teams and maintaining product standards. That second half is what makes the first half fundable.
154,000 hours
Reclaimed annually across 525 shops, worth approximately $3.5 million in time value, and reinvested into coaching and product standards.
YOOBIC customer story, Pret A Manger, September 2024.
Step 5: Design the attribution before you start
Attribution is the part ops leaders underestimate, and also where finance pushes hardest. It has to be designed upfront, because you cannot reconstruct a baseline afterwards.
Before-and-after comparison is the default, and also the weakest method available for proving a performance outcome. It cannot separate the software from seasonality, promotional calendars, local wage movements or a general market recovery. Worse still, retailers usually pilot in their weakest stores. Extreme performance is statistically transient, so those stores tend to improve anyway. Crediting that to the software is a mistake finance has seen before.
Matched-store testing fixes this, though it takes planning. Run the technology in one group of stores. Compare against a control group over the same period, matched on format, square footage, sales density, foot traffic and catchment demographics. The difference-in-differences approach then measures change in the treatment group against change in the control group. That strips out anything affecting both.
For example, LEMAIRE did a version of this well. They split their boutiques into cohorts based on execution signals. High performers were stores hitting above 90% on two of three leading indicators. Comparing cohorts year on year, high performers grew sales 17.7% against 7.5% elsewhere. The point is not the number. It is that the number came from a comparison rather than a claim.
UNTUCKit applied the same logic to training. They compared certified stores against non-certified ones to isolate the effect on units per transaction and conversion. Both approaches depend on having consistent execution data underneath, which is what store visits and audits generate as a by-product.
Step 6: Budget for what is not on the invoice
Cost underestimation kills approved projects, which is arguably worse than outright rejection.
Three costs sit outside the licence fee, so build them in. Data readiness comes first, and it is the largest. Gartner has warned that projects without AI-ready data face abandonment. Data work routinely absorbs most project resources before anything reaches a store. Second is change management and training, which is where adoption is won or lost. Third is running cost, because inference, monitoring and retraining are recurring expenses that conventional software never carried.
Boston Consulting Group puts numbers on this split. Their 10/20/70 guideline allocates 10% of effort to algorithms and 20% to technology and data. The remaining 70% goes to people and processes, because that is where the change either sticks or does not.
Therefore build these into the model from the start. A realistic full cost with a smaller net benefit survives scrutiny. A large benefit set against a licence fee alone does not.
What a funded case looks like
Finally, the strongest cases pair a revenue lever with a cost lever, then show the method behind both.
For example, Michaels reclaimed 223,000 labor hours a year across 1,350 stores while lifting task completion by 30%. They reinvested that time into customer-facing zones. That produced $1.8 million in incremental revenue in year one. Voluntary frontline turnover fell 24%, worth more than $8 million in annual P&L savings.
Hugo Boss ran a focused pilot on AI-recommended priority actions. Store-level adoption of those recommendations produced a 3.2% sales uplift. A single lever, measured over a defined window, is often a better opening case than a broad one.
Meanwhile PureGym saved 43 hours per club per year on operational walkarounds. Across the network that is more than 26,000 hours, valued at over £333,000.
“We barely got time to do things once, never mind twice or three times.”
Gordon Macpherson, Group Productivity Director, Morrisons
QUOTE: Taking the hours spent on manual processes out of stores pays for the platform alone. Steve Zawlocki, Vice President IT, Mattress Firm.
Follow the six steps in sequence and the case largely writes itself. It will also be harder to reject, because every number in it can be checked. Other retailers have published what they measured, and their customer stories are a reasonable place to sense-check your own numbers.
Frequently asked questions
How to measure ROI with AI?
Measure AI ROI by comparing a treatment group against a matched control group over the same period, rather than comparing the same stores before and after. Before-and-after measurement cannot separate the effect of the software from seasonality, promotional calendars or general market movement, and it is especially unreliable because retailers tend to pilot in their weakest sites, which improve anyway. A defensible approach sets the baseline metrics before deployment, runs the technology in one group of stores, matches a control group on format, sales density, foot traffic and catchment, then measures the difference between how each group changed. Value the resulting gains at fully loaded labor rates and separate cashable savings that clear payroll from capacity gains that need a redeployment plan.