The problem

The client runs contract warehousing for around forty e-commerce brands across three facilities. Its leadership team had collected nine AI proposals from department heads and two outside vendors: a warehouse-floor copilot, demand forecasting, a customer-facing tracking chatbot, computer vision for pallet counting, automated carrier invoice auditing, and several smaller ideas. Each proposal came with its own optimistic ROI slide, and none of them used the same assumptions.

The budget covered one serious build this year. The CFO asked a simple question nobody could answer yet: which of these pays for itself first, using our numbers instead of the vendors'?

What we actually did

Week 1: measuring the baseline before modeling anything

Most AI ROI models fail at the baseline, not the projection. We pulled twelve months of time-tracking, ticketing, and accounts-payable data and measured what each targeted process actually cost today: hours per week, loaded labor rate, error rate, and the downstream cost of each error. For carrier invoice auditing, for example, two analysts spent roughly 60% of their week reconciling carrier bills against contracted rates, and the finance team estimated overbilling slipped through on a meaningful share of invoices nobody had time to check.

Week 2: a single model with shared assumptions

We rebuilt all nine proposals in one spreadsheet model with the same inputs: loaded labor cost, realistic automation rates (we used the lower end of what we see in production, not demo accuracy), build cost, and ongoing run cost. Run cost included inference (estimated from real document and message volumes multiplied by per-token pricing, with a 30% buffer for retries and prompt growth), plus hosting, monitoring, and a maintenance allowance of roughly 15% of build cost per year.

Week 3: sensitivity analysis instead of a single number

Every proposal got a low, expected, and high case, and we flagged which single assumption each result depended on most. The warehouse copilot, for instance, only paid back if floor staff adoption exceeded 70%, which the operations team considered unrealistic across three shifts. Demand forecasting looked strong on paper but depended on clean SKU-level history the client didn't yet have for half its brands.

Week 4: a ranked roadmap and a pilot definition

Carrier invoice auditing ranked first by a clear margin: high document volume, structured inputs (invoices and rate contracts), a directly measurable dollar outcome (recovered overbilling), and no dependency on behavior change from floor staff. We wrote a pilot definition with pass/fail criteria agreed up front: extraction accuracy on line items, match rate against contracts, and dollars recovered per month.

A design decision worth calling out
We ranked projects by payback period and by how directly the outcome could be measured, not by strategic appeal. A project whose value shows up as dollars on an invoice is far easier to defend at a budget review than one whose value is "better decisions." That measurability was worth as much as the raw ROI figure.

Challenges and tradeoffs

Results

The client funded the invoice auditing build as the year's one AI project. A follow-on engagement delivered it in ten weeks, and at the seven-month mark the measured recovered overbilling plus analyst time redeployed to exception handling had covered the full build cost. The other eight proposals stayed in the model with updated assumptions, so the next budget cycle starts from measured data rather than new slides.

What we'd do differently

We would run the one-week manual time study at the start by default rather than as a fallback. Every ROI model we build is only as good as its baseline, and a short, direct measurement is cheaper than debating a number pulled from a coarse system.