Fifteen percent of supply chain leaders say that their AI initiatives have met or exceeded expectations. That's from JBF's Q3 2026 survey of 215 supply chain leaders, and the same survey shows most organizations never built a business case around measurable outcomes at all.
Three things separate a business case that can answer an ROI question from one that can't: setting a measurable expectation before starting, documenting a baseline and success criteria, and returning afterward to check results against the original case. Seventeen percent of respondents never clear the first checkpoint, only 11% clear the second, and only 3% clear the third, so most AI spending in this category gets approved without ever building in a way to know, later, whether it worked as intended and delivers expected value.
The pattern predates AI. JBF's Logistics Strategy Gap Survey found the same discipline gap in traditional technology decisions, with business cases treated as a formality instead of a document anyone expected to check against results. AI cost behaves differently and can move constantly, introducing complexity and uncertainty.
A Cost That Moves With You
Plenty of logistics technology already prices on volume rather than seats. A TMS might run on tiers of annual freight spend under management, $50 to $100 million in one band, $101 to $150 million in the next, with the rate for each tier set at signing and volume checked once a year, typically in arrears. Moving to a new tier takes a material shift in the business like an acquisition, not the normal month-to-month swings in shipment counts.
Most AI tools priced by API call, token, or compute-minute skip that structure, since there's no tier and no annual checkpoint. Every call generates its own charge as it happens, so the total tracks usage in real time instead of resetting on an annual schedule.
A brokerage team pilots an AI assistant to answer WISMO, or "where's my truck," inquiries and projects annual cost from steady volume. When a weather event shuts down a major corridor and exception volume triples for two weeks, the assistant does exactly the work it was built for, but at a cost that has also tripled in real time. A volume-tiered TMS contract would absorb that spike until the next annual review, but the AI bill doesn't.
The comparison most teams reach for, cost per token or minute against cost per hour of labor, holds up well when an AI tool replaces a very specific task an employee used to do. Multiply the hours saved by a loaded rate, subtract what the tool costs to run, and the answer tells you ROI. That comparison depends on a tightly scoped baseline with current-state time measured against AI-delivered time at matching volume. Apply that same math to a tool that produces a faster or better decision than a human, and the logic gets squishy. A decision that used to take a day and now takes twenty minutes doesn't have a clean hourly rate, because time saved wasn't the value driver.
Where the Real Value Hides
Task automation earns its ROI case first because it's the easiest one to build. Coca-Cola's North America operating unit cut its response time on WISMO inquiries from a 90-minute SLA to seconds after piloting FourKites' Tracy AI agent, and returned hundreds of associate hours a year in the process, per the company's published case study. That's measurable value in time saved – but where is that reclaimed time spent and were overall costs reduced? How is the value of providing the customer base with accurate WISMO information determined?
Start with the hours, because that is the number everyone reaches for first. Hundreds of associate hours a year multiplied by a loaded hourly rate produces a clean savings figure, and it belongs in the case. It only counts as savings if the labor line moved. If headcount, overtime, and contract labor all held steady a year later, those hours weren't removed from the cost structure, they were redeployed, and the business case owes an answer to where. Team members who spend reclaimed time on carrier escalations, appointment recovery, or new account onboarding are producing real value, but it's a different claim with a different metric behind it. Decide before the pilot which claim you're making: a lower cost to serve, or the same cost redirected to higher-value work that you can point to and measure.
The second question is what the response time itself is worth to the customer buying the service. Coca-Cola's customers went from a 90-minute SLA to an answer in seconds. Ninety minutes is long enough that a customer calls back, escalates, or plans around not knowing; seconds or minutes is short enough that the question stops being an event in their day. That difference lands in metrics the transportation team usually doesn't own: customer satisfaction scores, escalation and callback volume, and over a longer horizon, retention of customer business. None of those convert to dollars as cleanly as an hourly rate, which is exactly why they get left out of the case.
The Renewal Problem You Can't See Yet
The tiered pricing described earlier is also usually locked for the life of a multi-year agreement, escalators included, so the buyer knows the full rate schedule at signing. Moving to a materially different bill takes a major event, and even then it's visible well in advance. That predictability is something logistics buyers have come to expect from enterprise technology contracts, whether the underlying meter is seats, freight spend or shipments, or transaction volume.
AI doesn't have an equivalent track record yet. Multi-year negotiated schedules are far less common, especially for mid-market buyers without a custom agreement, and current per-unit rates are understood to sit below what it costs providers to deliver the service, subsidized by competition for market share that won't run indefinitely. No one, including the providers, has priced this category at a market-clearing rate long enough to know what a normal adjustment looks like at renewal.
The honest claim here concerns uncertainty, since there's limited history to say how a renewal reset would land, which is a harder risk to plan for than a known escalator. A defensible projection treats the renewal rate as an open variable to test, the same way procurement already tests a tier-boundary scenario for volume-priced software, just without a pre-negotiated ceiling to fall back on.
Building a Case You Can Defend
A defensible AI business case rests on the same three points regardless of what's being automated: document a baseline before the initiative starts, state a specific claim about what should change and by how much, and set a fixed point to check that claim against what happened next. The survey numbers above show exactly where most organizations skip these points, and in what order.
For consumption-priced tools, the projection needs one more layer, modeling cost as a function of usage instead of a flat number and treating the renewal rate as a variable to test rather than an assumption to make. A projection that only holds at the pilot's tested volume will misrepresent the initiative's cost curve the moment usage changes.
About the Author
Tara Buchler is Principal, Strategy at JBF Consulting, bringing more than 20 years of experience at the intersection of logistics operations and enterprise supply chain software. She partners with shippers to design and implement pragmatic, high-impact strategies that align business goals with advanced technology solutions.
Tara’s unique perspective blends vendor-side product leadership, hands-on implementation expertise, and operational insight—allowing her to provide objective advisory services rooted in real-world experience. Her background includes senior roles at e2open, BluJay Solutions, and LeanLogistics, where she helped shape TMS, visibility, and parcel execution capabilities for global shippers.
FAQs
Traditional enterprise software cost is negotiated and locked for the contract term, so the return calculation mostly has to solve for one uncertain variable, the benefit. AI priced by token, API call, or compute-minute makes cost variable too, since it moves with usage in real time instead of resetting on a schedule. A business case has to model two moving variables instead of one, against a pricing history that doesn't exist yet.
JBF's Q3 2026 survey of 215 supply chain leaders found that only 11% document a baseline and success criteria before an AI initiative starts, and only 3% go back afterward to check results against the original business case. Only 15% report that AI has met or exceeded expectations to date.
Treat the renewal rate as an open variable in the projection. Volume-priced software typically locks its rate schedule for the life of the agreement, with tier changes triggered by major events rather than routine growth. AI doesn't have that track record yet, since current per-unit pricing is subsidized by providers competing for market share, and no one has priced the category at a market-clearing rate long enough to know what a normal adjustment looks like at renewal.
Not on its own. Hours saved multiplied by a loaded rate is a starting figure, not a result. It only becomes savings if the labor line actually moves, so check headcount, overtime, and contract labor at the checkpoint. If those read flat, the hours were reinvested rather than removed, and the case has to name the higher-value work they went to and how its output gets measured. The other half of the value sits with the customer: a WISMO answer that arrives in about two minutes instead of ninety shows up in satisfaction scores, escalation and callback volume, and over time in retention and freight awarded at the next bid. Baseline both before the tool goes live.
