AI experiments rarely become waste because someone deliberately approved a bad investment. More often, a pilot launches with shared enthusiasm but no single owner for its total cost or commercial outcome. The tool remains available, invoices renew, infrastructure usage grows, and nobody has both the authority and the evidence required to stop it.
That is a material governance problem. Retail Dive reports a broader finding that one-quarter of AI spending is wasted, while a separate report describes Kohl’s testing an AI shopping assistant for product discovery. Together, these examples frame the decision facing retail and e-commerce operators: experimentation can be useful, but deployment is not proof of value. Every initiative needs an accountable owner, explicit economics, measurable outcomes, and an agreed route to expansion or shutdown.
Make one person accountable for the full cost
An AI initiative can involve merchandising, e-commerce, customer service, technology, data, finance, legal, and procurement. Shared participation is sensible; shared accountability is not. When ownership is distributed across a committee, each team sees only part of the economics. Marketing may see engagement, technology may see uptime, and finance may see a vendor invoice. Nobody sees the complete investment against the complete return.
Assign one business owner who controls the use case and is accountable for its outcome. For a product-discovery assistant, that may be the head of e-commerce or digital product rather than the technology team. Pair that person with a finance partner and a technical owner, but keep one named executive responsible for the decision to continue, modify, scale, or retire the system.
Ownership should include the vendor contract, implementation spending, model and cloud consumption, internal labor, monitoring, security, support, and any incremental content or data work. If an expense is necessary to operate the capability, it belongs in the initiative’s cost view—even when it sits in another department’s budget.
Establish a baseline before measuring AI
Teams cannot demonstrate improvement without documenting what happened before the intervention. Before launch, define the target journey, eligible users, comparison group, measurement period, and baseline performance. For a shopping assistant, the baseline could include conversion rate, search exits, average order value, customer-service contacts related to product selection, page latency, and the cost of conventional search and support.
Separate one-time costs from recurring costs. One-time costs may include integration, testing, data preparation, design, and staff training. Recurring costs can include software subscriptions, usage charges, model inference, cloud infrastructure, observability, maintenance, evaluation, and human review. Contract minimums and committed cloud capacity should be counted even when adoption is lower than expected.
This prevents a common comparison error: presenting only a model’s per-query charge while excluding the operating system around it. Finance should maintain a monthly total-cost view, with actual spending compared against the approved budget and original business case.
Measure unit economics and business outcomes
A useful AI dashboard connects operational consumption to an economic unit. For a retail assistant, start with cost per assisted session: total recurring operating cost divided by sessions in which shoppers meaningfully use the assistant. Also track cost per incremental conversion, cost per support contact deflected, and incremental gross profit after AI costs.
The denominator matters. Counting everyone who sees an assistant icon can make adoption appear stronger and unit cost lower. Define an assisted session using a meaningful action, such as submitting a request and receiving a response. Distinguish attempted sessions from completed sessions, and report failures, abandonment, and repeat requests.
Business outcomes should be measured against a credible control group wherever possible. Compare assisted and unassisted journeys among similar shoppers rather than assuming correlation is causation. Core measures can include conversion lift, reduction in search exits, change in average order value, support deflection, return or cancellation behavior, and gross margin. Operational guardrails should include latency, availability, response failure, escalation, and cost per session.
What most people miss
Usage is neither value nor success. High interaction can indicate genuine usefulness, but it can also reflect novelty, confusion, or a poor existing search experience. Similarly, support deflection is valuable only if customers resolve their needs without creating more returns, complaints, or downstream contacts.
Operators should also monitor mix effects. An assistant may increase conversion while steering customers toward lower-margin products, expensive fulfillment choices, or items with high return rates. Measure incremental contribution, not revenue alone. Review customer-experience indicators alongside economics so that savings are not achieved by making the journey worse.
Use stage gates instead of open-ended pilots
Each initiative should move through funded stages: discovery, limited pilot, validated deployment, and scaled operation. Every stage needs a spending cap, a fixed review date, evidence requirements, and a named decision-maker. A pilot without an expiry date is effectively a quiet production commitment.
At approval, set thresholds in three categories. First, adoption: enough eligible customers must use the capability to evaluate it reliably. Second, performance: the experience must meet standards for latency, reliability, and safety. Third, economics: the initiative must achieve an agreed improvement in contribution, productivity, or customer experience within a defined period.
Use a monthly operating review for costs, usage, incidents, and vendor consumption, plus a quarterly investment review for strategic value and continued funding. Material budget overruns, deteriorating unit costs, or missed milestones should trigger an off-cycle review. Procurement should align contract terms with these gates by seeking pilot periods, usage visibility, export rights, renewal notice, and an exit path rather than accepting an early long-term lock-in.
Apply the framework to a Kohl’s-style assistant
For a product-discovery assistant such as the one Kohl’s is testing, begin with a narrow decision statement: does conversational assistance help eligible shoppers find suitable products and produce enough incremental contribution or service savings to cover its full cost?
Instrument the journey from prompt to purchase. Record assistant engagement, successful responses, product clicks, search exits, cart additions, conversion, order value, margin, returns, support contacts, latency, and cost per assisted session. Compare results with standard product discovery for similar traffic, devices, customer types, and categories. Segment results because value may be concentrated in complex categories while simple replenishment purchases see little benefit.
The owner should receive a compact scorecard showing budget versus actual cost, assisted sessions, unit cost, outcome lift, operational guardrails, and forecast payback. Expansion should require evidence that improvements persist beyond initial novelty and remain positive after infrastructure, vendor, labor, and downstream costs are included.
Practical AI cost-governance checklist
- Name one business executive accountable for cost, outcomes, renewal, and shutdown.
- Document the baseline journey and metrics before exposing customers to the tool.
- Build a total-cost model covering vendor, cloud, model, integration, labor, monitoring, security, support, and contract commitments.
- Define an assisted session and calculate cost per assisted session and per incremental outcome.
- Measure conversion, search exits, average order value, margin, support deflection, returns, latency, and failure rates.
- Use a control or credible comparison group rather than attributing every observed change to AI.
- Set a budget ceiling, ROI threshold, operational guardrails, and an expiry date before launch.
- Review cost and usage monthly and the investment case at least quarterly.
- Negotiate consumption visibility, renewal notice, data portability, and termination terms.
- Pause or retire the tool when it repeatedly misses thresholds, lacks sufficient adoption, causes customer harm, or cannot show a credible path to positive economics.
