My WooCommerce AI Agent Told Customers We Were Out of Stock. We Had 386 of Those Products on the Shelf.

Chat scene: a customer asks about a 2.5kg spool, the AI agent replies it is not carried, and a stamp reads 386 in stock at that exact moment.

The hard part of putting an AI agent on a WooCommerce store isn’t the AI. It’s what the AI is allowed to see. Here are the six questions I now ask every vendor — and the demo test that exposes them in about ninety seconds.

I run a real online store. Not a demo, not a side project — a WooCommerce shop that has processed tens of thousands of orders. So when I put an AI agent on it, I thought I understood the risk I was taking. I assumed the failure would be embarrassing: an odd tone, an invented return policy, a customer screenshotting something strange and posting it.

The actual failure was quieter, and it cost a lot more.

For a stretch of weeks, my agent was telling customers we didn’t carry products that were sitting in the warehouse. Not “let me check for you” — a clean, confident, perfectly polite no, we don’t have that. When I finally stopped guessing and audited what the agent could actually see, the number came back: 386 in-stock products were invisible to it.

The agent wasn’t hallucinating. It was answering honestly about a catalog that was missing a large piece of itself. Which is worse, because it means the model was working exactly as designed and the answers were still wrong.

Nobody warns you about this. Every vendor demo shows a bot answering a question beautifully. No vendor demo shows you the bot answering beautifully about a product it doesn’t know exists.

The real failure mode isn’t hallucination. It’s coverage.

The conversation around AI in ecommerce is stuck on the wrong risk. Everyone worries the model will make something up. In practice, on a live store, the model rarely invents a product out of thin air. What it does is answer confidently from an incomplete picture — and a confident answer built on a partial catalog is indistinguishable, to the customer, from a confident answer built on a complete one.

There is no error message. There is no log line that says this customer just got told no about something you had in stock. The conversation looks perfect. The sale just doesn’t happen.

A bot that hallucinates gets caught in a week. A bot with 60% catalog coverage can quietly cost you sales for a year.

Two AI agents answering the same product questionDiagram. A customer asks whether a larger size is available. An agent with full catalog access confirms four are in stock and makes the sale. An agent with a partial catalog denies carrying it, losing the sale silently with no error logged.The same question. Two agents. No error message.“Do you have this in a larger size?”Agent that can see the productYes — we have 4 left in Large.SaleAgent with a partial catalogSorry, we don’tcarry that size.Silent lossBoth conversations look flawless in your logs. Only one of them earned money.

This distinction matters because it changes what you should be evaluating. When you shop for a WooCommerce AI agent, you are not really buying a model. GPT-class models are broadly available and broadly competent; no vendor in this space has a meaningfully smarter one. (I went through this exercise across the main WooCommerce chat tools and the models were the least differentiating part.) What you are buying is a pipeline that keeps a language model tethered to the current state of your store — every product, every price rule, every stock level, updated on a schedule that matches how fast your shop actually changes.

That pipeline is unglamorous, it doesn’t demo well, and it is the entire product.

Why the stakes are higher than they look

It’s tempting to file this under “nice to have.” The research doesn’t support that reading.

The documented average shopping-cart abandonment rate is 70.22% — a meta-analysis of 50 separate studies spanning 2006 to 2025, published by the Baymard Institute.

Most of that abandonment is not fixable by anything you install. 42% of shoppers in Baymard’s own survey abandoned simply because they were browsing and not ready to buy. Extra costs at checkout account for another 40%. No chat widget solves either of those.

Reasons shoppers abandon carts at checkout, Baymard InstituteExtra costs too high (shipping, tax, fees) 40%; Delivery was too slow 20%; Didn’t trust the site with card details 19%; Site wanted me to create an account 18%; Checkout too long or complicated 17%; Website had errors or crashed 17%; Returns policy wasn’t satisfactory 13%; Couldn’t see total cost up-front 12%Why shoppers abandon at checkoutBaymard Institute, US shoppers · respondents could pick more than one reasonExtra costs too high (shipping, tax, fees)40%Delivery was too slow20%Didn’t trust the site with card details19%Site wanted me to create an account18%Checkout too long or complicated17%Website had errors or crashed17%Returns policy wasn’t satisfactory13%Couldn’t see total cost up-front12%Not one of these is fixed by adding a chat widget. The gap a good agent closes sitsearlier — before checkout, when a question goes unanswered.

But a meaningful slice of lost sales happens earlier, before checkout, when someone can’t get an answer. In a YouGov-fielded survey of 6,480 adults across the US, UK, France and Germany, 50% had abandoned a purchase in the previous six months because they couldn’t find enough information about the product, and 83% said they’d leave a site lacking thorough product content. (Syndigo, The State of Product Content 2024 — worth noting Syndigo sells product-content software, though the fieldwork and methodology are fully disclosed, which is more than most numbers in this category can say.)

And here is the finding that should decide your evaluation criteria. Two peer-reviewed studies, not vendor blog posts:

  • Live chat increases purchase probability by 15.99% — measured after correcting for the obvious bias that shoppers who already intend to buy are the ones most likely to start a chat. (Tan, Wang & Tan, Information Systems Research 30(4), 2019.)
  • Chat’s positive effect on conversion is larger when the on-page product information is thinner. (Sun, Chen & Fan, Production and Operations Management 30(5), 2021.)

In other words: the value of conversational assistance comes precisely from its ability to fill gaps in what the page already tells you. Which is exactly the capability an agent loses when it can’t see a third of your catalog. The thing that makes it worth installing is the thing that breaks first, and breaks silently.

The WooCommerce AI agent demo test: six questions, ninety seconds

Every vendor selling a WooCommerce AI agent will offer you a demo. Most demos are rigged in the honest sense — they’re built on a clean sample catalog, with the ten products the vendor knows work. Here’s what I do now instead. I ask for a demo pointed at my store, and then I ask six questions in this order.

1. Ask about something you know is out of stock

Pick a product that went out of stock in the last few days. Ask the agent if you can buy it.

Good: it knows it’s unavailable, and — if you support it — offers a restock notification or a genuine alternative.
Walk away: it cheerfully tells the customer to add it to the cart. This means the agent is reading a catalog export from some point in the past, not your live stock.

2. Ask about your least popular product, by name

Not a bestseller. Go into your admin, sort by units sold ascending, and pick something from the bottom of the list.

Good: it knows the product, its price, and whether it’s in stock.
Walk away: “I couldn’t find that product.” Your bestsellers were never the problem. The long tail is where an agent earns its money, because those are exactly the products a customer can’t find on their own.

3. Ask the price of something currently on sale

Anything with an active discount, a coupon rule, or tiered pricing.

Good: it quotes the price the customer will actually pay at checkout.
Walk away: it quotes the list price. Now you have an agent quoting one number and a checkout charging another — the fastest way I know to turn a sale into a support ticket.

4. Ask something your store genuinely cannot answer

Invent a plausible question with no real answer. “Is this compatible with the 2019 model?” for a product where you’ve never published compatibility data.

Good: it says it doesn’t know and offers to connect a human. Refusal is a feature, and it’s the single hardest behavior to build.
Walk away: a confident, plausible, invented answer. That’s the one that ends up in a screenshot.

5. Ask about a product you discontinued

Something you removed from the catalog weeks or months ago.

Good: it doesn’t offer to sell it. Ideally it points at whatever replaced it.
Walk away: it sells it enthusiastically. Sync that only adds and never removes is how you end up taking orders for things that no longer exist.

6. Ask the way your customers actually type

Not the clean product name. Use the nickname your customers use, in the language they use it in, with a typo, on a phone.

Good: it gets there anyway.
Walk away: it only works when you type the SKU. Nobody types the SKU.

Six-question scorecard for evaluating a WooCommerce AI agent demo1. Something you know is out of stock; 2. Your least popular product, by name; 3. The price of something on sale; 4. A question the store cannot answer; 5. A product you discontinued; 6. The way your customers actually typeThe 90-second vendor demo scorecardScreenshot this. Ask all six against your own store, in this order.01Something you know is out of stock02Your least popular product, by name03The price of something on sale04A question the store cannot answer05A product you discontinued06The way your customers actually typeAny single failure means the agent is reading a stale copy of your store.

One more question that isn’t a test, but is the one I’d ask before signing anything: how often does the catalog re-sync, and what happens when the sync fails? If the honest answer is “nightly, and nothing,” you now know exactly how long a stock error can run before anyone notices. In my case it ran a lot longer than nightly.

When it’s grounded properly, the upside is real

I want to be fair here, because none of the above is an argument against putting an AI agent on a store. It’s an argument about what to check first.

The trajectory in the market data is genuinely striking. Adobe Analytics tracks AI-referred traffic to US retail sites across more than a trillion visits. In July 2024, visitors arriving from AI sources were 43% less likely to convert than other traffic. By February 2025 that gap had narrowed to 9%. By May 2026, Adobe reported that AI-referred traffic was converting 54% better than non-AI traffic, had grown 138% year over year, and those visitors were spending 53% more time on site and viewing 23% more pages.

AI-referred retail traffic conversion, July 2024 to May 2026Bar chart. July 2024: 43 percent less likely to convert than non-AI traffic. February 2025: 9 percent less likely. May 2026: 54 percent better. Source: Adobe Analytics.How AI-referred retail traffic went from worst to bestConversion rate vs. non-AI traffic · Adobe Analytics, 1T+ visits to US retail sites0%-43%Jul 2024-9%Feb 2025+54%May 2026Jul 2024: 43% less likely to convert → May 2026: 54% better

Something changed, and it wasn’t shopper enthusiasm. What changed is that the systems got connected to real data. That’s the whole arc of this technology in ecommerce, compressed into two years.

The counterintuitive part: know when it should say nothing

There’s one more finding worth sitting with, and it complicates the picture in a useful way. In a randomized field experiment with over 6,200 customers, researchers found that undisclosed chatbots performed as well as proficient human agents — and four times better than inexperienced ones. But when the bot was disclosed as a bot before the conversation, purchase rates fell by more than 79%.

I want to be careful with this one, because it gets misused constantly. That study was outbound sales calls for a Chinese fintech company, not an on-site chat widget, and it is routinely illustrated with website-chatbot imagery it doesn’t actually support (Luo, Tong, Fang & Qu, Marketing Science 38(6), 2019). I am not suggesting you hide that your agent is an agent — I disclose mine, and I’d argue you’re obligated to.

But the underlying signal is worth taking seriously: people evaluate the same answer differently depending on what they think is producing it. Which means an agent’s credibility is fragile, and every wrong answer spends more of it than a right answer earns back.

The practical version of that lesson, for me, was learning that the most valuable thing my agent does is sometimes not answer. When a customer asks something outside what the store actually knows, the correct behavior is to stop talking and bring in a person. That was harder to build than any of the answering.

What actually happened to those 386 products

People ask how it got fixed, so: it wasn’t a broken integration. Nothing had crashed. The products had been excluded by a rule that was doing precisely what it had been told to do — the rule was just wrong about which products it applied to. Everything downstream of it behaved perfectly.

The fix took an afternoon. Finding it took weeks. That gap between the two is the entire point of this article: there was no alert, no failed job, no red number on a dashboard. The only symptom was sales that didn’t happen, and you cannot see those.

What I changed afterwards wasn’t the model or the prompt. I started checking coverage the way you’d check stock — as a number I look at, not a thing I assume. If you want to run the six questions above against a live catalog rather than a demo one, that’s what our setup does, and honestly you should run them against whatever else you’re considering too. The tests are the useful part, not the vendor.

What I’d tell myself two years ago

Gartner predicts that by 2029, agentic AI will autonomously resolve 80% of common customer service issues without human intervention. That’s a prediction, not a measurement, and predictions in this category have a poor track record. But the direction is not in question.

What is in question — and what nobody selling you software will lead with — is whether the specific agent you install can see your specific store, completely and currently. That’s not a model question. It’s a plumbing question. And it is the only question in this category where the answer is either yes or no, with nothing in between.

Run the six tests. Take the ninety seconds. Ask about the product at the bottom of your sales report, the one you’d forgotten you stock.

If the agent has never heard of it, neither will your customers.


Sources: Baymard Institute, Cart Abandonment Rate Statistics (updated Sep 2025) · Syndigo / YouGov, The State of Product Content 2024 (n=6,480, Apr 2024) · Tan, Wang & Tan, Information Systems Research 30(4), 2019 · Sun, Chen & Fan, Production and Operations Management 30(5), 2021 · Luo, Tong, Fang & Qu, Marketing Science 38(6), 2019 · Adobe Analytics retail traffic reports, Mar 2025 and Jun 2026 · Gartner press release, 5 Mar 2025.