ShopX Research

ShopX Research · September 2026

Where AI actually touches the money in a Shopify store

AI now sits at every stage of the funnel: discovery, ads, pages, assistants, checkout, retention. Six months of operator posts, screened one by one, every number traced to whoever said it first.

Summary

Ten findings, and the sentence each one needs after it

AI now touches roughly one shopping journey in nine, sends traffic that converts better than search, and has already rearranged who gets recommended. It is also smaller than the headlines suggest, measured in ways that flatter it, and quoted with the qualifier removed.

5,328posts screened, one by one
1,000kept as relevant and substantive
329distinct numbers, de-duplicated
87traced to a first-party or research source

The short version

  • Real, and still small. Similarweb puts AI in 11.4% of shopping journeys; John Lewis sees 2.5% of product searches arrive through agents; one ten-store Shopify sample attributes 0.10% of revenue to AI referrals. All three are true — they measure different things.
  • The traffic is unusually good. Adobe: AI-referred retail visits converted 42% better than non-AI traffic in March 2026 and 60% better in July. Shopify: about 50% better than organic search for sessions that land on a product page, with 14% higher order value.
  • Growth is decelerating as the base fills in. Adobe's year-on-year readings run 693% (holiday 2025) → 393% (Q1 2026) → 125% (spring) → 62% (July 2026). Shopify's own multiples fall from ~13x AI orders in Q1 to ~3x in Q2.
  • One July day changed discovery. On 10 July 2026 the share of ChatGPT Shopping recommendations pulled from merchant product feeds went from 8.26% to 61.54%, and the ten largest merchants' share of recommendations went from 22.5% to 41.8%. Feed quality became a distribution channel overnight.
  • Agents research; people still pay. Ipsos: 27% of AI-aware consumers use AI for product research, 9% let it buy. Visa: 23% trust GenAI to pay on their behalf, rising to 61% when a named payment brand handles it.
  • In-chat checkout underperformed the click-out. Walmart listed about 200,000 items in ChatGPT Instant Checkout and found in-chat purchases converted roughly three times worse than sending the shopper to Walmart.com.
  • The assistant numbers are comparisons, not lifts. Sparky users spend 40% more per order than non-users; Mylow users convert about 3x. Shoppers who choose an assistant are not the same shoppers as those who do not.
  • Two areas have no evidence at all. Across 5,328 screened posts, not one first-party or research-grade number showed measured lift for AI in landing-page/checkout optimization, or for AI imagery and virtual try-on. Vendor claims fill that space.
  • Ads are moving into the answers. Sponsored placements appeared in 6% of ChatGPT hotel answers in March 2026 and 24% by May; OpenAI reports a $1B annualised ads run rate inside 200 days. The organic AI visibility merchants are being told to chase is the surface being sold.
  • One number in wide circulation does not exist. "AI traffic converts 54% better, per Adobe" matches no Adobe release in this corpus. Neither does "Shopify: AI converts 80% better". Both travel widely.

This report was built the slow way. Every LinkedIn and X post in the corpus was read and scored by a model, one post at a time — 5,328 of them — and 1,000 kept as relevant and substantive. Every number in those 1,000 posts was extracted with its source, then 711 raw claims were consolidated into 329 distinct facts: the same figure repeated by forty accounts counts once. Each fact carries the organisation it originated with and a reliability grade, and the grading is the same for everyone, so a vendor's launch claim is rated low even when the vendor is a company we like.

What that process keeps producing is a gap between what happened and what was said about what happened. Shopify disclosed a 13x figure for one quarter from a small base and a 3x figure the next; the 13x is the one still circulating. Walmart's assistant premium is a comparison between two self-selected groups; it is quoted as a lift. Adobe published 42% in March and 60% in July; the number that spread fastest, 54%, was published by nobody.

So the chapters are arranged by where the money moves — discovery, agents, the Shopify stack, on-site experience, funnel, ads, retention, support, content — and each one states what is live, what has shipped but is unproven, and what is still a demo. Chapter 11 lists the numbers this report refuses to use, and gives you a way to produce your own instead.

The map

The funnel, stage by stage: what AI is actually doing at each one

Eleven surfaces, one funnel. This is where each AI application sits, who is running it, how strong the evidence is, and what a merchant can do about it this quarter. Every row is expanded in the chapter it points to.

Read the strength column first. Three stages have first-party or research-grade numbers behind them; three have nothing but vendor claims. That asymmetry is the most useful thing on this page — it tells you where to move budget and where to run your own test before believing anyone.

Funnel stageWhat AI is doing thereWho is running itEvidenceDo this quarter
Demand & discovery
Ch. 02
Shoppers ask an assistant instead of searching. Answers are assembled from merchant product feeds, not web pages, since 10 July 2026. ChatGPT, Google AI Mode, Gemini, Perplexity, Amazon's assistant. 722 of the top 1,000 US retailers have ChatGPT as their largest AI referrer. First-party / research Adobe, Shopify, Similarweb, Profound all measure it Fix the feed before anything else. Completeness and attributes are now distribution.
Ad exposure
Ch. 07, 12
Creative is generated rather than filmed; ads are also appearing inside AI answers. Operators running AI-UGC pipelines; ChatGPT Ads (self-serve in 31 EU markets), Amazon Sponsored Prompts, AI Mode sponsored cards. Named but secondhand ad load measured; creative results self-reported Budget for judging creative, not only making it. Volume is no longer the constraint.
Landing & PDP
Ch. 06, 13
Pages and variants built by model; audits generated in seconds; product copy and imagery produced at scale. Page-builder agents, n8n audit workflows, Shopify Rollouts/SimGym. Vendor or unsourced no first-party lift number exists in this corpus Treat as a cost saving, not a conversion gain, until your own holdout says otherwise.
Consideration on your store
Ch. 05
Assistants, quizzes and semantic search answer "which one is right for me" on the merchant's own surface. Walmart Sparky, Lowe's Mylow, Amazon's on-site assistant, Shopify-connected concierges, Rebuy/Nosto-style engines. Named but secondhand every headline is a user-vs-non-user comparison Deploy it, but measure with a holdout. Self-selection inflates every published figure.
Checkout
Ch. 03
The agent fills the cart; a human approves. In-chat payment exists but is retreating. UCP (Google/Shopify), ACP (OpenAI/Stripe), Meta Muse, Shopify Agentic Storefronts — on by default. First-party / research the one measured number is negative: Walmart's in-chat checkout converted ~3x worse Connect the rails, keep the checkout on your store, and log what the shopper authorised.
Support during purchase
Ch. 09
Agents answer product and WISMO questions, and increasingly propose operational actions for a human to approve. Gorgias Cortex, Fin, in-house agents; per-resolution pricing is now standard. Named but secondhand Gartner: 1 in 4 use cases show positive ROI Compute fully loaded cost per resolved case before expanding.
Retention
Ch. 08
Flows drafted, segments built and campaigns personalised by model, inside the tools merchants already pay for. Klaviyo (agents, MCP, personalization layer with holdout testing), Shopify Sidekick reaching into apps. First-party / research flows = 41% of email revenue from 5.3% of sends Point AI at flows first. That is where the revenue per send already is.
Catalog & content behind all of it
Ch. 10
Descriptions, images, video and attribute enrichment produced at a fraction of the old cost. Everyone, at wildly varying quality. Cost per asset collapsed; cost per approved asset did not. Vendor or unsourced no measured lift; the payoff that is checkable is retrieval Spend the saving on structured completeness, not on more pictures.
Measurement across the whole funnel
Ch. 11
The instruments moved: GA4 split out an ai-assistant medium mid-June 2026; rank stopped predicting AI citation. You, or nobody. First-party / research Ahrefs, GA4, Adobe all document the shift Re-cut your channel reporting before quoting anyone's percentage, including this report's.

Two things fall out of the table. First, AI's grip is tightest at the two ends — discovery and retention — and weakest exactly in the middle, on the pages merchants spend most of their time editing. Second, the stages with the best evidence are the ones where a platform publishes numbers because it is selling you the channel; the stages with no evidence are the ones where the tooling is cheapest to buy. Those two facts are not unrelated.

Chapter 01

AI is in about one shopping journey in nine, and under 1% of traffic at the largest retailers

The measured shares are small, the forecasts are enormous, and most of the distance between them is definitional. Consumer adoption reads anywhere from 19% to 70% depending on what the survey asked, and US agentic-commerce forecasts for 2030 differ by more than three times depending on what each house counts.

What holds up

  • Ambition and readiness are not close. 73% of consumer companies plan to deploy agentic AI within two years; 24% report even moderate adoption today (First-party / research Deloitte, State of AI in the Consumer Industry 2026 · post).
  • Shoppers are ahead of merchants. 65% of shoppers plan to use AI for holiday shopping; 8% of retailers feel very confident they are ready (First-party / research Narvar, 1,348 shoppers and 100 retail executives · post).
  • The 2030 forecasts are not measuring the same market. McKinsey puts US agent-orchestrated retail at $900 billion to $1 trillion; Bain puts US agentic commerce at $300–500 billion (First-party / research McKinsey · post; Named but secondhand Bain · post).
  • Where AI money is going inside companies, it is not going at revenue. 71% of marketers aim AI at productivity, 9% at revenue growth (Named but secondhand Epsilon 2026 · post).
  • Roughly one third of consumers across 12 markets now use AI somewhere in the purchase journey, about triple the level 18 months earlier. "Somewhere in the journey" is a long way from "bought" (First-party / research BCG Center for Customer Insight, 13,000+ consumers · post).

Goldman Sachs tracked the share of traffic that generative AI sends to five of the largest ecommerce platforms in the world — Alibaba, Amazon, JD, Coupang and MercadoLibre. Across 2024 and 2025 the lines sit at or below 1%. From March 2026 they bend sharply upward. They are still at or below 1%. That is the honest opening position for this report: the growth rates are real, the base is small, and a large share of the numbers circulating this year are one of those two facts wearing the other's clothes.

Share of traffic attributed to GenAI at Alibaba, Mercadolibre, Coupang, Amazon and JD, 2024–2026
Figure 1Share of traffic attributed to GenAI at Alibaba, Mercadolibre, Coupang, Amazon and JD, 2024–2026: still near or below 1%, but accelerating sharply since March 2026.Source: Goldman Sachs (per post) · via Arvind Srinivas, X 2026-08-31

Read the chart for its shape, not its level. The steepening from March 2026 is the thing worth planning around; the level says you have time to plan. One caveat travels with it: the post attributes the data to Goldman Sachs, but the image itself carries no source line, so treat it as named-but-secondhand.

What is actually measured

The most defensible market-level read comes from Similarweb and Statista, who put AI in 11.4% of US shopping journeys in 2026, up from 4.5% in 2024, measured on desktop. The same dataset carries the number that most retellings drop: 89% of shopping journeys that involve AI also involve search. Only 11% of them involve AI without search. AI is being added to the journey, not substituted into it, and that distinction decides whether your search budget is under threat or merely under-instrumented.

Similarweb State of Ecommerce 2026
Figure 2Similarweb State of Ecommerce 2026: AI touches 11.4% of US desktop shopping sessions, gen-AI referrals to ecommerce grew 203% YoY, and AI+Search journeys convert at 23%.Source: Similarweb · via Aleyda Solis 🕊️, X 2026-09-10

Similarweb's 23% conversion figure on the same card is a journey-level rate, not a session rate. It cannot be laid next to Adobe's or Shopify's session-level numbers, and Chapter 11 shows what happens when people do it anyway.

Stated preference points the same way. Semrush asked consumers to pick a single source they would keep before buying: 19.4% chose an AI chatbot, behind reviews at 37.1% and search engines at 36.7%. A fifth of shoppers naming an AI assistant as the one source they would keep is a real result. It is also a hypothetical question, and it is third place.

Semrush survey
Figure 3Semrush survey: forced to pick one pre-purchase source, 19.4% of consumers choose an AI chatbot, behind reviews (37.1%) and search engines (36.7%).Source: Semrush · via Semrush, X 2026-09-10

Generational data narrows it further. PYMNTS Intelligence found 41% of US millennials using ChatGPT for product discovery in July 2026 — behind Google at 57%, but already ahead of Amazon at 37%. The purchases themselves still happen largely in stores. Discovery has moved faster than checkout, which is why a chapter on discovery (02) has thirty-nine high-reliability facts and a chapter on checkout (06) has one.

PYMNTS Intelligence (July 2026)
Figure 4PYMNTS Intelligence (July 2026): 41% of US millennials use ChatGPT for product discovery, behind Google (57%) but ahead of Amazon (37%); purchases still happen largely in stores.Source: PYMNTS Intelligence · via Jason Helmer, LinkedIn 2026-09-08

The adoption number depends entirely on the question

Every headline share in the corpus is technically defensible and none of them are comparable. Here is the same phenomenon measured six ways:

SourceWhat was actually askedShare
SemrushForced to keep one pre-purchase source, chose an AI chatbot19.4%
BCG (13,000+ consumers, 12 markets)Used AI at some point in the purchase journey~1/3
Adobe survey, June 2026Used AI for online shopping that month41%
DeloitteHave shopped with AI at least once56%
Accenture Consumer PulseUse AI tools weekly — any use, not only shopping57%
Survey shared by WWU CEBR"Using AI" — definition not stated64%

The top and bottom of that table differ by more than three times, and the gap is almost entirely in the wording. Deloitte's 56% clears a very low bar — once, ever. Accenture's 57% is not about shopping at all. BCG's roughly one third is the most useful of the six because the base and the window are both stated, and because the same BCG report gives the number that matters more: 13% of consumers say they buy whatever the AI tool recommends. Presence in the journey is about a third of shoppers. Handing over the decision is 13%.

Handle with care

Three numbers in this chapter's ledger should not be quoted. An agency post claims agentic commerce will grow from $8 billion to $1.5 trillion in four years at a 45–60% CAGR; that CAGR does not produce that end point — you would need roughly 270% a year. A news post projects $385 billion of US agentic commerce by 2030 without naming a forecaster. A newsletter claims 88% of retailers have adopted AI and 7% have scaled it, also unattributed. All three are quotable-sounding and none can be traced.

Three times the disagreement, one word of explanation

US agentic commerce by 2030: forecasts differ by more than 3x

The houses define the market differently: spend agents complete, versus spend agents merely orchestrate.

$0B$250B$500B$750B$1,000BMcKinsey — high$1,000BMcKinsey — low$900BBain — high$500BUnattributed projection$385BBain — low$300B
Named but secondhandBain & Company; McKinsey; one unattributed news projection (low reliability). Both houses publish ranges; these are the ends of each range.
Show data table
McKinsey — high$1,000B
McKinsey — low$900B
Bain — high$500B
Unattributed projection$385B
Bain — low$300B

McKinsey's US figure and Bain's US figure describe the same country, the same year and the same technology, and differ by a factor of more than three. The explanation is one word: orchestrate. McKinsey's $900 billion to $1 trillion counts retail revenue an agent had a hand in directing. Bain's $300–500 billion counts a narrower agentic market. McKinsey's global equivalent, $3 trillion to $5 trillion of consumer commerce by 2030, is a scenario range, and the word "up to" falls off it in most of the posts that repeat it.

Share forecasts fragment the same way, and the base is always the trick. Gartner's 20% by 2030 is 20% of all transactions. McKinsey and Rye's 15–25% is of ecommerce transactions. EMARKETER's 15.8% is of US ecommerce sales, under a base case, up from 3.2% in 2026. Three shares that look like the same claim, three different denominators, all landing on the same year.

EMARKETER base-case forecast
Figure 5EMARKETER base-case forecast: AI platforms and assistants directly drive 3.2% of US retail ecommerce in 2026, rising to 15.8% by 2030. Dollar values redacted.Source: EMARKETER Forecast, July 2026 · via Sky Canaves, LinkedIn 2026-09-02

EMARKETER's curve is the most operationally useful forecast in the pack, for a reason that has nothing to do with its end point: it starts at 3.2% in 2026, so it tells you the size of the thing you are being asked to budget against today. The report's own analyst adds the detail that reframes the whole category: the majority of those AI-driven sales are expected to come through retailer-native assistants such as Alexa for Shopping and Walmart's Sparky, and fewer than half through platforms like ChatGPT and Gemini. If that holds, the surface merchants control matters more than the surface they are panicking about. Chapter 05 takes that surface apart.

Ambition is cheap; production is not

Consumer companies: agentic AI ambition against operational readiness

Share of surveyed companies. The gap between the first bar and the rest is the story.

0%20%40%60%80%Plan to deploy agentic AI within 2 years73%Report at least moderate adoption today24%40%+ of AI experiments in production23%Mature governance for autonomous agents20%
First-party / researchDeloitte, State of AI in the Consumer Industry 2026.
Show data table
Plan to deploy agentic AI within 2 years73%
Report at least moderate adoption today24%
40%+ of AI experiments in production23%
Mature governance for autonomous agents20%

Deloitte's survey of consumer companies is the cleanest picture of the gap in this chapter's ledger. 73% plan to deploy agentic AI within two years. 24% report even moderate adoption today. 23% have moved at least 40% of their AI experiments into production. 20% have mature governance for autonomous agents. And 82% have not redesigned a single job around AI capability — which is the number that predicts the other four, because an agent that nobody's role changed to accommodate is an agent nobody is accountable for.

65% of shoppers plan to use AI for their holiday shopping this year. 8% of retailers say they feel ready for it. I read that gap twice.

Sarah Willersdorf, Kinship AI — LinkedIn, 2 Sep 2026

Narvar's survey behind that quote is small on the retailer side: 100 executives against 1,348 shoppers. Treat the 8% as directional. The direction is corroborated by where the money is pointed inside marketing teams.

Epsilon 2026 data as summarised by a CRM consultant
Figure 6Epsilon 2026 data as summarised by a CRM consultant: 71% of marketers aim AI at productivity, only 9% at revenue growth.Source: Epsilon 2026 · via James Tyler, LinkedIn 2026-08-18

Epsilon's split is the quiet finding of this chapter. 71% of marketers aim AI at productivity; 9% aim it at revenue growth. Intuit's QuickBooks survey of small businesses, January 2026, lands in the same place from the other side: 78% report productivity gains, 43% report revenue increases. Firms are buying AI to do the current job faster, then reporting the results as if they had bought growth.

Who is actually shipping

Four operators in the ledger are past the pilot stage, and none of them got there through a shopping assistant. Walmart's US ecommerce grew 24% in Q2 FY27 while US comparable sales grew 2.6%. The gap is real and widely quoted, and Walmart does not attribute it to AI, whatever the listicles do. Alibaba's merchant-facing ecommerce AI agents took more than 60,000 paid users in five months, which is the only clean paid-adoption number in the chapter. Shopee's in-house 245-billion-parameter commerce model, Compass, went from 3 billion to 340 billion API tokens a month in eight months. Tesco consolidated 250 separate AI initiatives into one group strategy backed by a 6,000-person technology team — the least glamorous item on the list and probably the most transferable.

Against that, the headline multiples are decelerating exactly as a small base predicts. Shopify reported AI-referred orders up about 13x year over year in Q1 2026 and about 3x in Q2 2026. Both are official; both are correct for their own quarter; neither is a slowdown in absolute volume. Chapters 02 and 04 take those numbers apart properly.

65% vs 8%

Shoppers planning to use AI for holiday shopping, against retailers very confident they are ready. Holiday 2026.

First-party / research Narvar (1,348 shoppers, 100 retail executives) · post

73% → 24%

Consumer companies planning agentic AI within two years, against those reporting even moderate adoption today. 2026.

First-party / research Deloitte, State of AI in the Consumer Industry 2026 · post

~1/3

Consumers using AI somewhere in the purchase journey across 12 markets, roughly triple the level 18 months earlier. 2026.

First-party / research BCG Center for Customer Insight (13,000+ consumers) · post

3.3x

Spread between the highest US agentic-commerce forecast for 2030 ($1T, McKinsey) and the lowest ($300B, Bain). Different definitions, not different confidence.

Named but secondhand Bain via LinkedIn · post

89% of merchants say they are actively preparing for agentic commerce. Agents are involved in about 3% of transactions today. That gap is being filled with a bet.

Nikki Baird, retail analyst — LinkedIn, 8 Sep 2026

Neither of Baird's two numbers can be traced to a named source in this corpus, so do not quote them as measurements. Quote the sentence after them. The preparation is real, the transaction share is small, and the space between is being filled with capital allocated on forecasts that disagree by a factor of three.

Operator note

Do not budget against a 2030 forecast whose definition you cannot restate in a sentence. Budget against one number you can produce yourself: the share of your own sessions and revenue that arrives from AI surfaces this month, with the baseline named. Most stores cannot produce it — AI referrals land in "direct" or "other" and stay there. If you cannot isolate that channel in your own reporting before the holiday quarter, that is the finding, and it is a cheaper problem to fix than any of the ones in the rest of this report.

Chapter 02

Discovery moved into the answer, then into the feed

AI still sends retailers a low single-digit share of their traffic, and its growth rate is falling fast. What changed in 2026 was not the volume but the mechanism: on 10 July, ChatGPT stopped reading the open web for most of its product recommendations and started reading merchant feeds.

What holds up

  • Growth is real and decelerating. Adobe measured AI-referred traffic to US retail sites up 393% year over year in Q1 2026 and 62% in July 2026 — each against its own year-earlier base (Adobe Analytics).
  • AI visitors convert better, but the baselines differ. Shopify: about 50% better than organic search, for sessions landing on a product page, Q1 2026. Adobe: 42% better than all non-AI traffic, March 2026. Neither is a controlled test.
  • One day reshuffled ChatGPT. Feed-sourced product recommendations went 8.26% → 61.54% on 10 July 2026, and the ten largest merchants' share of recommendations 22.5% → 41.8% (Profound, ~1.75M prompts).
  • Ranking stopped being the ticket. AI Overview citations coming from top-10 organic results fell from 76% to 38% in a year, across 863,000 searches (Ahrefs).
  • Almost nobody uses AI alone: 89% of shopping journeys involving AI also involve search (Similarweb).

How much traffic is actually at stake

Every retelling of AI discovery opens with a growth rate, and the growth rates are genuinely large. They are also shrinking each quarter, because the base is no longer near zero. Adobe's successive readings of one metric — AI-referred visits to US retail sites, year on year — run up to 4,700% (July 2025), 693% (holiday 2025), 393% (Q1 2026), 125% (April–June 2026), 62% (July 2026). Quoting the 2025 number in 2026 is the most common mistake in this corpus.

AI-referred traffic to US retail: growth is slowing as the base grows

Year-on-year growth, successive Adobe readings. Each period is measured against its own year-earlier base.

0%1,250%2,500%3,750%5,000%Jul 2025 (up to)4,700%Holiday 2025693%Q1 2026393%Apr–Jun 2026125%Jul 202662%
First-party / researchAdobe Analytics / Adobe Digital Insights, 2025–2026. Periods differ in length; the Jul 2025 figure is an "up to" reading.
Show data table
Jul 2025 (up to)4,700%
Holiday 2025693%
Q1 2026393%
Apr–Jun 2026125%
Jul 202662%

The absolute shares are the sobering half. Bernstein's reading of Similarweb data puts generative-AI referrals at 0.8–1.4% of total traffic at major US retailers in August 2026 — while the same analysis shows them at 22–39% of referral traffic, a denominator that excludes direct, organic and paid. That gap is why Bessemer's widely-shared "15–20% of referral traffic from AI chat" is compatible with a store where AI barely registers in the totals.

Bernstein/Similarweb
Figure 7Bernstein/Similarweb: despite rapid growth, gen-AI referrals are still only 0.8–1.4% of total traffic at major US retailers as of Aug 2026.Source: Similarweb, Bernstein analysis · via Arvind Srinivas, X 2026-09-12

First-party disclosures land in the same range. John Lewis says product searches arriving through AI agents went from 0.3% to 2.5% of its product searches in a year. Target's CEO told the Q2 2026 earnings call that traffic from external AI platforms grows 3.5x faster than the industry average, but is still small, with no absolute figure given. At the low end, Ethercycle's study of ten Shopify stores found AI referrals drove $75K of $76.5M in revenue — 0.10% — over January–June 2026, on last-click attribution that undercounts influence on a purchase AI did not close.

John Lewis: share of product searches arriving through AI agents

Eight times bigger in a year — and still one search in forty.

0%0.63%1.25%1.88%2.5%A year earlier0.3%Sep 20262.5%
First-party / researchJohn Lewis, reported by Reuters, September 2026.
Show data table
A year earlier0.3%
Sep 20262.5%

What those visitors are worth

The case for caring about 2% of searches is that they are not an average 2%. Shopify's Q1 2026 commerce data found sessions landing on a product page from AI platforms converted about 50% better than the same kind of session from organic search — average gap 56%, holding in 23 of 25 categories — with 14% higher average order value. About half of AI-referred sessions start on a product page at all, against roughly 20% for organic search. That last number is the mechanism behind the other two: the assistant does the browsing and hands over a shopper who has already chosen.

Adobe Digital Insights
Figure 8Adobe Digital Insights: revenue per visit from AI referrals to US retail sites went from 84% below non-AI traffic (Oct 2024) to 33% above it (Dec 2025).Source: Adobe Digital Insights · via Alex Groberman, X 2026-04-20

Adobe measures something different and lands in the same direction: AI-referred retail traffic converting 42% better than non-AI traffic in March 2026, 60% better in July 2026, revenue per visit 53% higher. Read the baselines, not the headline. Shopify compares AI with organic search on product-page-landing sessions only; Adobe compares AI with every other channel a retail site tracks, at site level. In both, the AI group is made of people who chose to shop through an assistant. That is a selection effect, not a measured lift — nobody in this corpus randomised anything.

Handle with care

"Shopify says AI sessions convert ~80% better" appears in newsletters with no link. Shopify's published Q1 figure is ~50%; the 80% looks like a restatement of a secondhand Q2 reading of ~1.8x after a product-page landing. "AI traffic converts 54% better" is attributed to Adobe, but no Adobe release in this corpus says 54%. And Similarweb's 23% is a journey-level rate from a clickstream panel; it cannot be set beside Adobe's or Shopify's session rates.

Journeys that combine AI with search convert best

Journey-level conversion rate. This is not comparable with session-level rates from Adobe or Shopify.

0%6.25%12.5%18.75%25%AI + search23%AI only12.5%
First-party / researchSimilarweb x Statista, State of Ecommerce 2026.
Show data table
AI + search23%
AI only12.5%

Similarweb's finding is the most useful for planning precisely because it is not a channel comparison. Journeys combining AI with search convert at 23% against 12.5% for AI alone — and 89% of AI-using journeys include search anyway. AI appeared in 11.4% of shopping journeys in 2026, up from 4.5% in 2024, and it usually appears mid-journey: 76% of those sessions, against only 23% at the start.

That is the practical shape of the channel. AI is rarely the first touch and rarely the last. It is the research step that used to be a comparison blog, a Reddit thread and four open tabs, compressed into one answer — which is why it arrives with intent attached, and why last-click attribution will systematically understate it.

Which assistant sends them

ChatGPT is still the channel for practical purposes, and it is slipping at the edges. In the ReFiBuy / Digital Commerce 360 AI1000, top-1,000 US retailers whose largest AI referral source is ChatGPT fell from 844 to 722 between Q1 and Q2 2026, while Gemini went 16 → 32, Perplexity 8 → 21, Claude 1 → 15. These are counts of retailers, not shares of traffic.

Whose AI sends the traffic: primary AI referrer of the top 1,000 US online retailers

Number of retailers whose largest AI referral source is each platform. ChatGPT still dominates, but it lost 122 retailers in one quarter.

Q1 2026Q2 2026
02505007501,000ChatGPT844722Gemini1632Perplexity821Claude115
First-party / researchReFiBuy and Digital Commerce 360, AI1000 report, Q1–Q2 2026. Counts of retailers, not traffic share.
Show data table
Q1 2026Q2 2026
ChatGPT844722
Gemini1632
Perplexity821
Claude115

The same index makes a more durable point about who wins. Only 10 of the 100 largest online retailers made the AI1000's top 100, and the new number one, Nixon, ranks 722nd by online sales. 71% of Shopify's AI-attributed orders in 2025 came from long-tail, specialised products. Size on the old shelf does not transfer.

10 July 2026: retrieval moved to the feed

Then the mechanism changed under everyone. Profound analysed roughly 1.75 million ChatGPT shopping prompts and found the share of product recommendations retrieved from integrated merchant feeds went from 8.26% to 61.54% in a single day, with the release of ChatGPT 5.6. Concentration followed: the top ten merchants' share of recommendations rose from 22.5% to 41.8%, distinct merchants shown fell from 13,524 to 10,607, and in the days around 10 July, 450 Profound-tracked brands lost at least a third of their ChatGPT Shopping visibility while 67 gained a third.

One day in July 2026: ChatGPT Shopping switched from the open web to product feeds

Share of product recommendations retrieved from integrated merchant feeds, and the share of recommendations going to the ten largest merchants.

9 July 202610 July 2026
0%20%40%60%80%Recommendations from feeds8.26%61.54%Share held by top-10 merchants22.5%41.8%
First-party / researchProfound, about 1.75M ChatGPT shopping prompts, July 2026.
Show data table
9 July 202610 July 2026
Recommendations from feeds8.26%61.54%
Share held by top-10 merchants22.5%41.8%

Profound tracks its own customer base, and a second reading of the same event matters. Novi ran 1,500 queries across 13 product categories and concluded that feeds decide rendering, not selection: which products the model picks still comes from how it understands them across the web, while the feed supplies the verified price, variant IDs, imagery and stock that let a pick be drawn as a shopping card.

Product feeds do not determine which products the model selects or recommends.

Kimberly Shenk, CEO, Novi — LinkedIn, 4 Sep 2026
Novi's counter-reading of ChatGPT Shopping payloads
Figure 9Novi's counter-reading of ChatGPT Shopping payloads: feed IDs populate variant and offer data on shopping cards, while initial product selection still comes from other sources.Source: Novi · via Novi, LinkedIn 2026-09-03

Both readings point the same way for a merchant: if your product is not in a feed ChatGPT ingests, your best case is a text mention beside competitors who get a card with a price on it. Shopify and Etsy catalogs are already integrated with no application needed, and Profound puts Shopify at about 35% of all feed-integrated retrievals — the specific reason a small Shopify merchant is not automatically locked out of a shelf that just got more concentrated.

The shelf got shorter, and more expensive

Google's side changed shape rather than plumbing. Productrise tracked 2 million listings over 23 days in August 2026: AI Mode shows 3.9 products per answer against 27.8 in the classic shopping carousel, only 1.28% of listings overlapped daily between the two surfaces, and on 49.6% of shared products the top seller differed. The lead offer in AI Mode averaged 21.6% more expensive than classic search for the same product. This is one small firm's study, of the lead listing only, and deserves that caveat every time it is quoted — but the direction matches Ahrefs finding only 38% of AI Overview citations now come from the organic top ten.

The cheapest offer stopped being the default winner.

Lindsey V., growth and performance marketing lead — LinkedIn, 4 Sep 2026

That is the honest read of a four-product shelf selected on undisclosed criteria. A brand that only ever won on price has lost its mechanism; a brand with reviews, complete attributes and third-party coverage has a shot it never had in a 28-product carousel sorted by price.

What a merchant can actually control

Very little of the ranking, and most of the inputs. Google's Merchant Center AI performance report, in beta and expanding through 2026, names the AI shopping intents a merchant appears for and the product attributes shoppers ask about that the feed does not contain.

Merchant Center lists grouped AI shopping terms and the product attributes shoppers ask about that the feed lacks (here
Figure 10Merchant Center lists grouped AI shopping terms and the product attributes shoppers ask about that the feed lacks (here: year, language).Source: Google Merchant Center · via Brodie Clark, X 2026-07-14

Feed work is the only lever here with a before-and-after measurement attached, and the measurement is mixed: in a four-feed audit, adding missing product images lifted organic clicks by 18 to 23 percentage points, fixing vague colour values did nothing measurable, and category fixes were too few to judge. Treat "complete the feed" as a portfolio of bets, not one bet.

Three controlled feed fixes
Figure 11Three controlled feed fixes: more product images lifted organic clicks +18 to +23pp; fixing color had no effect; category fixes were too few to judge.Source: not shown · via John Caiozzo, X 2026-09-08

It also helps to see what an assistant actually receives. Pulled from ChatGPT's retrieval stream, a product page arrives as raw meta tags plus a flattened markdown render — no layout, no alt text, no links. Design decisions do not survive the trip. Title text, attribute values and structured data do.

What ChatGPT's retrieval actually receives from a page, captured from its SSE stream
Figure 12What ChatGPT's retrieval actually receives from a page, captured from its SSE stream: raw meta tags plus a stripped markdown render with no links, alt text or layout.Source: Peec AI · via Metehan Yesilyurt, X 2026-09-07
  1. Get into a feed, then check what it renders

    Shopify and Etsy catalogs reach ChatGPT without an application. Confirm your products come back as shopping cards, not only as text, with correct price, variant and stock.

  2. Fill the attributes shoppers ask about

    Merchant Center's AI report names the missing ones for your own catalog. Images first — that is the fix with measured click impact behind it.

  3. Work the third-party surface, not only your own pages

    Avenue Z's beauty study of 565 ChatGPT citations found 96.1% came from third-party sources, 3.9% from brand-owned content, editorial alone 63.2%. Ahrefs' study of 75,000 brands found brand mentions predict AI citation about 3x better than backlinks — correlational, but consistent with it.

  4. Instrument it before you argue about it

    Shopify's sessions-by-referrer report separates AI sources. Baseline conversion and AOV by referrer now, so the next platform change is something you measure rather than read about.

Operator note

This channel is fragile in a way search never was. Search Engine Land's read of 6.77 million sessions found a single AI product change can halve AI referral traffic overnight, and the 10 July shift moved 450 tracked brands by a third or more within days. Size the work accordingly: fix the feed and the third-party surface, which pay off across every assistant, before buying anything sold specifically to rank inside one of them.

8.26% → 61.54%

ChatGPT Shopping recommendations retrieved from integrated merchant feeds, in one day, 10 July 2026, across ~1.75M prompts.

First-party / research Profound · post

~+50%

Conversion gap, product-page-landing sessions from AI platforms versus organic search, Q1 2026: average gap 56%, holding in 23 of 25 categories.

First-party / research Shopify Q1 2026 commerce data · post

2.5%

Share of John Lewis product searches arriving through AI agents, September 2026, up from 0.3% a year earlier.

First-party / research John Lewis, reported by Reuters · post

76% → 38%

Google AI Overview citations coming from top-10 organic results, 2025 to 2026, across 863,000 searches and 4M links.

First-party / research Ahrefs · post

The conclusion is smaller and more concrete than the discourse suggests. AI discovery is not replacing search: search still drives about a third of Shopify storefront sessions and grew 1.3x over two years, and almost every AI journey passes through it anyway. What AI has taken is the comparison step — the place where a shopper decides which product to want.

OpenAI doesn't need the checkout button if it owns the moment the shopper decides what to buy. Whoever owns intent owns the customer.

Mark W Lewis, founder, Netalico — LinkedIn, 10 Sep 2026

Chapter 03

Agentic commerce has rails, a fee fight, and almost no volume

The first in-chat checkout to reach mainstream merchants was switched off in March 2026, after Walmart measured purchases completed inside ChatGPT converting about three times worse than sending the same shoppers to Walmart.com. Almost everything shipped since is a protocol, a blueprint or a beta.

What holds up

  • The one merchant-side conversion number here from a party with money at stake is negative: Walmart's in-chat Instant Checkout converted about 3x worse than click-out to its own site, 2025–Mar 2026 (Walmart EVP Daniel Danker, via news reports). One retailer.
  • Consumers research with AI and pay with a brand they know. Ipsos: 27% of AI-aware consumers use AI for product research, 9% let it buy autonomously (Ipsos, 2026). Visa: 23% trust GenAI to pay, 61% when Visa is named as handler (Visa Trust Index, n=2,065).
  • Preparation ran ahead of volume: 89% of merchants say they are preparing for agentic commerce; agents are in about 3% of transactions (compiled from Adyen, Harris Poll and Bain, Sep 2026; the 3% definition is unpublished).
  • Both live protocols leave you as merchant of record, where chargeback liability attaches. Anthropic's launch numbers (carts up to 35% larger, +60% completion) are early-partner vendor claims, rated low, with no disclosed sample.

Three trades get filed under "agentic commerce". An AI recommends and a human buys. An AI assembles the cart and a human taps pay. An AI holds a credential and buys alone. The first is a real channel, the second is shipping, the third is a rounding error — and the distance between them is this chapter.

Agent-mediated influence is scaling faster than agent-mediated checkout.

Molly Schonthal, agentic commerce consultant — LinkedIn, 18 Aug 2026

What actually transacts today

The useful question is not which protocol wins but which surface can take an order this quarter. Sorted that way the field thins fast.

SurfaceWhat it does for a merchant todayStatus
Google AI Mode / GeminiCheckout inside the answer, over UCP; you stay merchant of recordLIVE
ChatGPTDiscovery and comparison over ACP, then handoff to your site. Instant Checkout ended 4 Mar 2026LIVE, discovery only
Square in ChatGPT and ClaudeRestaurant ordering; eligible US Square sellers opted in automaticallyLIVE, narrow
Meta MuseUS launch 8 Sep 2026. Browser, Shopify-over-UCP with Shop Pay, or Stripe LinkSHIPPED, six days old
Microsoft CopilotUCP onboarding in Merchant Center; Copilot Checkout in betaSHIPPED, unproven
Claude Commerce AgentsForkable blueprint for an agent on your site and back officeREFERENCE IMPLEMENTATION
Agent-to-agent payments (x402, MPP)Machines paying machines, no human in the loopDEMO

The checkout that died

OpenAI and Stripe launched Instant Checkout in September 2025 and ended it on 4 March 2026. Shopify's count, secondhand, was roughly twelve merchants actually using it; an unattributed post says about thirty integrated with sales near zero. Both can hold — "integrated" and "in use" are different tests. It reportedly took about 4% from merchants. One week before the shutdown, OpenAI and Amazon announced a $50B partnership; reading cause into that sequence is inference, not reporting.

ChatGPT Instant Checkout confirmation
Figure 13ChatGPT Instant Checkout confirmation: a Glossier fragrance bought inside the chat, with Glossier as seller of record.Source: OpenAI; Glossier · via Alex Groberman, X 2026-03-20

That screen is what a working agentic purchase looked like — a Glossier fragrance bought in the chat, Glossier as seller of record. It worked, and it still lost to a redirect. The technical problem was solved a year ago; the conversion problem was not. ChatGPT now discovers and hands the shopper to you, which one fashion operator called affiliate marketing with a chat interface on top.

The protocol map, minus the war

UCP, co-built by Google and Shopify, covers discovery, cart, checkout, order status and identity linking across four transports; its Tech Council added Amazon, Meta, Microsoft, Salesforce and Stripe in April 2026. ACP, from OpenAI and Stripe, is now mostly a discovery format with a handoff. AP2 handles rules-based agent payments. Muse speaks several at once.

UCP architecture
Figure 14UCP architecture: services, capabilities (catalog, cart, checkout, order, identity linking), extensions and four transports (REST, MCP, A2A, embedded) linking agent and merchant platforms.Source: Google / Universal Commerce Protocol · via Tushar K., LinkedIn 2026-04-26

The diagram is worth ten minutes because it shows where your obligations sit: capabilities you must expose, transports you may choose. It says nothing about adoption, and adoption is the weak leg.

UCP has 51 live merchants. Growth is stagnating. The infrastructure is real. The adoption isn't.

Philip Qiu, co-founder, Henry Labs — LinkedIn, 5 May 2026

That count is unattributed and four months stale — a snapshot, not a running total. The harder evidence against the war framing is what the machine-to-machine rails move.

Bernstein/Artemis
Figure 15Bernstein/Artemis: real agent-to-agent payment volume on x402 peaked at $5.1M in Nov 2025 and was under $1M a month through 2026; MPP near zero. Still proof-of-concept.Source: Artemis, Bernstein analysis · via Arvind Srinivas, X 2026-09-12

Bernstein's read of Artemis data puts real agent-to-agent volume on x402 at a $5.1M peak in November 2025 and under $1M a month through 2026, MPP near zero. The same agency posts that count ten competing agentic protocols are counting them over a flow smaller than one mid-size store's year.

Handle with care

A much-reshared agency post prices ACP at 7.2% per order in combined platform and processing fees, and claims brands running both protocols capture up to 40% more agent order volume. Both are unattributed. The 7.2% is most likely the ~4% platform fee plus ~3% card processing restated as one number; the 40% has no visible basis. Do not build a case on either.

Fees, and who holds the bag

Where a company names its fee, it is modest. Square charges its standard ~2.9% + $0.30 on restaurant orders placed inside ChatGPT and Claude, no added marketplace commission, against the 15–30% delivery marketplaces take. Instacart inside Gemini shows the shopper a 4% service fee at handoff.

Instacart inside Gemini builds a grocery cart from a meal-plan request, then hands off to "Checkout on Instacart" with a 4% service fee shown
Figure 16Instacart inside Gemini builds a grocery cart from a meal-plan request, then hands off to "Checkout on Instacart" with a 4% service fee shown.Source: Google Gemini / Instacart · via 🤖 Matthew Bouchner, LinkedIn 2026-06-22

Watch what that Gemini screen does: it builds the basket in Google, then bounces to "Checkout on Instacart". The disclosed fee is the price of the handoff, not of the checkout. Agent assembles, owner of the payment relationship closes — that is the pattern earning money in 2026.

Liability did not move with the interface. ACP states the merchant remains merchant of record and keeps refunds, chargebacks and compliance; UCP takes the same position. What changed is that the buyer is now a probabilistic system, and most order records were never built to record what the human authorised the agent to do.

When a customer disputes what their agent bought, who produces the evidence?

Richard Emanuel, Claude and n8n partner — LinkedIn, 9 Sep 2026

Operator note

Before enabling an agentic channel, decide what you store per order: agent identity, the approval event, the constraint the shopper set (budget, delivery date), and the product state shown at the moment of consent. You need it at the first dispute, not the hundredth.

The trust ceiling is real; the surveys disagree about where

This corpus holds at least a dozen 2026 numbers for "would you let an agent buy", from 7% to 85%. They are not measuring the same act. Seven of them, and what each actually asked:

ValueWhat was actually askedSource
7%Trust an AI platform to run a purchase end to endCompiled: Harris Poll, Bain
9%Currently allow AI to purchase autonomouslyIpsos, AI-aware consumers
23%Trust GenAI to pay, no brand namedVisa Trust Index, n=2,065
42%Would delegate if capped at $250 with 7-day returnsRTB House, US millennials
61%Accept it when Visa is the named handlerVisa Trust Index, n=2,065
74%Would trust an agent over their best friend to buyAccenture, hypothetical
85%Open to collaborating with a shopping agentAccenture

Consumers use AI to research, not yet to pay

Two surveys, kept in their own pairs. The Visa pair asks the same people about the same act with and without a named payment brand.

0%20%40%60%80%Visa · accept when Visa is named as handler61%Ipsos · use AI for product research27%Visa · trust GenAI to pay (unbranded)23%Ipsos · let AI buy autonomously9%
First-party / researchIpsos "Shopping with AI" (AI-aware consumers); Visa Trust Index for agentic commerce (2,065 US consumers, 2026).
Show data table
Visa · accept when Visa is named as handler61%
Ipsos · use AI for product research27%
Visa · trust GenAI to pay (unbranded)23%
Ipsos · let AI buy autonomously9%

Read the pairs, never the spread. The Visa pair is the most instructive line in the chapter: same people, same act, 23% to 61% on the strength of a named payment brand. Bain's 3x higher trust in a retailer's own agent than a third-party bot points the same way. Trust attaches to whoever takes the money.

Meta Muse App Store listing
Figure 17Meta Muse App Store listing: the agent proposes an order on a merchant site and the user must tap Allow before Link-card payment; browser-based booking runs the same way.Source: Meta (App Store listing) · via Juozas Kaziukėnas, LinkedIn 2026-09-07

Meta's App Store listing for Muse shows the compromise everyone landed on: the agent proposes, the human taps Allow, then Link pays. Stripe says Muse checks out instantly at more than a million Link businesses — every one still behind a human tap. That is the honest state of the art, and why 3% of transactions and 89% preparing are both true.

Connect now, wait on the rest

Microsoft Merchant Center UCP settings
Figure 18Microsoft Merchant Center UCP settings: merchants must supply return policy and support contacts to be eligible for AI shopping, and can toggle Copilot Checkout (beta).Source: Microsoft Advertising · via Glenn Gabe, X 2026-04-26

Microsoft's Merchant Center states the real entry fee plainly: publish a return policy and support contacts and you are eligible for AI shopping surfaces. The work is hygiene, not integration.

  1. Now: the free default rails

    UCP through Shopify Catalog costs nothing and is on by default; Square sellers are opted into ChatGPT and Claude ordering automatically. The work is catalog completeness, live inventory and price, a return policy, a support contact.

  2. Now: the dispute record

    Log agent identity, approval event and stated constraint on every agentic order. One practitioner's read of the timing: Visa's VAMP excessive-chargeback threshold dropped from 2.2% to 1.5% on 1 April 2026 for US, Canada, EU and APAC merchants. Your room for a new dispute class shrank this year.

  3. Wait: integrations sold on fee arbitrage

    Anything priced against the unattributed 7.2%-versus-zero framing or promising 40% more volume for running both. If your platform already speaks UCP, a second direct integration buys little at 3% of transactions.

  4. Wait: agent-to-agent payment rails

    Under $1M a month across all of x402 in 2026. Revisit when a named merchant publishes a volume number.

Keep the vendor numbers in the vendor column

Anthropic's materials say retailers running Claude shopping agents saw carts up to 35% larger and shoppers 60% more likely to complete. Twenty posts here repeat them; most drop the "up to", several swap "completion" for "conversion". The underlying disclosure is one early partner self-reporting 30–35%, no sample size, no stated comparison group. A separate +40% conversion case in the same materials is credited to a startup called Vambe.

They claim 40% increase in conversions, and then when you find the source, it's some no-name startup called "Vambe" whose website currently reads "404: NOT_FOUND Code: DEPLOYMENT_NOT_FOUND."

Kelly Goetsch, President, Pipe17 — LinkedIn, 3 Sep 2026

That is not an argument against the blueprint, which is useful engineering, free and forkable. The discipline cuts both ways: unattributed posts claim Amazon's Buy for Me covers 100M products across 400,000 merchants, with 170+ merchants organising legal action and chargebacks up 28%. None of it is sourced. Both sets of numbers are selling something.

3x worse

Conversion on purchases completed inside ChatGPT via Instant Checkout, versus click-out to Walmart.com. One retailer, 2025 to Mar 2026.

First-party / research Walmart EVP Daniel Danker, via news reports · post

23% → 61%

US consumers who trust generative AI to pay for them, unbranded versus with Visa named as handler. 2026, n=2,065, Visa-commissioned.

First-party / research Visa Trust Index for Agentic Commerce · post

89% vs ~3%

Merchants preparing for agentic commerce, against the share of transactions agents touch. Sep 2026; the 3% definition is unpublished.

Named but secondhand Compiled: Adyen, Harris Poll, Bain · post

2.9% + $0.30

Square's fee on restaurant orders placed inside ChatGPT and Claude, no added marketplace commission, against 15–30% on delivery marketplaces. Jul 2026, US.

Named but secondhand Square announcement, via news · post

Chapter 04

Shopify switched most of it on for you

Agentic Storefronts, Catalog and the agent files arrived on by default; opting out is the only decision most merchants were given. And the multiple everyone still quotes — 13x — belongs to Q1 2026 and had fallen to 3x by Q2, for reasons that are arithmetic rather than a slowdown.

What holds up

  • Shopify disclosed AI-driven traffic up more than 8x and AI-driven orders up nearly 13x year over year in Q1 2026, then about 3x each in Q2 2026. Both are Shopify's own figures; the fall is a base effect (Q1, Q2).
  • Shopify integrations account for about 35% of feed-integrated product retrievals in ChatGPT Shopping, July–September 2026 — measured by Profound, not by Shopify (post).
  • Sidekick adoption is real and unglamorous: up 4x year over year in Q1 2026, daily active merchants up 3.6x in Q2, top request category SEO and meta tags — 1.1M asks in 30 days (Shopify, DotDev 2026).
  • No Shopify disclosure says what these channels are worth on one store. The only real merchant dashboard in the corpus: ChatGPT $5,673.68 in 30 days against $52,483.10 through the Shop channel (post).

Nobody installed this

In March 2026 Shopify began turning Agentic Storefronts on for eligible stores. Aleyda Solis noted at the time that the channel was active by default, with permissions sitting in Settings → Sales channels — an opt-out screen, not an install flow (X, 25 Mar 2026). Six months on, the same screen lists ChatGPT, Microsoft Copilot, Google and Shop, and since 8 September, Meta.

Shopify admin's Agentic page
Figure 19Shopify admin's Agentic page: merchants enable catalog access and policies, let Shopify manage ChatGPT, Microsoft Copilot and Shop channels, and add a Knowledge Base source for agents.Source: Shopify · via Brodie Clark, X 2026-05-14

The plumbing under that screen is what makes the default stick. Shopify Catalog syndicates structured product data to AI channels, and Shopify's help documentation states plainly that blocking crawlers in robots.txt does not stop it — different pipes. Every store also serves /agents.md, /llms.txt and /llms-full.txt without anyone creating them. Read the documentation below as the list of what a merchant cannot switch off with the tools they already know.

Shopify documentation
Figure 20Shopify documentation: Catalog syndicates structured product data to AI channels; blocking crawlers in robots.txt does not stop it; every store serves /agents.md, /llms.txt and /llms-full.txt.Source: Shopify Help Center · via Selimhan Tokgoz, X 2026-09-12

That default has a measurable consequence off-platform. Profound, studying ChatGPT Shopping retrievals from July to September 2026, put Shopify integrations at roughly 35% of feed-integrated product retrievals. Shopify's own matching claim — AI searches drawing on structured Catalog data convert at 2x the rate of AI searches relying on scraped or outdated web data, Q1–Q2 2026 — has never had its comparison set published (post).

Every new channel repeats the pattern. September's Meta launch is the clean example: products go to Meta by default via Catalog, so the merchant's decision is an opt-out they have to notice first.

Shopify changelog, 8 Sep 2026
Figure 21Shopify changelog, 8 Sep 2026: Meta joins Agentic Storefronts, and products are shared with Meta by default via Shopify Catalog — merchants must opt out, not in.Source: Shopify Changelog · via jaime, X 2026-09-12
SurfaceHow it arrivesWhat is left to decide
Agentic Storefronts channels (ChatGPT, Copilot, Google, Shop, Meta)On by default for eligible storesSwitch channels off; set catalog and policy permissions
Shopify Catalog syndicationOn by defaultOpt out per channel — robots.txt will not do it
/agents.md, /llms.txt, /llms-full.txtServed automaticallyNothing to install; read what it says about you
Sidekick app intentsPer app, only if the publisher ships oneApprove each data read
AI Toolkit, ChatGPT and Claude connectorsMerchant installsScope which tools need approval
UCP and Catalog for your own agentPublic since 18 May 2026Build on it, or ignore it

The 13x belongs to Q1 2026

Shopify said AI-driven traffic to merchant stores was up more than 8x and AI-driven orders up nearly 13x year over year in Q1 2026. Three months later, on the Q2 2026 earnings call of 5 August, the same two lines came in at about 3x each. Together they look like a collapse. They are not.

Shopify's AI growth multiples: Q1 was a small base, Q2 is the truer read

Year-on-year multiples disclosed by Shopify for two consecutive quarters.

AI-driven trafficAI-driven orders
3.75×7.5×11.25×15×Q1 202613×Q2 2026
First-party / researchShopify Q1 2026 commerce data and COO remarks; Shopify Q2 2026 earnings call, 5 Aug 2026.
Show data table
AI-driven trafficAI-driven orders
Q1 202613×
Q2 2026

A year-over-year multiple divides this quarter by the same quarter a year earlier. Q1 2025 was close to nothing, so almost any volume in Q1 2026 produced a large number. By Q2 2025 the base had grown, so Q2 2026 was divided by a bigger denominator. A multiple falling from 13x to 3x is fully consistent with absolute volume still rising — it is what the first year of a new channel looks like once the denominator catches up.

The practical version: a September slide reading "AI orders are up 13x", with no "Q1 2026" attached, is quoting a stale quarter as current. Three further versions circulate — 15x for full-year 2025 (secondhand, no link), 7x traffic and 11x orders from an older disclosure with no period attached, and an unattributed 385% for Sidekick sitting awkwardly beside Shopify's disclosed 4x. Only the figures with a quarter attached survive checking.

~13x

AI-driven orders to Shopify stores, year over year, Q1 2026. Traffic over the same period: more than 8x.

First-party / research Shopify Q1 2026 earnings call · post

~3x

AI-driven traffic and AI-driven orders, each, year over year, Q2 2026 — a bigger base, not a smaller channel.

First-party / research Shopify Q2 2026 earnings call, 5 Aug 2026 · post

~35%

Share of feed-integrated product retrievals in ChatGPT Shopping traced to Shopify integrations, Jul–Sep 2026.

First-party / research Profound · post

One real dashboard

Shopify's disclosures are multiples and rates. None says what the channel is worth on one store, and the Agentic Storefronts graphic that circulated after the June launch is no help: its store name and its $103,813 read as sample data in a launch image. The only real dashboard in this corpus is a client store Kurt Elster posted in May 2026 — ChatGPT $5,673.68 over 30 days, Copilot $301.79, next to $52,483.10 through the Shop channel. Read the channel lines, not the tile's headline total, which does not reconcile with them; the "0.27% of revenue" in Elster's caption cannot be checked from the screenshot.

A real merchant's Agentic Storefronts dashboard
Figure 22A real merchant's Agentic Storefronts dashboard: ChatGPT drove $5.7K in 30 days (+42%), Copilot $302, against $52K via Shop Channel — AI channels are small but growing.Source: Shopify admin · via Kurt Elster, X 2026-05-12

It's a small slice with a steep curve. The right approach here is "monitor monthly" not "overhaul everything."

Kurt Elster — X, 12 May 2026

At this size the channel differences matter more than the totals. Google, Copilot and Meta close checkout inside the AI chat, where browser pixels, bundles and subscriptions do not apply; ChatGPT redirects the shopper to the merchant's own Shopify checkout (Stack Architect, X, 24 Aug 2026). A store whose margin rests on post-purchase upsells is not "live on AI" in the same sense across all five. Nor was the rollout clean: in March, Ben Kennedy found every store he tested returning a "Product not available" error after launch was announced.

Sidekick's most common job is meta tags

Sidekick usage was up 4x year over year in Q1 2026 per Shopify's COO and the Spring '26 Edition, and daily active merchants grew 3.6x in Q2. At DotDev 2026 Shopify showed the request mix over 30 days, and its shape says more than the growth rate does.

Shopify DotDev 2026 slide
Figure 23Shopify DotDev 2026 slide: Sidekick requests by topic over 30 days. SEO and meta tags lead at 1.1M, ahead of fulfillment (716k) and inventory (646k).Source: Shopify (DotDev 2026 slide) · via M Asif Rahman, X 2026-07-23

SEO and meta tags lead at 1.1M requests, ahead of fulfillment (716k) and inventory (646k). Merchants are not asking an assistant for strategy; they are handing it the copy work they were already doing badly. The structural change is app intents: Sidekick now reaches into AfterShip, Consentmo, Klaviyo, Judge.me and Loop, asking permission before each read, and since an August changelog it opens the app's own page full-screen when it invokes one — moving app discovery out of the App Store and into the assistant (post). The supply side answered: App Store review went from 40 days to 4 (Atlee Clark, VP Partnerships, Shopify), and May 2026 alone brought over 2,000 new apps, 27% mentioning AI (App Store Pulse). What none of it captures is the failure mode operators complain about.

Build a segment in Shopify. Well you can't. Sidekick has to do it. Turns out that it can't build a simple segment of "Customers that have purchased only..."

Calvin — X, 8 Sep 2026

The connector is really a permissions screen

In May 2026 Shopify shipped connector apps for ChatGPT and Claude so merchants could run the store from a chat window they already had open. Harley Finkelstein's framing was a merchant survey finding 83% already use ChatGPT (LinkedIn, 4 May 2026) — method undisclosed, so read it as direction, not measurement. What deserves study is not the demo but what the connector exposes.

Shopify connector inside Claude
Figure 24Shopify connector inside Claude: 19 interactive and 4 read-only tools, with GraphQL mutations and shop switching gated behind 'needs approval' — how merchants scope what an agent may change.Source: Claude / Shopify · via James Lee, X 2026-05-22

Nineteen interactive tools, four read-only, with GraphQL mutations and shop switching held behind a "needs approval" gate. That list is the real governance surface of the whole stack: it decides what an agent may change on a live store without a human present. Evaluate that screen and the app-intent approvals; the marketing page can wait.

Handle with care

Shopify-flavoured numbers with nothing behind them, excluded here: Magic/Sidekick at "65% merchant penetration", Audiences cutting CAC 22–38%, native recommendations lifting AOV 12–25%; "~7 million stores got agents.md in May", quoted alongside agencies charging $3,000 a month to set up a file Shopify generates for free; and a rumoured $500–$1,500/month "Canvas" tier Shopify has not announced. The claim that AI-referred sessions "convert about 80 percent better than organic search" has no traceable source either — Shopify's published figure for that comparison is roughly 50%, Q1 2026, and only for sessions landing on a product page. See Chapter 11.

Operator note

All of this is already running on your store. The work is inspection, not installation.

  1. Open the Agentic channel list

    Settings → Sales channels. Check the Meta row specifically — default-on since 8 September 2026.

  2. Read your own attribution report

    AI-referred orders already land in the admin with the referral attached. Zero is a store to investigate, not a store without AI demand.

  3. Fetch your own agent files

    Open /agents.md and /llms.txt on your domain. That is the description agents see, and you did not write it.

  4. Decide checkout per channel

    If your margin needs post-purchase upsells, subscriptions or browser pixels, in-chat checkout is not the same product as ChatGPT's redirect to your checkout.

Chapter 05

The surface you own, and the ruler that flatters it

Retailer-owned assistants and on-site personalization are forecast to carry more of this year's AI-driven US retail sales than every external chatbot combined. Almost every headline number attached to them compares shoppers who opened the assistant with shoppers who never did — a description of who uses the tool, not of what the tool caused. Asked directly, shoppers rank personalization last among the things that decide where they buy.

What holds up

  • EMARKETER expects retailer-native assistants to drive 54.1% of US AI-driven retail ecommerce sales in 2026. It is a forecast, not a measurement — but it points at the surface merchants actually control (post).
  • Walmart's CEO told the Q2 FY27 call (August 2026) that Sparky users spend about 40% more per order than non-users, with Sparky users up 70% year over year. Amazon's CEO gave the same shape of number in its Q2 2026 call: US assistant users spend over 40% more than non-users (post).
  • Neither of those is a lift figure, and neither company claimed it was. No retailer in this corpus published a holdout or controlled test for its own assistant.
  • Merchant-scale personalization results land an order of magnitude lower and come from vendors and agencies: DFS +10% online conversion and +8% AOV, LOOKFANTASTIC +5% revenue, Vocca +14.49% AOV (post).
  • Shoppers do not weight this the way the industry does. In Digital Commerce 360’s 2026 holiday survey, recommendations and personalization came last at 9.9% among influences on where to shop, behind price and discounts at 81.8% (post).
  • The bottleneck is underneath the assistant: 52% of retailers still rely on out-of-the-box search and only 8% feel ready for AI agents to navigate their sites (post).

The one AI surface a merchant fully controls

The previous three chapters are about places somebody else sets the ranking. Your own domain is not one of them. You choose what search returns, you see the whole session, the customer record stays with you, and nothing about the experience depends on a protocol being ratified. EMARKETER's forecast puts the money there too: retailer-native assistants — Rufus, Sparky and their kind — are expected to drive 54.1% of US AI-driven retail ecommerce sales in 2026, more than every general chatbot combined. Shoppers report the same preference: Evercore's agentic commerce survey found 55% use a retailer's built-in assistant against 48% who use a general AI tool, and 66% were satisfied against 3% dissatisfied.

In practice the surface is less exotic than the launch videos. Constructor's Ask Cleo sits on the Rugs Direct search results page, answers "which of these is actually washable", and hands back product cards with an offer to narrow by shape — a retrieval layer with a mouth on it, anchored to a catalog the merchant already maintains.

Constructor's Ask Cleo on Rugs Direct opens over search results, explains washable rugs and returns product cards with an offer to narrow by theme or shape
Figure 25Constructor's Ask Cleo on Rugs Direct opens over search results, explains washable rugs and returns product cards with an offer to narrow by theme or shape.Source: Constructor (Ask Cleo on Rugs Direct) · via Constructor, LinkedIn 2026-09-10

The numbers everyone is quoting this quarter

Four disclosures set the tone of the whole conversation. Walmart's CEO John Furner told the Q2 FY27 earnings call in August 2026 that Sparky users spend roughly 40% more per order than customers who do not use it, and that Sparky users were up 70% year over year. Amazon's Andy Jassy said in the Q2 2026 call that US customers using its shopping assistant spend over 40% more than non-users, and that 350 million shoppers used it in the past year. Lowe's says shoppers who engage Mylow convert at about three times the rate of those who do not, across 25 million questions answered. Williams-Sonoma says its Olive assistant converts at 3x the normal rate, with revenue through it up 620% and engagement up 700% since January 2026.

Below the enterprise tier the same claim repeats with bigger multiples and thinner sourcing. Constructor reports that Rugs Direct and Lightopia shoppers who engage Ask Cleo add to cart at 8–9x the rate of other shoppers. Nosto claims a +92% conversion increase in A/B tests of an LLM assistant wired to its intent model; Loops AI claims +83% conversion and +77% add-to-cart.

Retailer AI assistants: how much more users spend than non-users

These compare shoppers who chose the assistant with shoppers who did not. Self-selection, not a measured lift.

0%10%20%30%40%Walmart Sparky (per order, Q2 FY27)40%Amazon AI assistant (US, Q2 2026)40%Albertsons (AOV, secondhand)26%
Named but secondhandWalmart and Amazon Q2 2026 earnings calls; Albertsons via newsletter (low reliability).
Show data table
Walmart Sparky (per order, Q2 FY27)40%
Amazon AI assistant (US, Q2 2026)40%
Albertsons (AOV, secondhand)26%

Read that chart as a chart of one methodology, not of three results. Every bar is the same subtraction: the average of people who used the assistant minus the average of people who did not. The Albertsons bar is the weakest — secondhand, its metric undefined, and described in at least one account as experience-versus-experience rather than user-versus-non-user.

Walmart CEO on the earnings call
Figure 26Walmart CEO on the earnings call: Sparky users up 70% year over year, and they spend 40% more per order — a correlation, not proof Sparky causes the lift.Source: Walmart earnings call transcript · via Wall St Engine, X 2026-08-20

What the gap actually measures

A shopper who opens an assistant to plan a week of dinners was already the bigger basket: further down the funnel, more likely logged in, doing a stock-up trip rather than grabbing one item. The assistant did not create her; it attracted her. The 40% is the selection effect plus whatever the tool contributed, and nothing in the disclosures separates the two.

Both can be true. Neither is a lift. Those groups are not the same shoppers. Whoever opens an AI assistant to plan a week of dinners was already the bigger basket. The number describes who uses the tool, not what the tool caused.

Manish Sharma, omnichannel commerce executive — LinkedIn, 9 Sep 2026

Handle with care

Sparky +40%, Mylow 3x, Olive 3x, Ask Cleo 8–9x, Amazon "over 40%", Rufus "60% more likely to buy" — every one of these is a user-versus-non-user comparison. The tell is the size: an 8–9x gap in add-to-cart rate is not a plausible causal effect of a chat widget, it is a near-perfect description of intent. Use them as evidence that engaged shoppers exist and are findable, never to forecast what an assistant will do to your conversion rate.

Walmart's own number has a second lesson in it. The 40% is an increase on the 35% premium the company gave in Q1 FY27, three months earlier. Retellings that quote a single figure lose the trend, and several restate the Q1 number as "baskets 35% larger", which is a different metric from spend per order. When a claim survives three retellings, check which quarter and which denominator it started in.

Walmart Sparky: the spend premium its users carry, quarter to quarter

Two disclosures three months apart. Retellings that quote only one number drop the trend.

0%10%20%30%40%Q1 FY2735%Q2 FY2740%
First-party / researchWalmart investor remarks and Q2 FY27 earnings call. Users self-select.
Show data table
Q1 FY2735%
Q2 FY2740%

The clearest illustration of the trap is a vendor slide. Loops AI reports A101 Ekstra visitors who use its assistant converting at 37% against 21% for standard visits. Printing both numbers is more honesty than most vendors offer, and it lets you check the arithmetic: 21% to 37% is a 76% increase, not the +83% quoted elsewhere in the same campaign, while a companion card calls assistant sessions 2x. Three numbers, one dataset.

Vendor-reported
Figure 27Vendor-reported: A101 Ekstra visitors who use the Loops AI assistant convert at 37% vs 21% for standard visits. Self-selected users, not a controlled test.Source: Loops x A101 · via Loops AI (a16z SR006), LinkedIn 2026-09-03

Survey evidence gets closer to causation without reaching it. Evercore found 57% of Amazon assistant users bought a product they had not been considering — a genuine discovery effect, self-reported. But in the same survey only 36.4% said the assistant raised their Amazon spending, while 44.5% said it was flat. Discovery moved. Total spend mostly did not. That is a shift in which product gets bought, not a shift in how much gets spent, and for a brand those are opposite problems.

Evercore ISI survey
Figure 28Evercore ISI survey: 33.4% have used Amazon's shopping assistant; 57% of users bought something they had not known about; 36.4% say they now spend more on Amazon.Source: Evercore ISI Proprietary Survey · via ReFiBuy.ai, LinkedIn 2026-09-08

What shoppers say decides where they shop

One finding cuts against the premise of the category. Digital Commerce 360's 2026 holiday report asked shoppers what would influence where they shop this season. Recommendations and personalization came last, at 9.9%. Price and discounts led at 81.8%, fast or free shipping at 54.6%; even "website experience" edged personalization out, at 10.3%. Stated preference is not behaviour, and nobody notices that a collection page was reordered — but it caps what this work can be sold as internally. A store with a thin margin and a slow courier does not have a personalization problem. Put the letters AI in the question and the answers warm up — a vendor survey with eTail Insights has 65% of shoppers planning to use AI for part of their holiday shopping, 78% if it were more personalized — but only the first question is attached to a purchase (post, post).

Maybe good personalisation isn't something customers are supposed to notice at all. It just makes those answers easier to find.

Liliia Maliutina, ecommerce consultant — LinkedIn, 2 Sep 2026

Peer-reviewed work goes further. A study of 299 fashion ecommerce customers in the Journal of Retailing and Consumer Services found no direct link between recommendation quality and post-purchase satisfaction; perceived value carried the relationship, and dissatisfaction — not recommendation quality — was the key predictor of returns. In a category where returns eat the margin, that moves the case downstream: the number to watch is the return rate, not the click.

The largest controlled effect anywhere in this chapter is not a recommendation at all. A randomized field experiment with 35,625 users, reported by Wang, Zhang, Gao and Tan in September 2026, varied one thing in an AI-generated summary of customer reviews: how strongly it framed them as customer consensus. Product clicks rose 24%; time spent reading fell 10%. It is an SSRN working paper, not yet peer reviewed, so hold the magnitude loosely — but it is the design nobody else here ran, and the lever it moved was wording, not ranking. Any store already showing an AI review summary is running that experiment on its customers without a control arm (post).

Personalization that is not a chatbot

The results at merchant scale come from ranking and recommendation work, not conversation, and they are smaller and more believable. DFS added AI recommendations and photo-led visual search and reports +10% online conversion, +8% AOV and a 3% drop in bounce. THG's foundation finder for LOOKFANTASTIC reports a 5% revenue uplift, with 96% of buyers through it purchasing a product they had never bought before. Anphonic reports personalized upsells lifting Vocca's AOV 14.49%. Dr. Squatch put recommendations on its order-tracking page and reports a 31.89% click rate and $32,978 in Q1 2026 revenue from them — a small number, and one of the few in this chapter with a currency symbol and a period attached.

Every figure in that paragraph is vendor- or agency-published, with no controls disclosed. Treat them as a plausible range for what good merchandising work returns — single digits to low double digits — rather than as benchmarks.

Vendor case study
Figure 29Vendor case study: UK sofa retailer DFS added AI recommendations and photo/Pinterest visual search, reporting +10% online conversion, +8% AOV and -3% bounce rate.Source: Rezolve Ai; DFS · via Rezolve Ai, LinkedIn 2026-08-20

The least-discussed version of this is silent. Cooee's category merchandising for Shopify scores each SKU on more than a hundred signals and reorders the collection page: featured and new items up, out-of-stock down. No chat, no persona, nothing for the shopper to opt into — which also means no user-versus-non-user number to quote, because everyone gets it. That is a feature, not a shortcoming, and it is why this kind of change is the easier one to test properly.

Cooee's AI category merchandising for Shopify ranks products by an SKU score built from 100+ signals, pinning featured and new items up and pushing out-of-stock down
Figure 30Cooee's AI category merchandising for Shopify ranks products by an SKU score built from 100+ signals, pinning featured and new items up and pushing out-of-stock down.Source: Cooee · via Shwetank Tamer, LinkedIn 2026-09-09

One operator has published the test. Matthew Bertulli of Pela says he spent a couple of all-nighters with Claude Code building a small version of TikTok's For You page over a collection page: it watched how a visitor scrolled and reordered what came next. One split test, the static page left as the control, profit per session up 28%. His own caveat: "One test isn't a benchmark. It might do nothing on your store." Unaudited — and still one of the few numbers here with a control group behind it, and the only one whose denominator was chosen before the test ran (post).

The biggest production version of the quiet kind also has no chat window. RecGPT — Alibaba Group and Renmin University of China, in a paper accepted by ACM Transactions on Information Systems — runs on Taobao's "Guess What You Like" homepage feed. It throws away plain clicks as noise and keeps purchases, favourites, add-to-carts, detailed views and searches; an LLM compresses that history into an interest profile matched to a taxonomy of 169 interests, refreshed every two weeks. Every LLM stage runs offline in batch; the online path is a retrieval model answering in roughly 25 ms. No revenue lift is published with it, and the architecture is the transferable part: the model is not in the request path, and the input is deliberate behaviour, not clicks (post).

The floor most stores have not poured

Algolia and The Retail Hive's barometer found 52% of retailers still rely on out-of-the-box third-party search tools, and just 8% feel ready for AI agents to navigate their sites on a shopper's behalf. The models are not the constraint; the catalog, the attributes and the query understanding underneath them are.

Nearly half say they'd shop more with grocers that use AI well. But AI won't save you if "birthday cake" returns birthday candles. It'll just be wrong faster.

John A. Stewart, Algolia — X, 8 Sep 2026

The same research found 44% of grocery shoppers have no go-to store, which is the commercial reason to care: the shopper is not loyal, and a bad result page is a live churn event. On the platform side the tooling is arriving fast — Shopify's in-house 0.8-billion-parameter model scaled buyer-profile generation from 2 million to 72 million profiles a day — but for most merchants this arrives through the Shopify App Store, where a search for "personalization" returns 4,008 apps and five of the ten fastest-climbing new apps in one week of September 2026 were upsell tools. Supply is not the constraint either.

54.1%

Forecast share of US AI-driven retail ecommerce sales going through retailer-native assistants in 2026, rather than general chatbots.

First-party / research EMARKETER forecast · post

+40%

Spend per order for Walmart Sparky users versus non-users, Q2 FY27 (August 2026). Users self-select; this is not a measured lift.

First-party / research Walmart Q2 FY27 earnings call · post

8%

Retailers who feel ready for AI agents to navigate their sites on a shopper's behalf; 52% still run out-of-the-box search.

Named but secondhand Algolia & The Retail Hive barometer · post

9.9%

Shoppers naming recommendations and personalization as an influence on where they shop this holiday season — last of the options offered. Price and discounts: 81.8%.

Named but secondhand Digital Commerce 360 holiday report, 2026 · post

Numbers to leave on the shelf

"76% of customers get frustrated without relevant products" and "71% expect personalization" circulate with no attribution. Both trace to a 2021 McKinsey study — five years old, pre-dating every system in this chapter. The companion "57% say shopping is still generic" has no source at all. A deck that opens with these is quoting a deck.

How to get a number you can act on

  1. Fix retrieval before you add a mouth

    An assistant reads the same index your search box does. If "birthday cake" returns candles today, the assistant will say the wrong thing more fluently. Attributes, synonyms and stock accuracy come first.

  2. Hold out, do not compare

    Withhold the feature from a random slice of traffic and compare slices, not users. Every spectacular number in this chapter is a user-versus-non-user comparison; every modest one came closer to a controlled test.

  3. Pick the denominator before you run it

    Conversion per session, per user, or per visitor who reached a product page are three different numbers, and Walmart's 35% drifted into "35% larger baskets" in the retelling precisely because nobody carried the denominator.

  4. Measure substitution, not just addition

    Evercore's 57% who bought something unconsidered against 44.5% whose spend stayed flat is the whole risk in one line: the assistant may be moving which of your SKUs sells rather than how much sells. Carry the return rate beside it: the 299-customer study found dissatisfaction, not recommendation quality, is what predicts a parcel coming back.

Operator note

For a Shopify store the honest sequence is search quality, then collection ranking, then recommendations on high-traffic non-product pages (cart, order tracking, post-purchase), and only then a conversational assistant. The first three are testable with a holdout inside a normal sprint. The fourth is where you will be tempted to quote somebody else's 3x.

Chapter 06

The funnel is where the evidence runs thin

Ten facts reach this chapter's ledger, up from seven. One is rated first-party, and it is still not about AI. A second sweep changed the shape of the evidence, not its volume: after a thousand posts with no control group anywhere, two experiments turned up that had one. Two is not a body of evidence. It is enough to stop writing zero.

What holds up

  • The ledger holds 10 facts: 1 high, 5 medium, 4 low, and still no measured lift number at high reliability — the thinnest evidence in the report. What it holds for the first time is a test with a control group.
  • One store ran the test. Pela rebuilt a collection page as a feed that reorders itself from on-page behaviour and split-tested it against the static page: profit per session up 28% (Matthew Bertulli, LinkedIn, 22 Aug 2026). One test, reported by the co-founder.
  • Agents look and do not buy: 87% of AI-agent traffic to retail sites lands on a product page, barely 2% reaches a checkout (HUMAN Security, via Sreekant Vijayakumar, LinkedIn, 8 Sep 2026).
  • Almost everything else operators published is production, not outcome: page builders, CRO audit agents, bundle prompts. One n8n CRO audit ran in 29 seconds on 2,381 tokens (Yumna Aziz, LinkedIn). The cost of an AI opinion is precise; its value is unmeasured.
  • Shopify put the measuring instruments in the admin in Winter '26 — Rollouts (native A/B testing) and SimGym (AI shopper simulation) (Stephen Taylor, LinkedIn, 9 Sep 2026). Shipped, not proven: no merchant result from either appears in this corpus.
  • Amazon researchers put a number on simulation itself: roughly 1,000 LLM personas built from real behaviour logs called 70–90% of A/B outcomes before either variant met a real user (AIDB summarising the paper, X, 9 Sep 2026).

What ten facts look like

Every other chapter in this report argues about which measured number to trust. This one cannot. The whole ledger still fits in a table, and none of the three new rows is first-party.

ClaimValueSourceReliability
Walmart price rollbacks, quarter on quarter11,000+ vs 7,200Walmart Q2 FY27 earningsFirst-party / research
Shoppers who would reconsider an AI-recommended purchase if the brand's site falls short; who say AI changed which brand they buy97%; 47%Contentsquare research, 2026Named but secondhand
Holiday shoppers who would trade a discount for a guaranteed delivery date70%Radial survey, holiday 2026Named but secondhand
Gross-margin gain from real-time dynamic pricing+12–14%Bonita Technologica (vendor)Vendor or unsourced
TUMLE ads conversion after CRO + AEO work; advertorial page ROAS1.1% → 2.8%; 1.7xagency and operator postsVendor or unsourced
Abandoned carts converted by chatbot upsell; cart recovery with AI-personalised email18%; 3–8% → 10–15%unattributed author claimsVendor or unsourced
pet&home abandoned-checkout revenue recovered by AI phone agents, first 14 days$4,157.70; $41 per $1Ava AI own case studyVendor or unsourced
Profit per session, Pela collection page reordered by on-page behaviour, one split test+28%Matthew Bertulli, Pela co-founderNamed but secondhand
A/B outcomes called correctly, pre-launch, by ~1,000 LLM personas built from behaviour logs70–90%Amazon researchers, via AIDBNamed but secondhand
AI-agent traffic to retail sites reaching a product page; reaching a checkout87%; ~2%HUMAN SecurityNamed but secondhand

Note the one high-reliability row. Walmart cut prices 11,000-plus times in the quarter against 7,200 the quarter before, disclosed on its own earnings call: a real number about promotional pressure that says nothing about AI. Two medium rows are surveys asking people what they would do; the three new ones are single-source. The four low rows are a pricing vendor, two agency cases, unattributed benchmarks, and a startup's first case study.

+28%

profit per session on a Pela collection page reordered live from shopper behaviour, against the static page as control. One split test, Aug 2026.

Named but secondhand Matthew Bertulli, Pela co-founder · post

11,000+

Walmart price rollbacks in Q2 FY27, against 7,200 the previous quarter. The chapter's only first-party number — and not an AI number.

First-party / research Walmart Q2 FY27 earnings · post

97%

of shoppers say they would reconsider an AI-recommended purchase if the brand's website falls short; 47% say AI has changed which brand they buy. 2026.

Named but secondhand Contentsquare research · post

70%

of holiday shoppers say they would trade a discount for a guaranteed delivery date, holiday 2026 — up on last year.

Named but secondhand Radial survey · post

Both survey numbers are stated intent, collected by companies that sell into the problem they describe. Neither is a conversion rate. They point at one thing: the site is where the sale closes, even when the recommendation happened elsewhere.

Production got cheap. Evidence did not.

The published work is almost entirely about making pages faster. ScaleBot shipped a landing-page builder that writes an on-brand lander from scratch or, in its author's words, "rip[s] a competitor page in 1-click", carrying over fonts, logo and colours.

ScaleBot's AI page builder editing a generated advertorial lander, with per-element AI image edits and a direct-edit panel
Figure 31ScaleBot's AI page builder editing a generated advertorial lander, with per-element AI image edits and a direct-edit panel.Source: ScaleBot · via Mike Futia, X 2026-09-10

Read that screenshot for what it proves. The build step is now minutes. It says nothing about whether the page sold anything: no session count, no conversion rate, no control. Every AI page-building post in this corpus stops there.

The audit side is the same story, better instrumented. An n8n workflow fetches a landing page, converts it to Markdown and has Gemini 2.5 Flash return prioritised CRO fixes.

n8n workflow that fetches a landing page, converts it to Markdown and has Gemini 2
Figure 32n8n workflow that fetches a landing page, converts it to Markdown and has Gemini 2.5 Flash return prioritised CRO fixes; one run took 29 seconds and 2,381 tokens.Source: n8n · via Yumna Aziz, LinkedIn 2026-08-19

The run shown took 29 seconds and 2,381 tokens. That is the asymmetry of 2026: the cost of a CRO opinion is known to four significant figures, its accuracy nowhere. A bundle app handles the same trade more carefully — each natural-language prompt lands as a logged set of structure and settings changes, so a merchant can see what the model altered.

Shopify bundle app with an 'AI Studio' panel
Figure 33Shopify bundle app with an 'AI Studio' panel: each natural-language prompt is applied as a logged set of structure and settings changes to the bundle.Source: Rich Bundle Builder · via North Js Tech, LinkedIn 2026-09-03

The logging is the part worth copying. If an AI edits your page, the change set has to stay readable afterwards, or you cannot attribute a result to it even when you run a test.

Producing things is getting cheap fast, but being believed is not.

The Agile Brand with Greg Kihlström — LinkedIn, 4 Sep 2026

One store ran the test

The exception is not a page builder. Matthew Bertulli, co-founder of Pela, spent two nights with Claude Code building what he calls a dumb little version of TikTok's For You page for the brand's bestsellers collection: it watched what a visitor scrolled past, stopped on and opened, then reordered what came next. Split-tested against the static page, profit per session went up 28%. The caveats are his own — one test, one store, no session count or window published, and the man reporting the number owns the company. What earns it the space is not the 28% but the design: the only claim here with a control group and a metric fixed before the test, on a surface nobody was optimising. The order of products on a collection page has not changed on most stores since the theme went in.

One test isn't a benchmark. It might do nothing on your store.

Matthew Bertulli, co-founder, Pela — LinkedIn, 22 Aug 2026

The other control group belongs to a study nobody named: a randomised field experiment on 35,625 users varied how strongly an AI summary of customer reviews presented them as showing consensus, and product clicks rose 24% (Agnes Kaczmarek, X, 14 Sep 2026). Best design in the chapter, weakest attribution. Clicks are not orders, and the dial that moved was how confidently a machine summarises other people's opinions.

A/B testing with a model in the loop

Shopify's Winter '26 Editions put two relevant tools in the admin: Rollouts, native A/B testing with scheduled theme changes, and SimGym, which sends simulated AI shoppers through a storefront before real customers see it. Both are shipped, not proven. No merchant in this corpus has published a result from either, and a simulation is a prior, not an outcome.

Be careful with simulated shoppers. One post summarises research finding that LLMs predict purchase intent at around 90% accuracy when asked to roleplay a customer, and "fail completely" when asked to rate a product one to five, because models default to safe distributions (CrazyShyyt, X, 8 Sep 2026). The method decides the answer, and buying the tool means inheriting whichever method it chose. Amazon's 70–90% is the strongest version of the claim and still only triage: it predicts which arm wins, not by how much, and one call in five is wrong. Point it at the bottleneck below — which of fifty variants gets the month's traffic — not at the test itself.

Handle with care

If you let a model call the winner, fix the metric first. Revenue per session is the seductive choice and the dangerous one: it hides refunds, returns, discount depth and support load, so a variant that lifts revenue while lifting returns wins on the model's number and loses on yours. AI removed the variant-supply bottleneck without touching the traffic bottleneck — fifty landers a week, still only enough sessions to settle one comparison a month. Let one of those be a removal: a Shopify brand lifted revenue more than 10% by deleting its checkout upsell (Anya Geimanson, X, 8 Sep 2026 — brand unnamed, so take the direction, not the number).

Checkout: the hard numbers still point down

The closest thing to a measured AI conversion result belongs to a channel rather than a page, and it is negative: after testing roughly 200,000 items, Walmart found purchases completed inside ChatGPT converted about three times worse than click-outs to its own site (Gagan Ghotra, X, 20 Mar 2026). Chapter 03 owns that story. The point here is that it took a retailer of Walmart's size to produce it.

September added a second, from the other end of the pipe. HUMAN Security's traffic data puts 87% of AI-agent visits to retail sites on product pages and barely 2% at a checkout (Sreekant Vijayakumar, LinkedIn, 8 Sep 2026). Beside the Shopify figure Chapter 02 owns — about half of AI-referred sessions land straight on a product page, against roughly 20% for organic — the shape holds: AI traffic arrives deep in the funnel and does not close. Two populations with opposite intent land on the same PDP, read through one conversion rate. Split PDP conversion by referral source before buying anything to fix it (Bharti Prasad, LinkedIn, 25 Aug 2026).

Against that, one vendor case study: Ava's AI phone agents called shoppers who had abandoned checkout at pet&home.

Vendor case study
Figure 34Vendor case study: Ava's AI phone agents called shoppers who abandoned checkout at pet&home, claiming $4,157.70 recovered in 14 days and 41x ROAS.Source: Ava (avasales.ai) · via Ava AI, X 2026-09-11

Handle with care

$4,157.70 in 14 days at $41 per $1 spent is a vendor's own first case study with no holdout. Some of those shoppers would have come back without a phone call, and nothing here separates them. A holdout of abandoners who get no call, the same window for both arms, and net revenue after refunds would make it usable. Cite it as a demo.

Measure it yourself

One split test is not a benchmark, so the only honest number is still your own. Here is the cheapest version that produces evidence, inside a fortnight.

  1. One surface, one change

    One PDP block, one checkout promise, one collection order. Record what the AI changed before it ships.

  2. Declare the metric before the test starts

    Contribution margin per session, not revenue per session, plus guardrails you agree to lose on: refund rate, return rate, support tickets per hundred orders, discount depth.

  3. Hold out 10%

    Rollouts gives Shopify merchants native A/B without a third-party tool. If a vendor is running the change, the holdout has to be yours, not theirs.

  4. Run the delivery-date test first

    It is the one idea here with survey weight behind it. Put a real date above the fold against your current discount field and let the margin number decide.

  5. Publish the loser

    Internally is enough. Four of this chapter's ten rows are vendor material because nobody writes up the tests that failed.

Operator note

Treat AI in the funnel as a cost-side win — pages and audits that took a week now take an afternoon — and the conversion-side win as proven once, in one store, by the person who built it. Copy Pela's surface rather than Pela's number: a collection page that reacts to what a visitor just did, judged on profit per session against the frozen version. Do not underwrite a forecast with numbers this chapter lacks. And read your product page the way the model does as well as the way a buyer does: as Naseem Haider put it, it has two jobs, to persuade the human and remove ambiguity for the machine.

Chapter 07

Ads moved inside the answer before anyone measured what they return

Buying space inside an AI answer went from a $200,000 minimum to a self-serve auction in six months, and the share of answers carrying a paid placement is climbing. What a merchant gets back is still documented by one anonymous advertiser and two small agency tests.

What holds up

  • ChatGPT Ads reached a $1B annualised run rate in under 200 days, per OpenAI (late Aug 2026) — a run rate, not recognised revenue. eMarketer forecasts under $1B of actual chatbot ad revenue for calendar 2026. OpenAI, via Clayton Wood · eMarketer, via Glenn Gabe
  • Sponsored placements filled 24% of ChatGPT hotel answers in May 2026, up from 6% in March (Comscore) — travel queries; no retail equivalent is public. post
  • Amazon's first conversion number for ads inside a shopping assistant: Sponsored Prompt clickers convert 48% more often and spend 21% more (Q2 2026 earnings). Clickers versus non-clickers — a selection gap, not a measured lift. post
  • 62% of ad buyers use generative AI for video creative, up from 51% (IAB, July 2026) — but across 578,750 creatives Motion still finds ~5% of creatives winning and absorbing 55% of spend. post
  • Meta flipped 10 Advantage+ Creative modifications to default-on in July 2026 — backgrounds, music, text overlays, animation — with no opt-in, per an agency founder who found them running on client accounts. A creative test in that window ranked ads the merchant never uploaded. post
  • No high-reliability ROAS number exists for a merchant running ChatGPT Ads: the public proof point is one anonymous ecommerce advertiser at 3x over 28 days, no methodology. post

Two separate things happened to ads in 2026 and operators keep discussing them as one. Paid placements began appearing inside AI answers — the surface merchants were just told to earn their way into. Separately, making an ad got cheap enough that volume stopped being a budget question. The first has revenue behind it and almost no merchant-side evidence; the second has evidence, most of it written by people selling the tool.

The inventory arrived before the evidence

In late August 2026 OpenAI said ChatGPT Ads had passed a $1 billion annualised revenue run rate in under 200 days. Read that as written: a run rate annualises a recent period, it is not money recognised for a year. eMarketer's forecast has all chatbots together — ChatGPT and Google AI Mode included — generating under $1B in 2026 ad revenue, against OpenAI's own projection of $2.5B. Both can be true; only one describes invoices.

A Sponsored Saatva listing sits inside a Google AI Mode shopping answer, styled like organic picks and next to 33 cited review sources
Figure 35A Sponsored Saatva listing sits inside a Google AI Mode shopping answer, styled like organic picks and next to 33 cited review sources.Source: Google · via Barry Schwartz, X 2026-05-26

The revenue line matters less to a merchant than the shape of the placement. In Google AI Mode the sponsored listing sits inside the answer, in the model's own visual language, beside the citation panel a brand spent the quarter trying to enter. There is a Sponsored label and no ad block to scroll past: the unit competes for the same attention as an organic recommendation because it is rendered as one. Not every surface renders it that way. Through August 2026 ChatGPT had one placement, directly below the generated response, labelled sponsored and laid out like a search result — adjacent, not woven in.

Ads are filling AI answers: sponsored placements in ChatGPT hotel results

Share of answers carrying a sponsored placement. Travel queries; retail ad load was not measured.

0%6.25%12.5%18.75%25%Mar 2026Apr 2026May 202624%
Named but secondhandComscore, spring 2026. Context: OpenAI reports a $1B annualised ads run rate in under 200 days.
Show data table
Answers with a sponsored placement
Mar 20266%
Apr 202614%
May 202624%

Comscore's hotel series is the only public measurement of ad load inside AI answers. It is travel, not retail, and three points is a short series — take the trajectory as the finding, not the level. A surface that launched ad-free was carrying a sponsored placement in roughly one answer in four within a quarter.

Buying it is now a Tuesday afternoon

Access moved faster than proof. In February 2026 ChatGPT Ads meant a $60 CPM and a $200,000 minimum; by May it was self-serve CPC bidding with no minimum. On 31 August 2026 the Ads Manager opened for direct buying in 31 European markets, and in India on 4 September. Amazon Ads now has a pilot letting selected US advertisers buy ChatGPT inventory through Amazon DSP — a retail-media team gets in without ever talking to OpenAI.

How ChatGPT Ads access changed in six months
Figure 36How ChatGPT Ads access changed in six months: $60 CPM with a $200K minimum in February, self-serve CPC with no minimum by May, wider geographic rollout by August.Source: Clayton Wood · via Clayton W., LinkedIn 2026-09-03

The console is familiar on purpose — campaign, ad group, ad; conversions, feeds, audiences. A team running Google Shopping can operate this second feed-driven auction in an afternoon, which is the trap: a mature console implies a mature channel.

ChatGPT Ads opened self-serve in 31 European markets on 31 August; Bidnamic flags that the pixel defaults consent to true and that the 3x ROAS claim lacks methodology
Figure 37ChatGPT Ads opened self-serve in 31 European markets on 31 August; Bidnamic flags that the pixel defaults consent to true and that the 3x ROAS claim lacks methodology.Source: Bidnamic · via Ingvar Kraatz, LinkedIn 2026-09-04

It's unproven at scale. OpenAI's only public proof point is one anonymous e-commerce advertiser at 3x ROAS over 28 days. Vendor claim, no methodology.

Ingvar Kraatz, Co-Founder / COO at Bidnamic — LinkedIn, 4 Sep 2026

Bidnamic's second warning costs more. By OpenAI's own documentation the conversion pixel initialises consent to true unless told otherwise, and hashes customer data detected in site forms. Installed without a consent platform behind it, that is GDPR/PECR exposure on a test budget.

What it returns, honestly

Start with what buyers are told to expect. The most specific public set was published in August 2026 by paid specialist Nina Welt: CPM around $25–45 in scaled US segments, down from the early $60; a CPC benchmark near $1.72 against OpenAI's advice to open bids at $3–5; CTR around 0.7% blended, some campaigns at 1–2.5%; conversion rates from 0.2% to 5.8%. Collected expectations — Vendor or unsourced — not a measurement: no account count, no period, no method. The structural read is the durable part — at that CTR the unit performs like display while the user context reads like search. Price it as search and you switch it off inside a month.

The merchant-side evidence is two small tests pointing opposite ways. Windmill Strategy ran ChatGPT Ads for 30 days against industrial and manufacturing buyers: CPCs lower than expected, engagement solid, zero conversions. Planit, across three anonymised brands, reported organic ChatGPT traffic up 42% while ads were live (+10%, +41%, +164% by brand) — before-and-after, no control group, measured while AI referral traffic was growing everywhere. Both still beat the vendor 3x, because both name their method.

+48%

Shoppers who click a Sponsored Prompt in Amazon's AI shopping assistant convert more often, and spend 21% more, Q2 2026. Clickers vs non-clickers.

First-party / research Amazon Q2 2026 earnings call · post

5% / 55%

Share of ad creatives that win, and the share of spend they absorb, across 578,750 creatives and $1.29B of spend.

Named but secondhand Motion · post

Amazon's are the strongest numbers here and still need care. They came on the Q2 2026 call, past the $19.8B advertising headline. The 48% describes the shopper as much as the placement: whoever asks an assistant a question and clicks the answer is already deep in a decision. It proves intent and volume are real on that surface, not that the ad created the purchase. It is also the placement a seller controls least. Sponsored Products Prompts left beta on 25 March 2026 as a billable CPC unit, only automatic and broad-match campaigns qualify, and a seller can opt out of specific prompts but cannot write one, bid on one or give it a budget; exposure is read after the fact in the Prompts report. No holdout, no line item to pause — so the merchant cannot turn Amazon's 48% into a number of their own.

The ad load lands on the visibility you were just told to chase

Google AI Mode on mobile inserts sponsored merchant price offers (Back Market, eBay) directly beneath an organic MacBook Air recommendation
Figure 38Google AI Mode on mobile inserts sponsored merchant price offers (Back Market, eBay) directly beneath an organic MacBook Air recommendation.Source: Google · via Barry Schwartz, X 2026-05-26

On mobile the sponsored merchant offers sit directly beneath the organic recommendation — same answer, same scroll, paid row under the earned one. Every point of ad load is answer real estate no feed, schema or citation strategy wins back. After AI Overviews, Smarter Ecommerce measured median Google Shopping impressions falling from about 1.85M to about 1.4M while median CTR rose from about 1.20% to about 1.55%: fewer impressions, better qualified. Vendor analysis over ~175B impressions, confounded by everything else Google shipped that year.

Amazon is charging my company money to answer customer questions and then directing our customers to inferior competitor products who have paid Amazon. We never signed up for Rufus ads.

Molson Hart, Founder and CEO of Viahart — X, 21 Apr 2026

Creative: the cost collapsed, the hit rate did not

Adoption is measured here, not claimed. IAB put 62% of ad buyers using generative AI for video creative in July 2026, up from 51%. XR Extreme Reach found 26% of marketers using digital replicas and 22% synthetic AI talent. ADWEEK reported research putting 88% of advertisers on AI in creative production; that study is unnamed.

The economics of testing did not move. Motion's analysis of 578,750 creatives and $1.29B of spend found roughly 5% of creatives are winners absorbing 55% of spend, and drew the useful conclusion: low hit rates are a statistical feature, not a quality signal. Cheaper production multiplies the denominator without improving the ratio. It moves the constraint from making an ad to judging one.

A Brazilian TikTok Shop account where one presenter, reportedly AI-generated, sells many tools; top videos reach 1
Figure 39A Brazilian TikTok Shop account where one presenter, reportedly AI-generated, sells many tools; top videos reach 1.3M–4.1M views.Source: TikTok Shop · via rewind, X 2026-09-12

Volume does work somewhere. A Brazilian TikTok Shop account runs what the poster reads as one AI-generated presenter across dozens of products, top videos at 1.3M–4.1M views; the AI claim is his, unverified. That is media arbitrage with no brand to damage, and it now carries disclosure duties: New York's GBL §396-b took effect in June 2026, and Amazon began requiring sellers to tag AI-generated people in product media the month after.

Two findings cut across the volume play from opposite sides. IAB research reported in September 2026 found 56% of recent AI-shopping users preferred recommendations carrying creator perspectives, and 65% felt more confident when credible creator reviews informed the answersecondhand, and stated preference rather than observed behaviour. Against it, a small blind study by Kristie Marsela, matching AI iterations against the source photographs of one D2C brand, put detection accuracy at 51.08%, with the AI images scoring higher on perceived quality; sample size not stated. The scarce input is the perspective, not the pixels.

The platform is editing the thing you are testing

The sharpest change in creative testing this year was not a generator. In July 2026 Meta flipped ten Advantage+ Creative modifications to default-on — backgrounds, image templates, music, 3D animation, text overlays among them — with no opt-in and no notification, as reported by agency founder Kody Nordquist after finding them running on client accounts. He reports it; Meta published no number for what the edits do. The consequence is arithmetic: if the platform can add music, swap a background or layer text after upload, the ad that won is not the file you uploaded, and what you concluded about hook or format belongs to a variant you never saw.

A batch of 30 videos that each differ in 6 ways isn't a test. It's a lottery with a spreadsheet attached. You get a winner and no idea why it won.

Michael Godzic — X, 1 Aug 2026

The same runs on the buying side. Google's AI Max widens matching and rewrites assets inside an existing Search campaign; agency Mabo ran it across 40 live client accounts for seven weeks and reported 9% of spend generating 14% of conversion value, account ROAS up 43% over the window — its own clients, no control group, everything else moving too. Google sells the measurement as well: Data Manager, now wired into Analytics and DV360, arrives claiming advertisers who connect offline and app data see an average 26% rise in incremental ROAS Vendor or unsourced. Google's own number, selection built in — advertisers able to connect offline sales are the ones who have offline sales worth connecting.

What survives this is testing discipline, not a bigger generator. One vendor framework at least counts the right thing: Creative Testing Velocity measures distinct ideas tested per unit of spend rather than renders or cost per video. The idea is free: if you cannot say how many separate ideas last month's spend bought, the answer is smaller than the file count.

Handle with care

The loudest AI-creative posts here are sales material, rated Vendor or unsourced in the ledger. A finished AI UGC ad "in 9 minutes for $0.68", 550 ads a day, CTV spots from $1M+ down to $5,000 — each posted by someone selling the workflow, none reporting what the rest of the batch did. The counterweight is in the same corpus: an analyst reviewing 120 DTC ads judged 90% AI slop, and an agency shipping 60 ads a month for one brand starts each with buyer research, not a prompt.

Operator note

Treat ChatGPT Ads as a fixed-budget awareness test, not a ROAS channel, until you have your own number: capped spend, a holdout region or a paused-and-resumed window, branded search and direct traffic read alongside last-click. Wire the pixel into your consent platform before it touches an EU or UK site. Before the next creative test, open Advantage+ Creative on every active campaign, switch off the enhancements you did not approve and screenshot the settings. On Amazon, read the Prompts report monthly — it is the only view you get of a placement you cannot budget.

Chapter 08

Flows already earn their keep. What AI adds on top is mostly unmeasured

Automated flows produce 41% of email revenue from 5.3% of sends, so AI is being aimed at the most valuable surface a merchant owns. Almost everything built on top of it is reported without a control group — and the one exception is a Klaviyo feature, still in preview, that shipped holdout testing alongside the thing it is meant to measure.

What holds up

  • Klaviyo's 2026 benchmarks, across more than 183,000 brands: automated flows generate 41% of email revenue from 5.3% of sends, at 18x the revenue per recipient of campaigns (Klaviyo 2026 Email Benchmarks, cited by James Tyler).
  • The headless move is real and dated: 260+ MCP tools and 490+ APIs, announced at K:BOS, Aug–Sep 2026. Salesforce got there earlier, adding 60+ MCP tools on 19 August 2026 (Brandon Caples, LinkedIn).
  • Klaviyo's new personalization layer is in preview, with holdout testing built in — the correct instrument arriving with the product instead of years after it (Einar Thor, K:BOS day one, 10 Sep 2026).
  • Only two AI-retention results in this corpus were measured against a withheld control: +70% browse revenue and 9x incremental ROI over a 30-day 50/50 test at Fellow, and +$19,700 incremental over 30 days on one live account. Both are vendor or agency self-reports, and neither has been replicated.
  • The sweep aimed at this chapter came back nearly empty: a fourth screening round targeting AI in email, SMS and lifecycle yielded 15 usable posts, against 47 for ads and creative and 50 for imagery — the thinnest of any funnel area (this report's own screening, Sep 2026). The tooling ships far ahead of anyone publishing a measured result.

The surface is worth the attention

Retention is the one chapter in this report where the base rate is well documented before AI shows up. Klaviyo's 2026 email benchmarks, drawn from more than 183,000 brands, put automated flows at 41% of email revenue off 5.3% of sends — 18 times the revenue per recipient of a broadcast campaign. The ratio matters more than the 41%. Flows win because they fire on something the shopper just did, which is why they are the first place to spend a personalization budget: the context is already there, and the volume is small enough that better messages are cheap to make.

Klaviyo: automated flows earn 41% of email revenue from 5.3% of sends

Flows are triggered by behaviour, so they are the surface where AI personalization pays first.

Share of sendsShare of email revenue
0%12.5%25%37.5%50%Automated flows5.3%41%
First-party / researchKlaviyo 2026 Email Benchmarks, 183,000+ brands.
Show data table
Share of sendsShare of email revenue
Automated flows5.3%41%

Read the chart as an allocation argument, not a performance one: a brand that puts almost all of its sending effort into campaigns is spending most of its attention on the low-yield side of the ledger. Every AI retention pitch in this corpus is, underneath, a claim that the same machinery can now be pointed at the high-yield side without adding headcount.

18x

Revenue per recipient from automated flows versus broadcast campaigns, across 183,000+ Klaviyo brands, 2026 benchmarks.

First-party / research Klaviyo 2026 Email Benchmarks · post

490+

APIs Klaviyo exposed to agents at K:BOS, alongside 260+ MCP tools, Aug–Sep 2026.

First-party / research Klaviyo K:BOS 2026 · post

Headless, and who got there first

At K:BOS in September 2026, Klaviyo opened the CRM to agents: 260+ MCP tools and 490+ APIs against live profile data, so a marketer or an agent can read, write and send without opening Klaviyo. In a same-week interview the CEO gave the figure as 400+ APIs — a rounded lower bound rather than a contradiction, but the kind of drift worth pinning to the announcement and its date.

An agent can read, write, and go live on its own. No need to open Klaviyo at all.

Christopher Di Censo, Klaviyo — LinkedIn, 10 Sep 2026
Illustration of Claude Code calling Klaviyo's MCP tools (create template, create campaign, assign template) to generate an on-brand, editable email template
Figure 40Illustration of Claude Code calling Klaviyo's MCP tools (create template, create campaign, assign template) to generate an on-brand, editable email template.Source: Klaviyo MCP / StoreHero · via 坂本@StoreHero, X 2026-08-18

That image illustrates the workflow, not a screen recording — treat it as a diagram. What is verifiable is the tool surface and its date, and the surface is no longer a differentiator: Salesforce shipped Headless 360 in April 2026 and added 60+ more MCP tools on 19 August. Driving your ESP from a chat window is table stakes, so the question for a vendor is not whether they have an MCP server but what their agent may do without a human.

Who approves the send

StoreCRM beta
Figure 41StoreCRM beta: a plain-language Sidekick request becomes a branching post-signup scenario (wait 1 or 2 days by spend), saved as an inactive draft for human review.Source: Shopify Sidekick / StoreCRM · via StoreCRM@LINEもメルマガも送れるShopifyアプリ, X 2026-09-11

The StoreCRM beta above is a real capture, and the design choice worth copying is the last step. A merchant describes a branching post-signup scenario in plain language — wait one day or two depending on how much the customer spent — and the system produces a draft, saved inactive, for a person to check before anything sends — the cheapest available guard against an agent with write access to a hundred thousand people.

Proposed AI-native retention team
Figure 42Proposed AI-native retention team: humans keep strategy, creative and send approval; six specialist agents under an orchestrator handle analytics, segmentation, copy, flows, planning and reporting.Source: Grok Bots · via David Dokes, LinkedIn 2026-08-18

The org chart above is a proposal, not a deployment: one operator's seven-agent retention department, built to see what was possible. The annotation is the finding — humans keep strategy, creative direction and approval of every send, and the agents are grounded in warehouse data so they cannot invent metrics. Nobody here is running an unattended lifecycle programme, and nobody claims to be.

What a holdout changes

The most useful sentence of the K:BOS week was not about tools.

Personalization is getting a new layer, in preview, to coordinate audiences, content, product recommendations, timing and channels. With holdout testing to measure whether it makes a difference.

Einar Thor, Pipar\TBWA — LinkedIn, 10 Sep 2026

The layer is narrower than its name. Part of it is Klaviyo's Marketing Analytics renamed; what is new is a separate predictive model per decision — who to send to, what to show, when, on which channel — written back onto the customer profile, so a marketer can read what the system believes about a person. Predicted repurchase timing becomes a flow trigger. Some of it, the new segment layers included, was not live on announcement day (Jordan O'Connor's K:BOS breakdown, 12 Sep 2026).

A holdout is a randomly withheld slice of the audience that never gets the treatment, so the comparison runs against people who would otherwise have been targeted the same way. It is the difference between attributed and incremental revenue, and the measurement almost every AI marketing claim in this report is missing. Across eleven chapters this is the only vendor building the correct instrument into the feature instead of leaving the customer to construct it — and it is a preview feature, not a proven result. Nothing has been published about what its holdouts show.

Two operator results in the corpus used that design. At Fellow, a 30-day 50/50 test of AI-personalised messages inside the brand's existing Klaviyo flows reported 70% more browse revenue and 9x incremental ROI. On another account, a 30-day trial of an AI inbox-placement and subject-line tool reported $19,700 incremental.

The lift is measured against a holdout group, so this isn't platform attribution. It's a real incremental number.

Justin Guimera, retention agency — LinkedIn, 16 Mar 2026

Both are published by the party selling the tool: the design is right and the sample is not, because no one publishes the 30-day test that came back flat. Take them as existence proofs — user-level AI personalization can clear a holdout on some accounts — not as planning numbers.

Multi-agent loop for email optimisation
Figure 43Multi-agent loop for email optimisation: critic, rival author and synthesizer produce variants, three blind judges pick a winner weighted by revenue per send, then a 5–10% holdout test.Source: @shannholmberg (autoreason) · via Shann³, X 2026-04-25

The loop above is the pattern worth stealing from the vendor pitches: generate variants, judge them blind on revenue per send rather than on taste, then put the winner in front of a 5–10% holdout before it becomes the default. The agent count is decoration; the holdout at the end makes the rest of it legible.

What gets claimed without one

A fourth scraping round was aimed at exactly this chapter and came back with 15 qualifying posts, fewer than any other area of the funnel. None was measured against a control. The imbalance is the finding: Klaviyo shipped agents, an MCP surface and a predictive layer inside one quarter, and almost everyone reporting what that did to revenue is selling it.

The widest number in the corpus shows the shape. Win-back emails across ecommerce are put at about $0.11 per send; win-back flows where AI continuously updates products and messaging are put at $5 to $10 per email. That is a 45- to 90-fold gap with no period, no account count and no baseline for the treated group, and it comes from the co-founder of the company selling the optimisation, relayed on LinkedIn (Pietro Saccomani, 4 May 2026). Keep it as the ceiling of what a vendor will say out loud, and for nothing else.

The cart-recovery claims cluster the same way: +73% abandoned-cart revenue in 30 days at an unnamed medical supply brand (CustomersAI, 3 Jun 2026), 20% cart recovery at one wine merchant on WhatsApp, $4,157.70 in fourteen days from a phone agent. In each, a new channel arrived at the same moment the AI did, and nobody separated the two.

ClaimValueHow it was measuredReliability
Fellow, AI-personalised Klaviyo lifecycle emails+70% browse revenue, 9x ROI30-day 50/50 holdoutNamed but secondhand
AI inbox placement and subject lines, one account+$19,70030-day holdoutNamed but secondhand
"What stopped you?" agent reply instead of a discount~+10%Method not disclosedNamed but secondhand
Predicted-CLV segmentation, Klaviyo accounts+15–25% RPRCorrelational: accounts that use it vs accounts that do notNamed but secondhand
Founder AI voice-note SMS, subscription brand+53% second-order ratePractitioner report, no method statedVendor or unsourced
Win-back flows continuously updated by AI$5–10 per email vs $0.11Vendor CEO figure, no period or account countVendor or unsourced
AI phone agent on abandoned checkouts, one store, 14 days$41 per $1Vendor case study, no controlVendor or unsourced

Handle with care

The predicted-CLV number is the one most likely to be repeated wrongly. Fewer than 20% of Klaviyo accounts use predicted CLV for segmentation, and those that do see 15–25% higher revenue per recipient — but that compares accounts that chose to configure it with accounts that did not. Teams sophisticated enough to run predictive segmentation are doing ten other things right. The gap is real; the causal claim inside it is untested. The same caution applies to every "users of X outperform non-users of X" number in this report.

Two audits in this sweep put the binding constraint further upstream. One mid-size store — Shopify Plus, Klaviyo, live automation flows — still had only 36.7% of its database active and sent everyone the same message at the same time; the author's ordering is segmentation first, predictive marketing last (Olimpia Valentina Chiriacescu, 6 Jul 2026). An agency that audits Klaviyo accounts puts the typical one at roughly 30% of what the platform already does: three to five flows, a weekly campaign, agentic optimisation never switched on. Both are single-operator reports, and both point where the predicted-CLV gap points.

Output of an AI Klaviyo flow audit
Figure 44Output of an AI Klaviyo flow audit: it flags underperforming or missing flows (welcome, cart, browse abandon, winback) and ranks fixes by estimated revenue. Figures are illustrative.Source: not shown · via Kinga Dow I AI Klaviyo Email Marketing, LinkedIn 2026-04-05

Production speed is where AI in retention is least disputed and least interesting. Audit outputs like the one above are routine now: 90 days of flow data in, a ranked list of gaps out. The dollar figures on that screen are a model's estimate of opportunity from an unknown account, not booked revenue, and they circulate as if they were. Claims that a flow now takes 5–10 minutes instead of 2–3 hours are probably true, and they are a cost story: one practitioner reports agencies re-quoting retention retainers from $4,100 to under $3,000 a month. If AI only makes the same programme cheaper to run, the saving lands with the client.

Speed is also where the sharpest dissent sits. An email agency spent a year testing the category — copy generators, design assistants, segmentation models, send-time optimisation — and published a sentence no vendor deck contains.

AI is incredibly good at making you faster. It's terrible at making you interesting.

eCom Email Marketer — LinkedIn, 15 Apr 2026

Ten subject line variations in seconds, they report, all ten sounding like one template. A practitioner's read, not a measurement — but from someone with no tool to sell, and it matches the only controlled comparison the round turned up. Digital Applied paired 100,000 cold emails, half AI-written and half human, October 2025 to April 2026: the variable that moved inbox placement was not authorship. Three days between sends landed 93%, one day landed 71%, copy unchanged (relayed by William Doom, 9 Sep 2026). Cold outbound is not permission email, so leave the digits there. The shape carries: the writing got faster and the schedule around it mattered more.

What none of this fixes

One consultant estimate in the corpus is worth keeping precisely because it is unproven: roughly 7 in 10 first-time ecommerce buyers do not return within a year, and Zalando's order frequency has stayed near 5 orders per customer per year for six years, with SHEIN near 4. That is a practitioner's framing, not research, and the framing survives even if the digits do not: none of the AI work above gives a customer a reason for a second purchase. It makes the second ask faster, cheaper, better timed and better worded. That is worth money. It is not the same problem.

Operator note

Do the boring version first, in this order. Check the split of sends and revenue between flows and campaigns in your own account against 5.3% / 41%; if your flows are below that share of revenue, the fix is flow coverage, not AI. Then, before you buy any AI personalization tool, ask the vendor to run it against a 10% holdout for 30 days and report incremental revenue, not attributed revenue; a tool that cannot hold out cannot tell you whether you bought anything. Keep every agent's output as an inactive draft with a named approver until you have three clean cycles.

Chapter 09

Support and operations: the spend that has to prove itself

Gartner scored 432 customer-service AI use cases: one in four returns positive ROI, another 42% cannot be scored at all. The ones that pay share a shape: wired into the systems behind the chat window, metered per resolved case, stopped at a human before anything irreversible.

What holds up

  • Only one in four of 432 customer-service AI use cases produces positive ROI; 42% are unclear — Gartner, 2026 (post).
  • The volume is real: ASOS says AI agents handle about half of inbound care requests — but never defines "handled" (post).
  • Deflection is no longer free: Gorgias adds $0.90–$1.00 per AI Agent resolution on top of the ticket fee (post).
  • Inventory, fraud and forecasting carry the cleanest numbers here, nearly all published by the vendor selling the agent.
  • Where it works, the agent investigates and drafts; a human approves the action.

The ROI number, and what it measures

Gartner's analysis of 432 customer-service AI use cases found 25% with positive ROI and 42% unclear. Companies run nearly five use cases on average, commit 13% of the service budget to AI, and 42% cannot say what value those systems create. Yet 56% of service leaders expect their incentives tied to AI outcomes this year — pay attached to a number four in ten cannot produce.

Customer-service AI: only a quarter of use cases show positive ROI

Of 432 use cases. The remaining share is not reported, so it is left out rather than guessed.

0%12.5%25%37.5%50%Unclear ROI42%Positive ROI25%
First-party / researchGartner analysis of 432 customer-service AI use cases, 2026.
Show data table
Unclear ROI42%
Positive ROI25%

The rest of the 432 was not reported, so it is not drawn. Most customer-service AI is unscored, not proven bad.

The gap often starts before deployment: the goal is simply "add ai" · use cases are chosen by volume · containment becomes the headline metric · measurement ends at deflection.

Alexander Abstreiter, AI vendor — LinkedIn, 4 Sep 2026

He sells this software and says so in the same post. The diagnosis still holds.

"Handled" is not "resolved"

ASOS's playbook is the largest production claim in the corpus: about 50% of inbound care requests handled by AI agents, plus 90% company-wide Copilot adoption, reached in phases rather than one launch. It also leaves the load-bearing word undefined. A ticket an agent touched, one it deflected and one the customer thought was solved are three different denominators.

Below it sit vendor cases. Shopify's agentic use-case page reports Klaviyo's Customer Agent resolving 75% of LifeStraw's inquiries without a human, agent-recommended sales up 111% in 90 days — no pre-AI baseline, no scope of inquiry types. Gorgias ambassador Max Wallace reports 20% of conversations automated in the first weeks at Tommy John, disclosing the relationship in his first line. Then the round numbers passed around as common knowledge: ~30% of tickets are "where is my order", AI deflects ~80% of those, every vendor claims ~60% resolution.

CommerceAgentBench
Figure 45CommerceAgentBench: 13 models on 107 real commerce workflows; the best (Claude Opus 5) completes 61.7%. Even top agents fail about four in ten operational tasks.Source: Accio Research (commerce-agent-bench.site.accio.ai) · via Dipti Sharma, X 2026-08-26

The benchmark above is the corrective. Across 107 real commerce workflows the best of 13 models completed 61.7%, on a leaderboard Accio runs itself, with domain task counts summing to 117 against a stated 107. Directional, not exact, and still: top agents fail about four operational tasks in ten. A resolution rate above 70% is a claim about a ticket mix, not capability.

The anecdotes circle the same hole. An agency reports a store's support email halving, 100 to 50 a day. An operator watched cost per interaction drop from $6 to $0.50, then read the reviews, which were furious. A vendor says its "dumber" agent, which hands off below 70% confidence, drew zero complaints in 14 days against three for the smarter one. All tiny, all self-reported, all about what most rollouts never instrument: the customer's outcome after deflection.

Handle with care

"30% of tickets are WISMO", "AI deflects 80% of WISMO", "~60% resolution": unattributed in every post carrying them (example), and your helpdesk exports the real share in a morning. Gartner's 13% of service budget is an average; the much-shared "median 12%" is a different statistic from a different period.

25%

Of 432 customer-service AI use cases, the share with positive ROI; another 42% unclear, 2026.

First-party / research Gartner · post

~50%

ASOS inbound care requests handled by AI agents in production, 2026. "Handled" undefined.

Named but secondhand ASOS CTO Przemek Czarnecki · post

$0.90–$1.00

Gorgias AI Agent price per resolution, on top of the billable ticket, 2026.

Named but secondhand Gorgias pricing page · post

Per-resolution pricing changes the arithmetic

Gorgias bills per billable ticket — $10 a month for 50, up to $900 for 5,000 — and the AI Agent adds $0.90 to $1.00 per resolution on top, so a chat the AI closes is billed twice. One missing number now decides it: your fully loaded cost per resolved case today. A thousand AI resolutions a month is about $1,000 in resolution fees plus ticket fees, against the staffed hours removed. The meter is spreading: in September 2026 Salesforce closed its acquisition of Fin, ex-Intercom, 30,000+ customers.

Operator note

Log four numbers for 30 days before switch-on, on the ticket types the agent will take: resolution rate, escalation rate, loaded cost per resolved case, refund or CSAT outcome. Without the "before", the vendor dashboard is your only evidence, and it reports deflection.

The operational pattern: the agent drafts, a human approves

The clearest working deployments here are not chatbots. They are agents that stop at an approval gate.

Gorgias Cortex in Slack audits ~100 human-handled tickets, finds a missing retention step, drafts a Skill update with evidence and risk tier, and waits for approval
Figure 46Gorgias Cortex in Slack audits ~100 human-handled tickets, finds a missing retention step, drafts a Skill update with evidence and risk tier, and waits for approval.Source: Gorgias (Cortex agent) · via Romain Lapeyre, X 2026-09-07
Gorgias's Cortex agent in Slack audits a merchant's pending automation rules and help-center drafts, recommending which to enable, but changes nothing without human approval
Figure 47Gorgias's Cortex agent in Slack audits a merchant's pending automation rules and help-center drafts, recommending which to enable, but changes nothing without human approval.Source: Gorgias · via Romain Lapeyre, X 2026-09-09

Both captures show one loop in Gorgias's Slack agent, Cortex: it audits roughly a hundred human-handled tickets, finds a retention step the AI had been skipping, drafts the Skill update with evidence and a risk tier, then waits. In the second it reviews pending rules and help-centre drafts and recommends which to enable, changing nothing itself. Polar's Operator sells the same contract, "it acts when you approve", and its $3k/day of Meta spend on a broken URL is a self-report of what such a loop finds.

A practitioner's AI support workflow
Figure 48A practitioner's AI support workflow: rules route tickets, AI drafts replies, code checks promises, a second AI reviews daily, and any refund needs human approval.Source: none shown · via David, X 2026-09-07

Above is the small-merchant version, posted free by a store operator: rules route the ticket, AI drafts the reply, code verifies the promises in it, a second AI reviews the day's answers, and any refund needs a person. His 0.45% dispute rate over 9,000+ tickets is in the post, not the screenshot. Copy the architecture, not the numbers.

Approval gates scale only while the action is rare and expensive. Albert Heijn's AI produces over a billion automated demand forecasts a day, 17,000 products across 1,200 stores, 50 days out, and nobody approves a billion forecasts. There a human approves the purchase order, the exception, the policy. Amazon is cited secondhand as reporting 35% fewer stockouts from AI inventory tools; the podcast citing it asks the right question, which is when a human should approve the order.

Returns, inventory and fraud: the cleanest numbers, and who published them

Shopify's agentic use-case page carries the tidiest outcomes here: Signifyd's fraud agents taking Cymbiotika to 98% order approval with 93% fewer chargebacks, Prediko saving Kuppa Joy 10 hours a week with 90% fewer order errors. Vendor-selected cases on the platform's marketing page, no baseline stated. SAP's Salling Group figure, up to 66% less time from supplier to shelf, is labelled a projection by SAP itself — never quote it as a result.

Returns suit the approval pattern best and are measured least. AfterShip Returns became one of the first Sidekick app extensions, so a merchant can ask "show me returns that need my approval" inside the Shopify admin. SHIPPED, not proven: no time-saved or resolution figure exists for it, and Dr. Squatch's 25% drop in WISMO from predicted delivery dates is unattributed.

A consumer agent (Muse) runs an Amazon support chat in the browser and gets a promised $70 refund issued — merchant support teams will increasingly face agents, not people
Figure 49A consumer agent (Muse) runs an Amazon support chat in the browser and gets a promised $70 refund issued — merchant support teams will increasingly face agents, not people.Source: Muse; Amazon · via PREEM//The Unofficial Muse Guy, X 2026-09-13

The queue is changing from the other side too: above, a consumer agent opens Amazon's support chat for its owner and extracts a $70 refund a supervisor had promised. One tap for the customer, unlimited polite retries for your team; policy that reads clearly to a person and vaguely to a machine will be tested by machines first.

What separates the deployments that pay

The money went into the layer that talks. Much less went into the layer it has to reach: entitlements, the returns rule, live stock and the pricing engine running quietly on a maintenance line. […] Fund the conversation second. Fund what it reaches first.

George Mudie — LinkedIn, 13 Sep 2026
  1. Connect the systems of record first

    An agent wired to the returns rule and live stock gives the real figure. Disconnected, it gives a confidently wrong one just as fast.

  2. Choose by commercial meaning, not ticket volume

    ASOS's CTO: "the biggest mistake a company can make is investing in use cases that are not commercially meaningful."

  3. Meter the resolved case, not the conversation

    Under per-resolution pricing, resolution rate is a cost driver too. Track cost per resolved case before and after, AI line included.

  4. Instrument what happens after the handoff

    Confidence thresholds, a real route to a person, and outcomes tracked on deflected tickets: refunds, repeat contacts, review sentiment. Deflection alone produced the 42%.

Chapter 10

Product imagery got 50x cheaper and nobody measured what it sold

A fourth scraping round aimed straight at this area produced exactly one controlled measurement — and it is about the sentence a model writes over your reviews, not about pictures. The payoff that is actually checkable still sits upstream of the image: complete, structured product data is what the AI channels retrieve.

What holds up

  • This is the thinnest evidence base in the report. Of the ten facts in this chapter's original ledger, five are secondhand and five vendor or unsourced; none first-party or research grade. A fourth screening round aimed at this area added one research-grade item, and try-on returned four usable posts — the thinnest return of any area swept.
  • That one result is about text, not imagery: in a randomized field experiment on 35,625 users, an AI review summary framed as customer consensus raised product clicks 24% (Wang, Zhang, Gao & Tan, Sep 2026, SSRN working paper, post).
  • Asset production cost has genuinely collapsed. One global FMCG company cut digital asset creation from $100 to under $2 each, per a consultant interview (post). The company is not named.
  • The only number attached to a measurement system is Google's: high-quality product images drive up to a 6% lift in Shopping ad clicks, Google internal data (post). "Up to" is doing work in that sentence.
  • The virtual try-on figures everyone quotes — +50% conversion, 3x add-to-cart — come from 2026 studies whose publisher is never named, and they compare shoppers who used try-on with shoppers who did not (post). L'Oréal's 3x and Irisphera's 3.5x arrive the same way, with no period and no baseline.
  • Hard image rules arrive before any of this is proven: from 31 January 2027 Google Merchant Center enforces a 500x500 px minimum, up from 100x100, and recommends 1,500x1,500 (post).

Say the evidence is thin, because it is

Across the whole corpus, 87 of 329 de-duplicated numbers trace to a named first-party disclosure or a research firm with a stated method. In this chapter's original ledger that count is zero. Imagery, product copy and try-on are where operators spent real money in 2026 and where the industry published the least measurement. Read everything below against that.

We tested that rather than repeat it. A fourth scraping round, screened to the same rubric, was pointed at exactly these areas. Try-on, searched for on purpose, yielded four more qualifying posts: a vendor announcing Zara's entry, a listicle about L'Oréal, a note on fit-memory startups, a vendor's own usability observations. None controlled — the thinnest return anywhere in the corpus.

One controlled experiment — and it is about words, not pictures

The sweep did turn up one thing this chapter did not have: a randomized field experiment. Across 35,625 users reading AI-generated summaries of customer reviews, researchers varied one element — how strongly the summary asserted that customers agreed with each other. Consensus framing raised product clicks 24% and cut reading time 10%. The paper is Wang, Zhang, Gao & Tan, "When AI Speaks for the Crowd", September 2026, an SSRN working paper, not peer reviewed, and the post relaying it is its only record in this corpus (post). Treat the effect size as provisional; treat the design as the point. Every other number in this chapter compares self-selected users with everyone else — this one holds the shopper constant and changes only the machine's sentence, and the arm that made people read less sold more. That summary is product content somebody else writes about you, in wording you do not set.

Cost per asset fell. Cost per approved asset did not.

The production story is the most credible thing here, because the people who did the work tell it. A designer at Every Man Jack produced 2,800 new product images in a couple of months, white-background main shots through to generative lifestyle imagery, syndicated through the company's PIM into Amazon, Walmart.com, Target.com and Kroger.com. A separate practitioner reports 5,000 AI image edits for about $200 — and a week of human QA on top. That second half is the number to plan around: generation is nearly free, verification is a salaried person's week.

The failure mode is not a bad picture. It is drift across a set.

By pose seven, something shifts. The jawline softens. The proportions stretch. A PDP built from ten different women isn't a catalog. It's a casting error nobody caught.

Walid M Samuel, visual production at Pixofix — LinkedIn, 9 Sep 2026

Every Man Jack's VP of marketing says the same from the CPG side: models still botch small captions on packaging, so designers stayed in the loop rather than being replaced by one upload.

Try-on: the loudest numbers are the weakest

The commercial case is easy to state. Online fashion returns run 20-40%, wrong size the leading cause; US retail returns total about $816 billion (NRF) and cost 20-30% of revenue. Only the $816B and the cost share are sourced; the rest circulates unattributed.

Hobbyist personal-shopper prototype built with Codex
Figure 50Hobbyist personal-shopper prototype built with Codex: complete outfits picked from the user's taste and past purchases, rendered on the user as an 'AI styling preview'.Source: THREAD (prototype) · via Nika Sakandelidze, X 2026-09-08

What try-on returns in exchange is where the evidence gives out. The capability is real and now buildable by one person in a weekend, as the prototype above shows. The results attributed to it are not. A widely shared post cites a 2026 study of 1.2 million fashion shoppers at 50% higher purchase conversion and 3x more cart additions among try-on users, and a second of 577 Shopify stores at 3.1x add-to-cart, 90% of them on mobile. Neither publisher is named anywhere in the corpus, and both compare people who chose the feature with people who ignored it — a shopper who uploads a photo of herself to see a dress on her body has already decided she wants the dress.

What 2026 changed is the scale of deployment, not the quality of the proof. Zara entered virtual try-on, the clearest mainstreaming signal in the corpus. L'Oréal's are the largest numbers anyone quotes — 1 billion virtual try-ons, 20 million skin diagnostics, 3x higher conversion — reaching us through a marketing listicle on X, no period, no baseline, no L'Oréal document behind them (post). Irisphera, posting off Zara's move, claims 45,000+ products processed, 5-15 second renders, up to 3.5x conversion — and "20,000+ successful purchases with zero returns" (post). Zero returns on twenty thousand fashion purchases is not a measurement. When a vendor's headline number is impossible, read the ones beside it as advertising too.

Handle with care

The return-rate claims are no better: AI sizing "reportedly" cutting returns 30-40%, AR try-on about 20% — early pilots self-reported by the brands and vendors running them, no named study behind either. Keep them out of a business case. Google announced in March 2026 that its standalone try-on app Doppl would shut down, folding the technology into Search and Shopping listings — a signal about where try-on lives, not what it earns.

The part with a checkable payoff

One place this work pays off without trusting anyone's study: retrieval. Chapter 02 documents the shift — on 10 July 2026 the share of ChatGPT Shopping recommendations sourced from merchant feeds jumped from 8.26% to 61.54% in a single day (Profound, ~1.75M prompts), and Shopify reports AI searches using its Catalog convert at twice the rate of those relying on scraped data. Retrieval runs on structured attributes, not art direction.

CMS image editor with one-click AI actions to generate product alt text and image descriptions
Figure 51CMS image editor with one-click AI actions to generate product alt text and image descriptions.Source: Sanity · via Brodie Clark, X 2026-07-12

Adobe's Content Visibility Checker puts average language-model visibility on retail product pages at 66% (post): a third of what you wrote, shot and priced never reaches the software doing the shortlisting. Syndigo's State of Product Experience 2026 survey (8,736 consumers, six countries) has the shopper side — 83% leave for a competitor when the information they need is missing, 26% are more likely to buy what an AI recommends (post). Coverage and attitudes, not lift: they size the leak, they do not price the fix.

That reframes what "good imagery" means. Google's own guidance is unusually concrete for this chapter: crisp, multi-angle assets, up to 6% more clicks, and a 500x500 px floor — up from the legacy 100x100 — enforced by disapproval from 31 January 2027, with 1,500x1,500 recommended and secondary angles supplied through additional_image_link. Alt text is that job in text. Catalogue audits are it again, with the same evidence problem. A brand strategist put a few live product pages through an AI auditing tool and got back 511 inconsistencies: a missing video on a Caudalie page, a blurry image and mismatched pricing at London Victorian Ring Co, wrong descriptions on LG pages — drift no merchandising team reads its way through at catalogue scale. The recovery figures she quotes (Caudalie €700k recovered, LVR +7% conversion, LG $900k captured) come from the tool's vendor, in a post offering readers free access to it. ELEMIS reports an Akeneo agent turning one launch's three-hour product-data investigation into 30 minutes, a human approving before anything went live; the source is an Akeneo employee. Run the audit. Treat the numbers beside it as marketing.

It has forced us to really critique our imagery — and evolve our focus from aesthetics to retrievable, qualifying information. Better not just for agents, but human shoppers too.

Nicklaus Hasselberg, VP Marketing & Omnichannel, Every Man Jack — LinkedIn, 18 Aug 2026

Disclosure became a legal surface in 2026

Two rules landed while the measurement question was still open. A New York law (GBL §396-b) took effect in June 2026; Amazon began requiring sellers to tag AI-generated people in product photos and video the following month, and TikTok and Meta label realistic AI content automatically. In the EU, the AI Act requires visible disclosure of AI-generated or manipulated imagery, media published before 2 August grandfathered; MANGO, Adidas and Boozt were among the first storefronts seen complying. Tolerance is its own variable: one performance marketer wrote that AI models on an apparel site "completely threw me off from considering a brand".

Operator note

Spend the imagery budget where the return is checkable. Fix the feed first: resolution floor, multi-angle coverage, real alt text, complete attributes including the plain-language use case a model needs to place your product in an answer. Then run try-on as your own experiment — hold out a cohort, measure conversion and return rate on your store, ignore the published multiples. Nobody in this corpus has done that work for you.

up to +6%

Click lift on Google Shopping ads from high-quality product imagery, per Google internal testing, 2026.

Named but secondhand Google internal data · post

$100 → <$2

Digital asset creation cost per unit at one unnamed global FMCG company, a 95% cut that funded 50x more personalisation, 2026.

Named but secondhand Consultant interview, Karmesh Vaswani · post

+24%

Product clicks when an AI review summary framed customer opinion as consensus, against the weaker framing, in a randomized field experiment on 35,625 users, Sep 2026. SSRN working paper, not peer reviewed.

First-party / research Wang, Zhang, Gao & Tan, "When AI Speaks for the Crowd" · post

+50% / 3x

Conversion and add-to-cart for virtual try-on users vs non-users, 2026 studies whose publisher is never named. Users chose the feature; this is not measured lift.

Vendor or unsourced Unnamed 2026 VTO studies · post

Chapter 11

The numbers that do not survive checking

Of the 21 claims in this chapter's ledger, three trace to a first-party or research source and fourteen are unattributed, misquoted, or in one case satire that travelled. The most-quoted number in AI commerce exists in nine published versions, and the two largest have no source at all.

What holds up

  • Pew Research Center: 49% of US adults use AI chatbots, only about 29% trust their answers (post).
  • The Harris Poll, 3,000 consumers: 78% assume brands pay to be recommended by an AI; 72% accept AI shopping help only if they make the final call (post).
  • Deloitte: 20% of consumer-industry companies have mature governance for autonomous agents — just as those agents are handed payment credentials (post).
  • Every assistant-user figure in this report — Sparky +40% per order, Mylow ~3x conversion — compares people who chose the tool with people who did not.

Nine numbers, nine different comparisons

"AI traffic converts better" moved more budget in 2026 than any other claim, and nothing in this report has less agreement about what it means. Nine versions circulated, from +31% to +100%, and no two measure the same thing.

"AI traffic converts X% better": nine numbers, nine different comparisons

Every figure is a lift over a different baseline. The last one has no traceable source.

0%25%50%75%100%Netalico · 9 Shopify Plus stores vs organic100%Shopify Q2 2026 · vs search, PDP-landing (secondhand)80%Adobe Jul 2026 · vs all non-AI traffic60%"Adobe 54%" · no matching release54%Shopify Q1 2026 · vs organic search, PDP-landing50%Ethercycle · 10 Shopify stores vs Google46%Adobe Mar 2026 · vs all non-AI traffic42%Visibility Labs · ChatGPT vs non-branded organic31%Adobe panel, holiday 2025 · vs other channels31%
Named but secondhandAdobe Analytics; Shopify commerce data; Visibility Labs; Ethercycle; Netalico. The 54% figure is attributed to Adobe in posts but matches no Adobe release in this corpus.
Show data table
Netalico · 9 Shopify Plus stores vs organic100%
Shopify Q2 2026 · vs search, PDP-landing (secondhand)80%
Adobe Jul 2026 · vs all non-AI traffic60%
"Adobe 54%" · no matching release54%
Shopify Q1 2026 · vs organic search, PDP-landing50%
Ethercycle · 10 Shopify stores vs Google46%
Adobe Mar 2026 · vs all non-AI traffic42%
Visibility Labs · ChatGPT vs non-branded organic31%
Adobe panel, holiday 2025 · vs other channels31%

Adobe's +42% (March 2026) and +60% (July 2026) compare AI-referred traffic with all non-AI traffic — paid search, email and affiliates included, a mixed bag holding some of the worst traffic a retailer buys. Shopify's ~50% (Q1 2026) compares with organic search only, and only for sessions landing on a product page: a much harder baseline, which is why the smaller number is the more impressive one. Visibility Labs' 31% is ChatGPT against non-branded organic. Ethercycle's +46% is ten stores; Netalico's ~2x is nine.

Adobe Analytics readings of AI-referred vs non-AI conversion
Figure 52Adobe Analytics readings of AI-referred vs non-AI conversion: -38% in Mar 2025, +42% in Mar 2026, +54% in May 2026. Only three points; the path between is unknown.Source: Adobe Analytics · via Viktorijan “Vick” Mucunski, X 2026-09-07

One publisher measuring one thing does not give a stable number either. Three Adobe readings of the same comparison run -38% (March 2025), +42% (March 2026), +54% (May 2026). The sign flipped inside a year. A plan built on the March 2025 reading would have written AI referrals off as below-average traffic. Three points is not a trend, and the path between them is unknown.

What this report refuses to use

Seven widely-shared figures did not survive checking. They are named rather than quietly dropped, because all seven are still circulating.

Number in circulationWhere it came fromWhy it is not used
AI converts 54% better / a "54% conversion rate"Two posts citing "a recent Adobe study"No linked Adobe release carries it, and a 54% conversion rate is not a number a retail site produces.
Shopify: AI sessions convert ~80% better than organicTwo newsletters, near-identical wording, no linkShopify's published Q1 figure is ~50%. The 80% is likely the secondhand Q2 "1.8x after landing on a PDP" with its qualifier stripped.
Claude 16.8% / ChatGPT 14.2% / Google organic 2.8% CVROne AI-visibility account on XNo source, no sample; a 5–6x channel gap exceeds every measured study.
AI traffic converts at ~6% versus ~1%A post relaying an unnamed siteUnverifiable, and six times the largest measured gap.
AI buyers 3x likelier to convert, at half the basketUnnamed research, shared secondhandShopify's measured AOV difference is +14%, not -50%.
Walmart +22% AI-driven ecom growth; Amazon -$1.5B delivery; Target +12% impulseAn unsourced listicleWalmart reports 24% ecommerce growth and does not attribute it to AI.
23–60% of consumers use AI to shopFive posts, no source or definitionNamed alternatives exist: Similarweb observed AI in 11.4% of journeys; BCG's 13,000-consumer survey says about a third.

Correlation wearing a lift's clothes

The second failure mode is quieter, because the sources are impeccable. Walmart told investors Sparky users spend about 40% more per order than non-users (post); Amazon reported over 40% for its assistant; Lowe's says Mylow users convert at roughly 3x non-users (post). Real disclosures, named executives, and not one of them a lift.

Walmart's and Amazon's '+40% per order' compare assistant users with non-users — self-selection, not lift
Figure 53Walmart's and Amazon's '+40% per order' compare assistant users with non-users — self-selection, not lift. Albertsons' +26% and +10% compare one experience against another.Source: Manish Sharma; Q2 2026 earnings calls · via Manish Sharma, LinkedIn 2026-09-09

A shopper who opens an assistant mid-basket is already further into a purchase than one who does not, and an earnings call cannot separate cause from consequence. The contrast in the figure is the useful part: Albertsons' +26% and +10% compare one experience against another — a comparison that can come out wrong. A case study of users versus non-users shows you an audience, not a result.

Small bases move fast, then stop

Shopify's AI metrics from two disclosures side by side
Figure 54Shopify's AI metrics from two disclosures side by side: Q1 blog (>8x sessions, ~13x orders) vs Q2 call (~3x traffic and orders). Different bases; don't mash quarters.Source: Shopify blog (May); Shopify Q2 earnings call (Aug 5) · via Clausius, X 2026-09-12

Shopify's two 2026 disclosures sit badly beside each other: a Q1 blog reporting over 8x AI sessions and about 13x AI orders, then a Q2 call reporting roughly 3x for both. Both are official and nothing broke in between — Q1 was measured off a near-zero base a year earlier. Multiples decay while absolute volume rises. The absolute share is the figure nobody quotes.

On a store we work with: ChatGPT is 0.27% of revenue, +42% MoM. It's a small slice with a steep curve. The right approach here is "monitor monthly" not "overhaul everything."

Kurt Elster, Shopify consultant — X, 12 May 2026

The plumbing moved under the dashboard

Google AI Overviews now cite the top-10 results half as often

Share of AI Overview citations that come from a page already ranking in the top ten.

0%20%40%60%80%A year earlier76%202638%
First-party / researchAhrefs, 863,000 searches and 4M AI Overview links.
Show data table
A year earlier76%
202638%

Rank stopped predicting citation. Across 863,000 searches and four million AI Overview links, Ahrefs found the share of citations coming from a page already in the top ten fell from 76% to 38% (post). A team still reporting "we hold position 3" is reporting a proxy that lost half its predictive power.

GA4 for one site
Figure 55GA4 for one site: ChatGPT sessions moved from 'referral' and '(not set)' to a new 'ai-assistant' medium in mid-June 2026; reports filtering on 'referral' would show a collapse.Source: Google Analytics 4 · via Ann Smarty, X 2026-09-12

The labels moved too. In mid-June 2026, ChatGPT sessions on this site stopped arriving as referral and (not set) and began arriving under a new ai-assistant medium; every saved report filtered on referral showed AI traffic collapsing on a Tuesday. Nothing collapsed. Then there is who is at the other end: Cloudflare puts bots ahead of humans at 57.5% of website traffic, and Lunio's 2026 report puts invalid ad traffic at 24.2% on TikTok and 7.6% on Google — discount that second one, because Lunio sells click-fraud protection.

Trust is a ladder, not a number

Trust figures look contradictory until you notice the surveys ask about different rungs. Researching with an AI, letting it build a cart and letting it spend money are three decisions, and the drop-off between them is steep.

Visa Trust Index (2,065 US consumers, May 2026)
Figure 56Visa Trust Index (2,065 US consumers, May 2026): only 23% trust GenAI with a payment, while 61% would trust Visa to handle it.Source: Visa Trust Index survey · via Bharat Melag, LinkedIn 2026-09-13

Visa's Trust Index puts the bottom rung at 23% of 2,065 US consumers willing to trust GenAI with a payment — and 61% when Visa is named as the party handling it. Trust transfers to an institution; it is not extended to the model. That is a design instruction: let the assistant research, and put a name the shopper already banks with on the payment step.

49% / 29%

US adults who use AI chatbots, versus those who trust their answers.

First-party / research Pew Research Center · post

78%

Of 3,000 consumers assume brands pay to be recommended by an AI; 45% want it disclosed.

First-party / research The Harris Poll, "The Algorithmic Aisle" · post

57.5%

Of website traffic is bots, ahead of humans at 42.5%. Site mix matters.

Named but secondhand Cloudflare · post

Handle with care

A cluster of UK consumer figures travelled this quarter: 60% would abandon an AI agent after one wrong answer, 10.6% trust AI recommendations, 64% want agents to shop for them. No post names the survey, and the last two cannot describe the same population. Read them as evidence that sentiment is unsettled, never as a planning input. Same for "95% of GenAI pilots show no P&L impact" — traceable to a 2025 MIT report none of the posts names, and not ecommerce-specific.

The costs that arrive without a dashboard

Agentic checkout moves a liability before it moves a number: under the Agentic Commerce Protocol the merchant stays merchant of record and keeps refunds, chargebacks and compliance. Meanwhile the tolerance narrowed. On 1 April 2026 Visa's VAMP excessive-dispute threshold dropped from 2.2% to 1.5% for US, Canada, EU and APAC merchants, with fees around $8 per disputed transaction — secondhand from a conference write-up, so check the Visa bulletin before it enters a model.

When a customer disputes what their agent bought, who produces the evidence?

Richard Emanuel, on an Anthropic agentic-commerce session — LinkedIn, 9 Sep 2026
Dashboard lifts of 68–112% in free-listing clicks shrank to +18 to +80pp after removing seasonality and an untreated control market; about two-thirds was the calendar
Figure 57Dashboard lifts of 68–112% in free-listing clicks shrank to +18 to +80pp after removing seasonality and an untreated control market; about two-thirds was the calendar.Source: not shown · via John Caiozzo, X 2026-09-08

Produce your own number

The way out is not a better source; it is a measurement you ran yourself. Here is what the correction usually looks like: dashboard lifts of 68–112% in free-listing clicks shrank to +18 to +80 percentage points once seasonality was removed and an untreated control market used. Two-thirds of the headline was the calendar — assume that error inside any uncontrolled before-and-after, including your own.

  1. Fix the labels before the funnel

    List every AI referrer visible in analytics and check how each is classified, then re-check after each platform release. A report filtered on referral silently stopped counting ChatGPT in June 2026.

  2. Take the bots out first

    Split crawler and agent hits from human sessions before computing any rate. A conversion rate over unfiltered sessions is not a conversion rate.

  3. Write the baseline next to the number

    Never record "+X%". Record "+X% versus organic search, sessions landing on a product page, Q3, n=___". Most conflicts above dissolve once the baseline is attached.

  4. Separate who chose from who was assigned

    Users versus non-users is an audience comparison. For a lift, hold out a random slice of eligible sessions, or compare two experiences head to head.

  5. Use a control market, strip the calendar out

    Ship the change in one market or collection, leave a comparable one untouched, and report the result in percentage points rather than multiples.

  6. Price the downside in the same review

    Dispute rate, per-dispute fee, AI-surface spend with no attributable orders and bot share belong on the page with the lift. An upside-only channel review is a marketing document.

Operator note

Monthly cadence, high reaction threshold. AI referral revenue at most Shopify stores is still under 1% on a steep curve: that justifies instrumentation and a named owner, not a replatform. Re-derive the figure each quarter — multiples off small bases fall even while the business grows, so the number you quoted in Q1 will be wrong by Q3 for reasons unrelated to your work.

Chapter 12

AI UGC and the creative pipeline, as operators actually run it

Strip the screenshots away and the AI-creative posts describe one workflow, assembled the same way by operators who have never met: pull the competitor ad library, script, storyboard, generate, fan out. Production cost has fallen out of the equation. The research at the front and the judging at the back are where the work moved.

What holds up

  • The pipeline is standardised. Operators on three continents describe the same five steps with interchangeable tools — that is a practice spreading, not a proven result. Machina · John · 余温
  • Adoption is one of the few measured things here: 62% of ad buyers used generative AI for video creative in July 2026, up from 51% (IAB); 26% of marketers use digital replicas (XR Extreme Reach, 2026). post · post
  • Every cost figure here is an operator's own report — $0.68 and nine minutes for a finished ad, $999 for six human-filmed ones. Nobody publishes what the rest of the batch did. Vendor or unsourced
  • Volume moved; the hit rate did not. Motion still measures ~5% of creatives winning and absorbing 55% of spend across 578,750 creatives — chapter 07's finding, and the reason a generator alone changes nothing. post

The pipeline as it is assembled today

Nobody starts at the generator. Every serious version starts in a competitor ad library.

it starts in the meta ads library, pulling the winning ads in the niche before writing a single line

Machina, AI ad operator — X, 9 Sep 2026

MAX gives the model product, audience, offer and example ads, plus browser access, then tells it to open the Meta ad library and search five named competitors. Mike Futia's scrapes competitor pages, watches every video with a model and returns a brief with ten concepts in the brand's voice; his TikTok variant turns winning videos into shot-ready briefs. Step one is research, and it is the step most brands skip.

After the brief the chain is short. Script and shot list come out of the same conversation; a storyboard image is generated first as the anchor — one reference image driving every beat and camera angle — which stops the shot-by-shot drift that makes AI video obvious. Then it is a model call: Seedance inside a Claude skill, Higgsfield over MCP for research, scripting, images and animation in one conversation, Gemini Omni for multi-shot ads with a consistent creator across scenes, Kling over MCP for talking head, product shots, narration and music in one pass. On their own timing a variant takes minutes, a new concept a working day including research.

Two details separate daily operators from demos. Prompting: Stephen Bishop runs three fixed clauses on every generation, skin texture first, and argues people switch models when they should fix the prompt. Editability: Higgsfield's Motion Designer returns an After Effects project where camera, lighting and typography stay editable — an asset you revise rather than regenerate and hope.

Framework for a DTC marketing knowledge layer
Figure 58Framework for a DTC marketing knowledge layer: performance, customer and brand data centralised so creative, retention, PDP and launch agents work from real evidence.Source: @shannholmberg · via Shann³, X 2026-04-21

The step nobody screenshots is the one Shann Holmberg draws: agents only produce output as good as what they were given. Performance data, customer language, brand rules and competitor context in one structured place, loaded before anything is written. Skip it and you get generic output, then blame the model.

Where the human still sits

Three places: choosing the angle, approving the cut, owning the account the variants go into. The clearest evidence that generation is not the hard part is a company reversing. Icon raised $9.2M for an AI admaker, then pivoted to running the human UGC workflow — creators, shipping, scripts, coaching, editing — at six filmed ads for $999, keeping AI for everything around the shoot.

A company that originally built an AI admaker ended up putting humans back into the part of the process that actually produces the creative.

Divyanshi Sharma — LinkedIn, 25 Aug 2026

Same shape in imagery, better sourced in chapter 10: 5,000 AI image edits for about $200, behind a week of human QA.

What is winning attention right now

Operator opinion, flagged as such. Song ads are the format of the moment, and the operator running them is their sharpest critic.

They retain attention. They get cheap clicks. They can absolutely crush CPA. But cheap acquisition and great acquisition are two different things.

Franky Shaw, ecommerce operator — X, 12 Sep 2026

Claymation is second, on an arbitrage argument: it reads as expensive, almost nobody runs it, and an eight-step workflow replaces the studio. Three months later it is taught in French and packaged into Chinese cross-border tooling — the half-life of a creative arbitrage. AI personas are third: one TikTok Shop account runs a single synthetic presenter across dozens of products, top videos at 1M–4M views by the poster's count. Long-form autopilot video is fourth. Static remix is fifth and least discussed, though most merchants could run it tomorrow: 45+ on-brand statics from one product and one angle, or Oliver Kenyon's tighter version, six carousel images where each has a defined job. That one sits on the PDP, not the feed, and it is the cheapest funnel test here.

Meta Business AI answers shoppers across WhatsApp/Messenger chats, Facebook and Instagram ads (suggested questions) and brand websites
Figure 59Meta Business AI answers shoppers across WhatsApp/Messenger chats, Facebook and Instagram ads (suggested questions) and brand websites.Source: Meta · via wetracked.io, X 2026-08-07

The funnel under the format is moving too. Meta now answers shoppers inside the ad unit as suggested questions, then carries the conversation into Messenger, WhatsApp or the brand's site. If that holds, the ad's job shifts from carrying the pitch to starting a conversation something else finishes — which makes the hook, not the close, the part worth generating fifty of.

The economics operators report

$0.68

Compute for one finished AI UGC ad, nine minutes end to end, he reports. A later post gives $0.80 and 18 minutes.

Vendor or unsourced Sam, operator post · post

$999

Icon's price for six human-filmed UGC ads, after pivoting away from AI-generated creative, Aug 2026.

Vendor or unsourced Icon, via Divyanshi Sharma · post

62%

Ad buyers using generative AI for video creative, July 2026, up from 51%. The one adoption figure with a named source.

First-party / research IAB · post

What it replaces, in their accounting: creator fees of $150–500 a video, or $300–2,000 at 30–50 videos a month; retainers of $5,000–20,000 a month for 15–20 videos; a tool stack one post prices at $10,000+ a month. None are audited; their consistency tells you the market price, not the return. One operator did better than a claim: he ran the same listing task three ways and scored the outputs himself. Copy the method, not his verdict.

The durable half: judging creative, not producing it

Take the video out and what remains is a research stack that is aging better. Teardowns that score every creative an account ever ran and say why. Systems that scrape Instagram and TikTok and rank clips by sales intent. Vendors turning the Meta ad library into a chat over three million ads. Workspaces giving a brand's agents one context library instead of five silos, like Shadow. Skill libraries too: a marketing pack for coding agents credited with more than 48,000 GitHub stars, and a product-photography workflow one person spent 15 hours and $478 to nail, then published free.

The most interesting build is a scoring loop: ten runs against a written eval, prompt rewritten, retested, winner kept. A hook writer scored 32/50, then 47/50. His eval, his scale, no outside check. The shape is right though: if ~5% of creatives carry 55% of spend, the scarce skill is a rubric that predicts which 5%.

Buyer-profile generation case
Figure 60Buyer-profile generation case: a distilled 0.8B student model scored 84.6 vs its large teacher's 83.0, with an 8x shorter prompt and daily output rising from 2M to 72M profiles.Source: Shopify (per post; not named in image) · via AlphaSignal, X 2026-09-07

Rubrics also industrialise in a way generators do not. Here a small distilled model scores 84.6 against its much larger teacher's 83.0 on buyer profiles, with an eight times shorter prompt and daily output rising from 2 million to 72 million: once judgment is written down, it gets cheap enough to run on everything. The attribution to Shopify comes from the post, not the slide — read it as a pattern, not a disclosure.

Where it breaks

Ratio, not volume. Making 500 ads instead of 50 multiplies the denominator. If you cannot tag variants and read which carry the spend, more creative is a larger bill.

Slop. A practitioner ranking ecommerce ad formats puts AI UGC in F tier, below repurposed TikToks, the same week other operators call it the whole business. Both are opinions; the disagreement is the finding.

Brand risk compounds. A synthetic presenter across dozens of unrelated products is arbitrage with nothing to damage; a brand whose customers meet the same invented face next quarter is not. A vendor selling "the world's first ad that nobody knows is an ad" is describing a disclosure problem as a feature.

Disclosure is already law. The EU AI Act treats AI-generated imagery as a deepfake and wants it disclosed visibly to EU shoppers; by mid-August 2026 Linda Bustos could find three ecommerce sites doing it — MANGO, adidas, Boozt. Put the label in the asset template now; retrofitting 500 files is the expensive version.

Handle with care

Almost every number here was posted by someone selling the workflow. The techniques repeat across competing operators, which is worth something; but no post reports the ads that failed, the batch behind the winner, or the spend that proved it.

What a merchant should copy

  1. Build the input layer first

    One place holding performance data, customer language, brand rules and five competitors' live ads.

  2. Buy the judging before the generating

    Tag every variant at upload so you can answer which 5% carry the spend. Without that, stop here.

  3. Start with statics, not video

    A six-image PDP carousel where each image has a job is a same-week test with a readable outcome.

  4. Run one fixed task three ways

    Same brief, same inputs, three tools, your own scoring sheet — the only cost comparison that means anything for your catalogue.

  5. Keep a person on the angle and the cut

    The model is good at twenty variants of a chosen angle. Choosing which angle deserves twenty is still the job.

Operator note

Copy the research step, not the generator. The operators with the most convincing output start in someone else's ad library and end at a scoring rubric. The model in the middle is what they swap every quarter.

Chapter 13

The rest of the funnel, as operators actually run it

Between the click and the second order sits most of a merchant's real work: pages, assistants, flows, tickets, listings. Operators are running AI through all five, and the shape repeats — the machine does the assembly in minutes, and a human keeps the one decision that costs money if it is wrong.

What holds up

  • Four practices have spread far enough to count, because unconnected operators describe the same steps: page and listing audits, flow drafting over an MCP connection, catalog rewrites through Shopify's AI Toolkit, and WISMO prediction. Repetition is evidence a practice is spreading, not that it works.
  • The only surface with a documented floor under it is retention: automated flows produce 41% of email revenue from 5.3% of sends (Klaviyo 2026 benchmarks, 183,000+ brands). Everything in this chapter is being aimed at that ratio.
  • Almost nothing here has a control group. The two exceptions are retention tests: +70% browse revenue over a 30-day 50/50 holdout at Fellow and +$19,700 incremental on one live Klaviyo account (Revamp, Justin Guimera) — vendor self-reports, but with the right design.
  • The constraint is readiness, not model quality: 8% of retailers feel ready for AI agents to navigate their sites, and 52% still run out-of-the-box search (Algolia & The Retail Hive). The quarter's clearest failure fits that shape and not a model's: Starbucks retired its computer-vision inventory count after nine months.

From the click: landing pages and PDPs

The cheapest thing to build right now is a page audit. Yumna Aziz published the whole pipeline in n8n on 19 August 2026: a form takes the URL, an HTTP node fetches the HTML, it is converted to Markdown, and Gemini 2.5 Flash returns prioritised CRO fixes with line-by-line rewrites — one run at 29 seconds and 2,381 tokens, she reports Vendor or unsourced. Mike Futia built the same idea as a Claude Code plugin covering technical SEO, product schema, Core Web Vitals and AI-search readiness, scored 0–100, and reports it replaces a $200/month Ahrefs subscription Vendor or unsourced.

None of them tell you which fix to ship. They read copy, not behaviour, and the ranking they return is an opinion with a number on it. Point an agent at marketing work and the answers come back "correct, but not much use" — which is why the popular remedy is a skills pack (marketingskills, ~50 marketing skills, 48,000+ GitHub stars) rather than a bigger model. Operators keep buying context, not capability. The page still decides the sale: Contentsquare found 97% of shoppers would reconsider an AI-recommended purchase if the brand's site falls short Named but secondhand — a survey about a hypothetical, so a direction, not a rate.

The best-documented PDP workflow here is Connor Gillivan's week inside Glara, an agent reading a live Shopify catalog rather than an uploaded export, where nothing publishes without approval and any change rolls back.

Glara had already drafted the fix, and it wasn't a rewrite. It was one line: "when to use." Mist onto face, pat in, don't rub. I read it, changed a word, approved it. That was the whole job.

Connor Gillivan — LinkedIn, 8 Sep 2026

On-site: assistants, finders and the handoff

Product finders are the older, better-measured half of this. THG's Foundation Finder for LOOKFANTASTIC reports 5% revenue uplift Named but secondhand — behind it, 3,400 foundation samples read under spectrophotometers against 6,100 shade options. That is what a working finder costs.

On-site shopping assistant (GreatStore
Figure 61On-site shopping assistant (GreatStore.ai demo on fashion store ISTA) takes a shopper selfie plus chosen products and returns styling advice with a generated look.Source: ISTA / GreatStore.ai · via Bhanu Sharma, LinkedIn 2026-09-07

The rest of the category looks like the image above: a selfie in, styling advice and a generated look out. It is a vendor demo, not a deployment, and no conversion number has been published for it. Screenshots run well ahead of evidence here.

Conversational add-to-cart
Figure 62Conversational add-to-cart: the concierge asks which colour before adding the large fleece, then shows the updated bag (2 items, $290).Source: Après Studio (Salesforce Agentforce Commerce demo) · via Aaron Hutten, LinkedIn 2026-09-03

This one is a real capture, and it is the moment that matters: the assistant asks which colour before adding the fleece, then returns a two-item, $290 bag. Once an assistant can write to the cart, the question stops being answer quality and becomes permissions and escalation. Max Wallace, who discloses he is a Gorgias brand ambassador, reports 20% of conversations automated weeks after launching its AI Agent and Shopping Assistant Vendor or unsourced. Which model sits behind that permission is a procurement question. COLIBRIX ONE benchmarked 20-plus ecommerce agents over 2.4 million runs: the top-ranked model costs 14 times more and answers 8 times slower than the runner-up for three points of quality, and when planted text told the agent to run an unauthorised command, only one of the twenty refused Vendor or unsourced. The ranking is theirs; what it measures is the point. On a surface that can write to a cart, latency and injection resistance are the specification. Josh Gonsalves built the small version in Notion Custom Agents: the shopping assistant and the support agent are one machine seen from opposite ends of the funnel.

Lifecycle: the week that changed

Retention is furthest along because the tools opened first. Jordan O'Connor spent four days testing the Klaviyo × Claude integration on a client doing over $10M a year, and his framing is what to keep: the model is no longer analysing the account, it is building inside it. Tyler Phillips published the copyable version on 21 April 2026: load the brand guide, product doc and CDN image URLs first, connect Klaviyo over MCP, then ask for a three-email abandoned-checkout flow for a named segment and let it build the templates and wire the flow IDs. A month of calendar in under an hour, one variable tested per campaign. Kumud Deepali Rudraraju's Omnisend post is an ad; the part worth taking is that its MCP connection is read-only, so an audit runs without write access to the list. O'Connor's later write-up of K:BOS, 12 September 2026, records the same move as shipped product: 260-plus Klaviyo MCP tools, so an agent can set a flow live without opening the app, and predictive timing pulled inside the flow — a trigger that fires when a customer is likely to buy again rather than on a fixed delay. Timing is the part a merchant cannot draft by hand.

41% / 5.3%

Share of email revenue from automated flows, and the share of sends they use, across 183,000+ Klaviyo brands, 2026.

First-party / research Klaviyo 2026 Email Benchmarks · post

+70%

Browse revenue at Fellow over a 30-day 50/50 holdout of AI-personalised emails sent inside existing Klaviyo flows; 9x incremental ROI.

Named but secondhand Revamp case study · post

8%

Retailers who feel ready for AI agents to navigate their sites on a shopper's behalf; 52% still use out-of-the-box search.

Named but secondhand Algolia & The Retail Hive · post

What changed in the week is throughput, not judgement. Chase Dimond's agency shipped 40 client flow emails in one week without adding a retention strategist — "the queue just stopped being the ceiling." Aim it at the dead flow, not the calendar: average revenue per send on a post-purchase win-back is $0.11, against $5–17 for brands doing it properly Vendor or unsourced.

Vendor demo of an agent answering 'why is retention low?' by pulling 90 days of cancellation data from a DTC stack (Shopify, Recharge, Klaviyo, Gorgias) in about four minutes
Figure 63Vendor demo of an agent answering 'why is retention low?' by pulling 90 days of cancellation data from a DTC stack (Shopify, Recharge, Klaviyo, Gorgias) in about four minutes.Source: Hyperagent · via Eli Weiss, X 2026-05-19

The demo above answers "why is retention low?" by pulling ninety days of cancellations out of Shopify, Recharge, Klaviyo and Gorgias in about four minutes. The figures on screen are demo data — judge the shape, not the result. A general-purpose model cannot do it, because — as one email operator puts it — it "doesn't know your welcome flow is working and your win-back is quietly dead". Shann Holmberg's email walkthrough starts from the account's real open rate and revenue per send. Copy that order of operations.

After the sale is still the funnel

Support is a conversion surface that only ever shows up in the retention line. Dr. Squatch is reported to have cut WISMO contacts by 25%, at 94% delivery-date accuracy, by predicting delivery problems from carrier data instead of answering questions about them (Malik Usman, 7 Sep 2026) Vendor or unsourced. The same brand worked the other side of that page: recommendations placed on the order-tracking screen itself, clicked by 31.89% of customers and credited with $32,978 and 766 orders in Q1 2026 (Jalpesh Patel, 10 Sep 2026) Named but secondhand. A denominator and a period, which is more than most figures here carry — but no holdout, so it sizes the surface, not the lift.

Deflection is not the end of it: the answer itself is a place to recommend. Zak Cassady-Dorion's agency has Klaviyo's Customer Agent settle an address correction and then recommend a product off that order, and reports roughly a 10% lift on abandonment flows after swapping the discount code for a real reply to "what stopped you?" Vendor or unsourced. Shopify's write-up of LifeStraw puts a figure on both halves — 75% of inquiries resolved without a human, agent-recommended sales up 111% over 90 days Vendor or unsourced. The resolution rate is the number every vendor publishes and the wrong one to optimise alone. Between 25 April 2025 and 31 January 2026, e-commerce was the largest of 31 categories on India's National Consumer Helpline, at 47,743 refund grievances — behind a growing share of them, a bot with no visible route to a person Named but secondhand. A ticket closed fast and a problem solved are different events.

Handle with care

Deflection is billed. Gorgias charges per billable ticket and adds roughly $0.90–$1.00 per AI resolution on top, so a chat the agent resolves is charged twice Vendor or unsourced. Model the deflection rate against that per-resolution price before counting savings.

A seller reports being billed for shoppers clicking Rufus questions on its own listings
Figure 64A seller reports being billed for shoppers clicking Rufus questions on its own listings: 1,815 interactions, $2,941 spend, 70% with no sales and no incrementality data.Source: Seller post on Facebook (unnamed) · via Michael Patrón, X 2026-04-20

On somebody else's surface it gets worse, and the screenshot above is the warning: a seller reports being billed when shoppers click the assistant's suggested questions on his own listings — 1,815 interactions, $2,941, roughly 70% with no sale attached. On your own site you set the price of an assistant conversation. On a marketplace, somebody else does.

Merchandising, pricing and the catalog

Shopify's AI Toolkit put a write-capable agent in the terminal. Mike Futia's version is one prompt that reads the catalog, rewrites descriptions and pushes them live; Edward Deng's list of seven jobs adds re-ordering collections by what converts and post-purchase sequences written from order history. One item there is dangerous: inventory-aware urgency copy, rewritten automatically, is a claim about stock that has to be true. The remedy operators use is to shrink what the agent may touch: in Maxim Lazovsky's supplier pipeline the model normalises titles, extracts weights and drafts descriptions into a Shopify draft, but SKU, barcode, price, inventory and variant IDs stay protected, and an anomalous change is held for review automatically. ELEMIS runs the same split through Akeneo — the agent proposes a fix with its expected impact, a person approves — and reports one launch investigation dropping from three hours to thirty minutes Vendor or unsourced.

Google Merchant Center's new AI performance report splits AI-shopping share of voice by journey phase and query type; this merchant holds 1
Figure 65Google Merchant Center's new AI performance report splits AI-shopping share of voice by journey phase and query type; this merchant holds 1.0% vs competitors' 8.3%.Source: Google Merchant Center · via Brodie Clark, X 2026-07-14

This is the checkable payoff, and it points back to chapter 02's feed shift. Google Merchant Center is the first Google product to expose query data for AI Overviews and AI Mode, split by journey phase and query type; the merchant above holds 1.0% share of voice where a competitor holds 8.3%. That is a number you can act on without trusting anyone's case study. Marketplaces have the cheap version: Ruben Alikhanyan's audit is five buyer questions and fifteen minutes — he typed them into Rufus for a $4M brand's hero SKU and it was cited zero times. Pricing has no such instrument: the claim that AI Mode shows the same products at 21.6% higher prices is summarised secondhand from a study whose publisher is unnamed — a prompt to check your own feed, not a number to repeat.

The seller's posture

Agency case study
Figure 66Agency case study: personalised upsells on chocolatier Vocca's store reportedly lifted AOV 14.49%, influenced 14.93% of orders, and drove up to 12.66% of revenue via Rebuy.Source: Anphonic; Vocca; Rebuy · via Anphonic, LinkedIn 2026-08-21

Most pitches this quarter look like the card above: an agency case study, three precise percentages, one named brand, no control group, no denominator. Not dishonest, not evidence. What a deck never contains is a retirement. Starbucks spent nine months counting backroom shelves with iPad Pros and computer vision, then ended the programme in an employee newsletter Named but secondhand: reflective surfaces double-counted, staff were disciplined for the model's misses, and the inventory system underneath was 1990s AS/400 that could not hand it clean data. None of that was a model problem. Three buckets.

  1. Adopt now

    Anything that drafts for a human to approve and can be rolled back: page and listing audits, flow and segment drafting over MCP, catalog rewrites with the money fields locked, WISMO prediction, and recommendations on a page you already own, starting with the order-tracking screen. Worst case is a wasted hour. Simon Jackson's library of 50 store agents rests on the right premise — "most of them are checklists, not judgement" — and he reports a first run flagging 17 dead SKUs and $2,300 of discount leakage Vendor or unsourced. Start there.

  2. Pilot with a holdout

    Anything that claims revenue. Retention is the one surface where holdout testing is already normal; copy the design before you copy the numbers. Dr. Squatch's tracking-page result is the shape to repeat this way: a denominator, no control, one withheld cell would settle it. A 30-day 50/50 split costs nothing but patience, and it is the only thing that turns someone's case study into your own.

  3. Ignore until someone publishes a controlled number

    On-site assistant lift, AI dynamic-pricing margin claims, and every user-versus-non-user comparison in this report — two more landed this quarter, Loops at A101 Ekstra, +83% conversion, and Constructor's Ask Cleo, 8–9x add-to-cart, each comparing shoppers who chose the agent with shoppers who did not. Buy an assistant for deflection if that maths works; never on the strength of a gap between people who chose the tool and people who did not.

Operator note

Before adding another tool, ask where the context lives. Every workflow here that works shares one step — Glara reading the live catalog, Klaviyo driven over MCP, a stack synced into one brand library before any agent touches it (Jackson Corey, 2 Sep 2026). A tool that cannot read your account will keep producing answers that are correct and useless.

Appendix

Six months on one line, and the words everyone is using differently

What actually shipped between March and September 2026, dated, and a plain-language glossary — because half the disagreements in this corpus are two people using one term for two different things.

Timeline

Only events with a date a merchant can check. Quarterly figures are filed under the day they were disclosed, not the quarter they describe.

DateWhat happenedWhy a merchant cares
27 Feb 2026OpenAI and Amazon announce a $50B partnership; OpenAI winds down Instant Checkout, the in-chat buy button, in March.The first in-chat checkout closed before it proved itself. Selling into the answer is not the same as checking out inside it.
1 Apr 2026Visa's excessive-dispute threshold (VAMP) drops from 2.2% to 1.5% across US, Canada, EU and APAC.Agent-placed orders raise dispute risk while the merchant stays merchant of record. The tolerance for that risk just narrowed.
5 May 2026Shopify's Q1 disclosures: AI-driven traffic up 8x, AI-driven orders up ~13x, Sidekick usage up 4x.The multiples that still circulate. They are Q1, off a small base — see Chapter 04.
12 May 2026Shopify publishes the conversion read: AI product-page sessions convert ~50% better than organic search, AOV 14% higher.The most-quoted merchant-level number in the report, and the one with the clearest baseline.
10 Jul 2026ChatGPT Shopping switches retrieval: feed-sourced recommendations jump 8.26% → 61.54% in a day; top-10 merchant share 22.5% → 41.8%; 450 tracked brands lose a third or more of their visibility.The single most consequential day in this corpus for discovery. Your feed became your shelf.
5 Aug 2026Shopify Q2 call: AI traffic and orders both ~3x year over year; daily merchants using Sidekick up 3.6x.The honest read of the trend once the base fills in.
6 Aug 2026Klaviyo Q2: $371M revenue, +26% YoY, 205,000+ customers; agent and MCP surface expands.Lifecycle tooling is where AI reached general availability first.
19 Aug 2026AI1000 Q2: 972 of 1,000 retailers change rank, median move 35 places; ChatGPT-led retailers fall 844 → 722 as Gemini, Perplexity and Claude gain.Visibility in AI answers is volatile in a way search rank was not.
23 Aug 2026Amazon: assistant users spend 40%+ more. Target: AI-platform traffic small but growing 3.5x faster than the industry.Both are user-vs-non-user comparisons. Read Chapter 05 before repeating them.
31 Aug – 4 Sep 2026ChatGPT Ads Manager opens self-serve in 31 European markets, then India.Paid placement inside answers is now buyable by ordinary advertisers.
2 Sep 2026Amazon reports Sponsored Prompt clickers convert 48% more often and spend 21% more; ad revenue $19.8B in Q2.Retail media moved inside the assistant.
8 Sep 2026Meta launches the Muse shopping agent in the US; Shopify adds Meta as an agentic sales channel in the admin, on by default.A channel appeared in your admin without you switching it on. Check what it exposes.

Glossary

Where two sources in this report use a word differently, the disagreement is noted.

Agentic commerce

Shopping where software acts on the buyer's behalf: it searches, compares, assembles a cart, and in the strongest version pays. The forecasts disagree wildly because they draw the line in different places — McKinsey counts spend agents orchestrate (up to $1T in the US by 2030), Bain counts spend agents complete ($300–500B). When you read a number, ask which of the two it counts.

AI referral traffic

Visits arriving from an AI surface — ChatGPT, AI Mode, Gemini, Perplexity, Copilot. Measured inconsistently: as a share of referral traffic (Bessemer's 15–20%) or of all traffic (Bernstein's 0.8–1.4%). Same phenomenon, denominators 15x apart.

Feed-integrated retrieval

An AI assistant answering a shopping question from structured merchant product feeds rather than by reading web pages. After 10 July 2026 this is how most ChatGPT Shopping recommendations are sourced.

GEO / AEO / AI SEO

Optimising to be cited or recommended inside AI answers. The tactics that survive checking in this corpus are unglamorous: complete structured product data, entity clarity, third-party mentions. Rank alone predicts citation far less than it did — AI Overview citations from the top ten fell from 76% to 38% in a year.

UCP · ACP · AP2 · MCP

The plumbing. UCP (Universal Commerce Protocol, Google-led) and ACP (Agentic Commerce Protocol, OpenAI/Stripe) let an agent read a catalog and place an order; AP2 handles agent payment authorisation; MCP is the general connector standard that lets a model use a tool — the same standard Shopify and Klaviyo expose their admins through.

Agentic Storefront

Shopify's surface for agent traffic: catalog syndication, agent-readable policies, order attribution back to a channel. Enabled by default for eligible stores; blocking crawlers in robots.txt does not stop Catalog syndication.

AI UGC

Ad creative in the visual language of a customer testimonial, generated rather than filmed. Cheap enough to change the production question from "can we make one" to "can we judge fifty" — see Chapter 12.

Holdout

A slice of the audience deliberately left untreated so the effect of a change can be measured. Almost nothing in this corpus has one, which is why so many numbers here are comparisons rather than lifts.

Self-selection

The reason "assistant users spend 40% more" is not "the assistant adds 40%". People who open an assistant were already further along. Chapter 11 is mostly about this.

Method

How this was built, and what it cannot tell you

A corpus of public LinkedIn and X posts, screened post by post, with every number traced to whoever published it first. Good for reading what the market believes and checking whether it holds. Not a substitute for your own store's data.

5,328unique posts screened individually
1,000kept: on-topic and substantive
329facts after de-duplication
87rated first-party or research-grade
1,099images opened and classified
161figures that survived curation

Collection

Three sweeps of LinkedIn and X between 14 March and 14 September 2026, using 58 search queries covering every commercial surface: discovery and AI search, agentic checkout, personalization, recommendations, funnel and CRO, email and SMS, ads and creative, pricing, support, product content, try-on, merchandising, inventory, and Shopify's own AI features. That returned 6,840 raw posts, 5,328 of them unique after removing reposts and near-identical text.

Screening

No keyword filter decided what stayed. Every post was read and judged on four questions: is it substantively about AI applied to selling online; does it carry a concrete tactic, case or number; is it crypto, hiring, event promotion or an empty pitch; and which commercial surface does it belong to. Crypto was excluded outright — 177 posts, many of them using agentic-commerce vocabulary to sell a token. Off-topic accounted for 1,369, thin self-promotion 442, generic hype 339, hiring 100. What remained was ranked and the top 1,000 kept.

As a check on the screening, 21 verdicts were pulled at random and re-read by hand against the post text; all 21 held. That is a small sample. It catches gross failure, not drift.

The evidence ledger

Every number in the 1,000 posts was extracted with its claimed source, period, population and baseline: 711 raw claims. Those were then consolidated into 329 facts, because the same figure repeated by forty accounts is one piece of evidence, not forty. Each fact records the organisation it originated with, the posts that carried it, and a reliability grade applied uniformly:

  • First-party / research a named company disclosure (earnings call, official data release, documentation) or a named research firm or survey with a stated method and sample.
  • Named but secondhand the source is named but the post is a retelling, or it is an agency/merchant case with concrete specifics but no independent check.
  • Vendor or unsourced a vendor's own launch or marketing claim, an anecdote, or a number with no attribution at all.

Where facts disagreed, both were kept and linked, with the reason for the difference stated: a different quarter, a different baseline, a different population, or a misquote. Ninety-nine facts carry such a link.

Figures

Every image attached to a kept post was opened and classified — 1,099 of them — then the 382 plausible ones were re-examined against five tests: does it carry real information, will it read when printed, is the number in it sourced and consistent with the post, is it a genuine capture rather than a mockup, and is it a duplicate of a stronger image. 161 survived. Screenshots that turned out to be AI-generated illustrations carrying invented numbers were rejected, as were charts whose labels contradicted their own plotted values.

The second pass: the agentic-commerce corpus (Chapters 12–13a)

Chapters 12, 13 and 13a were built later and from a separate sweep, run on 14 September 2026, because the first corpus covered agentic commerce as one surface among fifteen and the subject had outgrown that share. Three pulls of LinkedIn and X returned 11,908 raw items, 9,105 unique after de-duplication. Two buckets were defined: posts whose own text contains the phrase "agentic commerce", and posts that do not but report a named platform's move in AI shopping, agent checkout or merchant AI tooling. Only posts carrying at least one image were considered, which narrowed the field to 3,780 candidates.

Those candidates were read one at a time — 2,728 of them — and graded on two axes: is the post or its author promoting crypto, and is the post substantive about agentic commerce or merely brushing against it. 249 were dropped as crypto and 587 as incidental (hiring ads, event promos, hashtag stuffing, stock chatter, and a large family of templated "I built a store with Claude in an hour" posts). The remaining pool was ranked and cut to 1,500: 500 phrase posts and 250 platform-move posts from each network.

All 1,840 attached images were opened and classified individually; 1,032 carried enough legible information to be usable and 31 are printed in these three chapters. Quotes in these chapters were machine-checked against the source post text: 26 that turned out to be paraphrase rather than verbatim were sent back and rewritten, and one claim that carried a figure absent from its cited post was removed outright.

Two differences from the first pass are worth stating plainly. Crypto was excluded far more aggressively here, because on X roughly a third of everything using the phrase was promoting a token — a stricter rule also cost the corpus a handful of legitimate posts whose authors work at crypto-native payment companies. And the monthly post counts in this second corpus cannot be read as market volume: the X sweep took the top posts per month and the LinkedIn sweep ranked by relevance, so the shape reflects the sampling, not the discourse.

What this method cannot do

  • It measures discourse, not the market. A corpus of posts over-represents whoever posts: vendors, agencies, consultants and people with something to sell. Quiet operators are absent.
  • It is English-first. Non-English posts were kept when relevant, but the queries were English, so the view of China, Japan, Korea and Southeast Asia is thin.
  • It cannot verify a first-party claim. When Walmart says its assistant users spend 40% more, this report records who said it, when, and what it compares. It cannot audit the number.
  • Six months is short. Several of the strongest findings — the July feed shift, the ads moving into answers — are weeks old at publication and may not hold.

Tooling

Collection ran through Apify actors for LinkedIn and X. All screening, number extraction, image classification and writing was done by Claude (Opus 5) under the rules above, with the outputs checked against the source posts. Both language editions were written independently from the same ledger rather than translated.

The ledger

Every first-party and research-grade fact used in this report, with its source, period, how many posts carried it, and a link to one of them.

ClaimValuePeriodSourcePostsPost
McKinsey estimates AI agents could orchestrate $3 trillion to $5 trillion of global consumer commerce by 2030.
forecast scenario, not measurement; posts rarely give the report title
$3T-$5Tby 2030McKinsey agentic commerce report5
McKinsey projects AI agents could orchestrate $900 billion to $1 trillion of US retail revenue by 2030.
forecast; some posts drop the 'up to' and state $1T as a point estimate
$900B-$1T (up to $1T)by 2030McKinsey agentic commerce report5
Narvar's survey found 65% of shoppers plan to use AI for holiday shopping, but only 8% of retailers feel very confident they are ready.
100 executives is a small sample
65% vs 8% (78% would use if more personalized)holiday 2026Narvar survey (1,348 shoppers, 100 retail executives)3
Deloitte found 56% of European consumers have shopped with AI at least once, crossing 50% adoption within 18 months.
'at least once' is a low bar
56%2026Deloitte consumer research2
Walmart US comparable sales grew 2.6% while US ecommerce grew 24% in Q2 FY27. +24% vs +2.6%Q2 FY27 (Aug 2026)Walmart Q2 FY27 earnings2
EMARKETER's base case has AI platforms and assistants directly driving 15.8% of US retail ecommerce sales by 2030, up from 3.2% in 2026.
forecast; base case scenario Chart period verified from the EMARKETER graphic itself (US AI-driven retail ecommerce sales, 2026-2030, July 2026 forecast); an earlier ledger pass recorded 2031.
15.8% (from 3.2% in 2026)by 2030EMARKETER forecast (base case)1
BCG's survey of 13,000+ consumers in 12 markets found nearly one-third use AI in their purchase journey, roughly triple the level 18 months earlier.
self-reported survey
~1/3 (3x in 18 months)2026BCG Center for Customer Insight survey (13,000+ consumers, 12 markets)1
Deloitte found 73% of consumer companies plan to deploy agentic AI within two years, but only 24% report even moderate adoption today. 73% vs 24%2026Deloitte State of AI in the Consumer Industry 20261
Only 23% of consumer companies have moved at least 40% of their AI experiments into production (Deloitte). 23%2026Deloitte State of AI in the Consumer Industry 20261
82% of consumer companies have not redesigned jobs around AI capabilities (Deloitte). 82%2026Deloitte State of AI in the Consumer Industry 20261
Alibaba's ecommerce AI agents for merchants attracted more than 60,000 paid users in five months. 60,000+five months to Sep 2026Nikkei Asia1
Shopify's Q1 2026 data: sessions landing on a product page from AI platforms converted about 50% better than from organic search, holding in 23 of 25 categories (average gap 56%).
like-for-like on PDP-landing sessions only; K:BOS session quoted 49%
~+50% (avg gap 56%; 23 of 25 categories)Q1 2026Shopify Q1 2026 commerce data (Shopify blog)11
Productrise tracked 2 million listings over 23 days in August 2026: Google AI Mode's lead offer averaged 21.6% more expensive than classic search for the same product.
single small-firm study; lead offer only; one post rounds to 21% and calls it a survey
+21.6%23 days, Aug 2026Productrise study (2M listings, 100,000 searches)10
John Lewis says product searches arriving through AI agents rose from 0.3% to 2.5% of its product searches in one year, roughly eightfold.
some posts misdescribe as 'traffic' or 'AI-influenced searches'
0.3% -> 2.5%one year to Sep 2026John Lewis (reported by Reuters)9
Adobe found that by March 2026 AI-referred traffic to US retail sites converted 42% better than non-AI traffic, with 48% longer visits and 13% more pages per visit.
site-level aggregate; AI visitors arrive later in the funnel (selection); some posts say 'vs paid search, email, affiliates' or 'every channel'
+42%March 2026Adobe Analytics Q1 2026 retail report6
Adobe Analytics, analysing over 1 trillion US retail site visits, found AI-referred traffic grew 393% year over year in Q1 2026.
growth off a small base
+393%Q1 2026Adobe Analytics (1 trillion+ US retail visits)5
Adobe measured AI-referred traffic to US retail sites up 693% year over year over the 2025 holiday season, and 805% on Black Friday.
one post rounds to 700%
+693% (+805% on Black Friday)holiday 2025Adobe Analytics holiday 2025 report5
Shopify's Q1 2026 data showed AI-referred orders had 14% higher average order value than organic-search orders. +14%Q1 2026Shopify Q1 2026 commerce data (Shopify blog)5
Similarweb: shopping journeys combining AI and search convert at 23%, versus 12.5% for AI alone, 84% higher.
journey-level conversion, not session conversion; not comparable to Adobe/Shopify rates
23% vs 12.5% (+84%)2026Similarweb State of Ecommerce 20265
Profound found the share of ChatGPT Shopping recommendations retrieved from integrated product feeds jumped from 8.26% to 61.54% on July 10, 2026, across ~1.75M prompts.
posts round to 8->62% or 8->65%; one says 'now ~65% of recommendations' (later level)
8.26% -> 61.54%July 10, 2026Profound research (~1.75M prompts)5
Adobe's August release found AI-referred retail visits converted 60% better than non-AI traffic in July 2026.
one post frames it as 'vs traditional search' over '11 months, 30B visits' - base differs by retelling
+60%July 2026Adobe Digital Insights (August 2026 release)4
About half of AI-referred sessions on Shopify land directly on a product page, versus roughly 20% for organic search (~2.5x). ~50% (vs ~20% organic; ~2.5x)Q1-Q2 2026Shopify commerce data (Q1 blog; Q2 update)4
Around July 10, 450 Profound-tracked brands lost at least 33% of ChatGPT Shopping visibility while 67 gained 33%+.
Profound's customer base, not all merchants; posts give denominators 687 or ~700
450 down / 67 up (of ~687 tracked)around July 10, 2026Profound research4
Adobe found AI-referred retail visits generated 53% more revenue per visit than non-AI traffic in July 2026. +53%July 2026Adobe Digital Insights (August 2026 release; Digital Commerce 360)3
Similarweb found AI appeared in 11.4% of shopping journeys in 2026, up from 4.5% in 2024.
panel-based clickstream
11.4% (from 4.5% in 2024)2026 vs 2024Similarweb x Statista State of Ecommerce 20263
Similarweb found 89% of shopping journeys involving AI also included search; only 11% used AI without search. 89%2026Similarweb State of Ecommerce 20263
Similarweb found AI referrals to ecommerce sites grew 203% year over year, faster than any other referral channel.
different panel and period from Adobe growth figures
+203%2026Similarweb x Statista State of Ecommerce 20263
Only 1.28% of listings overlapped daily between AI Mode and the Popular Products carousel; on 49.6% of shared products the top seller differed. 1.28%; 49.6%Aug 2026Productrise study3
Target's CEO said traffic from external AI platforms is still small but growing 3.5x faster than the industry average.
no absolute volume disclosed
3.5x fasterQ2 2026 earningsTarget Q2 2026 earnings call (CEO Michael Fiddelke)3
Adobe found AI-referred visits to US retail sites were up 62% year over year in July 2026.
decelerating growth as base grows
+62%July 2026Adobe Digital Insights (via Digital Commerce 360)2
Shopify said traditional search still drives roughly a third of storefront sessions and grew 1.3x over two years. ~1/3 (search sessions 1.3x over two years)Q2 2026Shopify Q2 2026 earnings call (Harley Finkelstein)2
Retailers whose primary AI referral source is ChatGPT fell from 844 to 722 of the AI1000 between Q1 and Q2 2026.
count of retailers, not traffic share
844 -> 722 of 1,000Q1 -> Q2 2026ReFiBuy & Digital Commerce 360 AI1000 report2
In the Q2 2026 AI1000, 972 of 1,000 retailers changed rank, the median moved 35 places and 641 moved 25+.
vendor index methodology
972/1,000 changed; median 35 places; 641 moved 25+Q2 2026ReFiBuy & Digital Commerce 360 AI1000 report2
Only 10 of the 100 largest online retailers made the AI1000 top 100; the new No.1, Nixon, ranks No.722 by online sales. 10 of 100 (No.1 Nixon ranks No.722 in sales)Q2 2026ReFiBuy & Digital Commerce 360 AI1000 report2
After ChatGPT Shopping's July 10 feed shift, the top-10 merchants' share of recommendations rose from 22.5% to 41.8% and unique merchants shown fell from 13,524 to 10,607. 22.5% -> 41.8%; 13,524 -> 10,607 (-20%)July 2026Profound research2
Ahrefs found the share of Google AI Overview citations coming from top-10 organic results fell from 76% to 38% in a year, across 863,000 searches.
one post attributes the 38% to AI Mode
76% -> 38%2025 -> 2026Ahrefs (863,000 searches, 4M links)2
Google AI Mode shows 3.9 products per answer versus 27.8 in the classic shopping carousel. 3.9 vs 27.8Aug 2026Productrise study2
BCG found 13% of consumers purchase whatever an AI tool recommends, and 43% feel overwhelmed by information.
self-reported
13% (43% overwhelmed)2026BCG Center for Customer Insight survey1
Similarweb's 2026 State of Ecommerce report found nearly 1 in 4 US shoppers ask AI before they buy. nearly 1 in 42026Similarweb x Statista State of Ecommerce 20261
Adobe reports traffic from AI sources to US retail sites grew 125% year over year between April and June 2026.
single post
+125%Apr-Jun 2026Adobe Digital Insights1
71% of Shopify's AI-attributed orders in 2025 came from long-tail, specialised products.
definition of long-tail not given
71%2025Shopify1
From Q1 to Q2 2026, AI1000 retailers with Gemini as primary AI referrer rose from 16 to 32, Perplexity 8 to 21, and Claude 1 to 15. Gemini 16->32; Perplexity 8->21; Claude 1->15Q1 -> Q2 2026ReFiBuy & Digital Commerce 360 AI1000 report1
The AI1000 Index Average fell from 42.0 in Q1 to 39.7 in Q2 2026.
proprietary index
42.0 -> 39.7Q1 -> Q2 2026ReFiBuy & Digital Commerce 360 AI1000 report1
The top 100 large, feed-integrated merchants win 50% more ChatGPT product recommendations than before the July 2026 change. +50%post July 10, 2026Profound research1
Profound classified 7.5M ChatGPT conversations: commercial intent rose from 13.9% to 19.2% in a year, an estimated 28 billion buying conversations annually.
annualised volume is an extrapolation
13.9% -> 19.2% (~28B buying conversations/yr)12 months to Sep 2026Profound (7.5M classified ChatGPT conversations)1
Profound found Claude searched the web in 93% of responses versus 13% for Claude Code, and only 1 in 5 brands appeared in both. 93% vs 13% (1 in 5 brands in both)2026Profound research1
Profound's analysis of 1.9B+ conversations across 50+ industries found 80% of industries have a different AI search leader in Europe than in the US. 80%Summer 2026Profound Index Report Summer 2026 (1.9B+ conversations)1
Pew found users clicked a result on only 8% of Google searches that showed an AI Overview.
Pew compared with 15% without AIO (not in post)
8%2025Pew Research Center browsing panel1
Semrush found 19.43% of consumers would choose an AI chatbot as their only pre-purchase information source, 27.37% of AI users and 44.39% of power users.
hypothetical choice
19.43% (27.37% AI users; 44.39% power users)2026Semrush consumer study1
PYMNTS Intelligence found Google is millennials' top product-discovery tool (57%), with ChatGPT second at 41%, ahead of Amazon (37%).
millennials only
Google 57%; ChatGPT 41%; Amazon 37%; YouTube 29%; Instagram 26%; Gemini 26%July 2026PYMNTS Intelligence1
Walmart found purchases completed inside ChatGPT via Instant Checkout converted about three times worse than sending shoppers to Walmart.com.
one retailer; one post generalises it to all merchants
3x lower2025-Mar 2026Walmart EVP Daniel Danker public remarks (news reports)6
Visa's Trust Index found only 23% of US consumers trust generative AI to make payments for them, rising to 61% when Visa is named as payment handler.
Visa-commissioned; branded-handler framing favours Visa
23% vs 61%2026Visa Trust Index for Agentic Commerce (2,065 US consumers)3
Ipsos found 27% of AI-aware consumers use AI for product research, but only 9% let AI make purchases autonomously.
base is AI-aware consumers, not all consumers
27% vs 9%2026Ipsos 'Shopping with AI' study1
Shopify reported orders from AI search/chat platforms up nearly 13x year over year in Q1 2026.
growth multiple off a small base; posts often drop the period
~13xQ1 2026Shopify Q1 2026 earnings call / COO / Shopify commerce blog14
Shopify reported AI-driven traffic to its merchants' stores up more than 8x year over year in Q1 2026.
small base
8xQ1 2026Shopify Q1 2026 earnings call / Shopify commerce blog9
Shopify said AI-driven traffic and orders to its merchants' stores both tripled year over year in Q2 2026.
Q1 multiples were much higher because Q1 2025 base was tiny
~3x eachQ2 2026Shopify Q2 2026 earnings call (Aug 5, 2026)6
Shopify says AI searches powered by its structured Shopify Catalog convert at 2x the rate of AI searches relying on scraped or outdated web data.
Shopify-reported; comparison set not public
2xQ1-Q2 2026Shopify (Q1 earnings, Spring '26 Edition)4
Shopify integrations account for about 35% of feed-integrated product retrievals in ChatGPT Shopping. ~35%Jul-Sep 2026Profound research2
Shopify said AI-driven orders come from new buyers at nearly twice the rate of other channels.
one post garbles as 'convert 2x'
nearly 2xQ1-Q2 2026Shopify Q2 2026 earnings call2
Shopify reported Sidekick usage up 4x year over year. 4xQ1 2026Shopify COO; Spring '26 Edition2
Shopify said daily active merchants using Sidekick grew 3.6x in Q2 2026. 3.6xQ2 2026Shopify Q2 2026 earnings1
Walmart's CEO said customers using its Sparky AI assistant spend about 40% more per order than those who don't, up from a 35% premium in Q1.
compares assistant users vs non-users; self-selection, correlation not causation
+40%Q2 FY27 (Aug 2026)Walmart Q2 FY27 earnings call (CEO John Furner)14
Walmart said Sparky users grew 70% year over year in Q2 FY27.
posts alternate 'users' and 'usage'; no absolute count
+70%Q2 FY27 (Aug 2026)Walmart Q2 FY27 earnings call11
Amazon's CEO said US customers using its AI shopping assistant spend over 40% more than non-users.
compares assistant users vs non-users; self-selection, correlation not causation
over +40%Q2 2026 earningsAmazon Q2 2026 earnings call (CEO Andy Jassy)2
Lowe's says shoppers who engage its Mylow AI assistant convert at roughly 3x the rate of those who don't; Mylow has answered 25 million questions.
compares assistant users vs non-users; self-selection, correlation not causation
~3x; 25M questions2026Lowe's SVP / company statements2
EMARKETER estimates retailer-native AI assistants (Rufus, Sparky etc.) will drive 54.1% of US AI-driven retail ecommerce sales in 2026.
forecast/estimate, not measured
54.1%2026EMARKETER forecast1
Amazon said 350 million shoppers used its AI shopping assistant in the past year. 350M12 months to Aug 2026Amazon Q2 2026 earnings call1
Target said AI-powered wish-list creation was up 50% ahead of back-to-school. +50%back-to-school 2026Target Q2 2026 earnings call1
Walmart ran 11,000+ price rollbacks in the quarter versus 7,200 the previous quarter.
not AI-specific
11,000+ vs 7,200Q2 FY27 vs Q1Walmart Q2 FY27 earnings1
OpenAI said ChatGPT Ads reached a $1 billion annualised revenue run rate in under 200 days.
run rate, not recognised revenue; one post says 'seven months'
$1B in under 200 dayslate Aug 2026OpenAI announcement5
In Q2 FY27 Walmart Marketplace grew 52%, Walmart Connect ads 43% and store-fulfilled delivery 40%; global ad revenue grew 38%.
global ad +38% figure comes from one unattributed post
+52% / +43% / +40% (global ads +38%)Q2 FY27 (Aug 2026)Walmart Q2 FY27 earnings3
eMarketer forecasts chatbots such as ChatGPT and Google AI Mode will generate under $1B in ad revenue in 2026, versus OpenAI's projected $2.5B.
forecast made mid-2026
<$1B vs $2.5B2026eMarketer forecast1
Amazon said shoppers who click a Sponsored Prompt in its AI shopping assistant convert 48% more often and spend 21% more.
clickers vs non-clickers; selection
+48% conversion; +21% spendQ2 2026Amazon Q2 2026 earnings call1
Amazon said brands using its Ads Agent saw 8% lower CPMs and 6% lower cost per acquisition.
company-reported
-8% CPM; -6% CPAQ2 2026Amazon Q2 2026 earnings call1
Amazon reported $19.8 billion in advertising revenue in Q2 2026. $19.8BQ2 2026Amazon Q2 2026 earnings1
IAB found 62% of ad buyers now use generative AI for video creative, up from 51%. 62% (from 51%)July 2026IAB1
Klaviyo opened its CRM to AI agents with 260+ MCP tools and 490+ APIs, announced around K:BOS 2026. 260+ MCP tools; 490+ APIsAug-Sep 2026Klaviyo K:BOS 2026 / MCP server release3
Klaviyo's 2026 benchmarks across 183,000+ brands: automated flows generate 41% of email revenue from 5.3% of sends, with 18x higher revenue per recipient than campaigns. 41% revenue from 5.3% sends; 18x RPR2026 benchmarksKlaviyo 2026 Email Benchmarks2
Klaviyo's Q2 2026 revenue was $371 million, up 26% year over year. $371M (+26% YoY)Q2 2026Klaviyo Q2 2026 earnings1
Klaviyo raised full-year 2026 revenue guidance to about $1.53 billion.
guidance
~$1.53BFY2026Klaviyo Q2 2026 earnings1
Klaviyo passed 205,000 customers in Q2 2026; customers above $50K ARR reached 4,477 (+36%), and non-Americas revenue grew 35%. 205,000+; 4,477 (+36%); +35%Q2 2026Klaviyo Q2 2026 earnings1
Salesforce added 60+ MCP tools on August 19, 2026. 60+Aug 19, 2026Salesforce1
Gartner's analysis of 432 customer-service AI use cases found only one in four produces positive ROI and another 42% have unclear ROI. 25% positive; 42% unclear2026Gartner analysis of 432 customer-service AI use cases3
Albert Heijn's AI produces over 1 billion automated demand forecasts daily for 17,000 products across 1,200 stores, 50 days ahead. 1B+ (17,000 products x 1,200 stores x 50 days)2026Ahold Delhaize report1
Pew Research Center finds 49% of US adults use AI chatbots, but only about 29% trust their answers. 49% vs ~29%2026Pew Research Center1
Only 20% of consumer companies have mature governance for autonomous agents (Deloitte). 20%2026Deloitte State of AI in the Consumer Industry 20261
Harris Poll's survey of 3,000 consumers found 78% assume brands pay to be recommended by AI, and 72% are comfortable with AI shopping help only if they make the final call. 78%; 72%; 45%2026The Harris Poll 'The Algorithmic Aisle'1

87 facts at "First-party / research". Lower-graded facts are cited in the chapters themselves, each carrying its grade inline.