Skip to content

Research · 2026-08-03

88.8% of specialty retailers let AI crawlers in. Most still can't be read.

We crawled 138 specialty retailers expecting to find doors closed to AI assistants. Permission turned out to be near-universal — and 32.6% of catalogs were still unreadable to an agent the site didn't recognise. 10 stores invite the named crawlers in writing and refuse an unrecognised one at the door.

88.8%
Admit every named crawler
87 of 98 stores publishing a reachable robots.txt allow all five.
22
Refused our crawler
Returned 403 or 429 to an unrecognised agent — 10 of them while inviting the named five in robots.txt.
33.3%
Median GTIN / MPN coverage
Among catalogs we could read, the share of products carrying an identifier an assistant can match on.

The finding

We expected to find retailers shutting AI crawlers out, and platform choice deciding who gets read. Neither is what the data shows.

Across 138 specialty retailers in consignment, dive, music, pet, quilting, permission is close to universal. Of the 98 stores whose robots.txt we could fetch, 87 — 88.8% — admit every one of the five AI crawlers we checked. 11 of 98 restrict at least one named crawler. The most-restricted crawler is `Google-Extended`, permitted by 89.8% of them.

Stores whose robots.txt permits each crawlerClaudeBot91.8%GPTBot91.8%Google-Extended89.8%OAI-SearchBot98%PerplexityBot99%
Share of stores whose robots.txt permits each named crawler, among the 98 stores publishing a reachable robots.txt.

What we did find is a gap between stated policy and actual behaviour. 22 of 138 stores (15.9%) refused our own crawler outright at the edge — HTTP 403 or 429 — and 10 of those same stores explicitly welcome the five named crawlers in the robots.txt they publish.

Access is granted by identity, not by policy. The written rule is open. The door opens for agents the edge recognises. An assistant with a known crawler walks in; anything else is turned away before it reads a single product.

That has a direct consequence for this study, and we think it is the more useful result: we could enumerate a product catalog for only 45 of 138 stores (32.6%). Everything below about product data describes that subset, and that subset is defined by who let us in.

What the readable catalogs look like

Among the 45 catalogs we could read, structured markup is close to solved and identifiers are not. Median JSON-LD Product coverage is 100.0%. Median GTIN or MPN presence is 33.3%.

The distinction matters because they answer different questions. Markup tells an assistant this page is a product. An identifier tells it which product — the thing it needs before it will confidently put a specific item in front of a buyer.

Median agent-readiness score, stores we could readshopify (n=36)79%closed (n=2)75.5%custom (n=7)87%unknown (n=0)not measured
Median agent-readiness score by platform group, over stores whose catalog we could read. Groups are sized very differently and the readable set is not random — see the limits below before reading anything into the gap.

Manifests at a well-known path

43 of the 113 stores we could check (38.1%) serve a UCP manifest at /.well-known/ucp. 35 of those are Shopify stores, where the platform serves the file on the merchant's behalf.

That is the entire claim: a file exists at a path. Whether any of these merchants participates in a given assistant's commerce programme is arranged off the public web, and a crawl cannot observe it in either direction. We record presence and stop there.

What the assistants actually answered

The 9 pairs whose catalogs we could both read were matched on category and catalog size, then asked one shared panel of buyer questions — 20 runs per store, with every store's own name and brands banned from the query text. Both members are scored against the same answers, so neither can win a question by being named in it.

PairCategoryShopifyOther platform Delta
dive-1dive 0% 20% (closed) -20 pts
dive-2dive 40% 0% (closed) +40 pts
music-3music 70% 15% (custom) +55 pts
music-4music 30% 0% (custom) +30 pts
music-5music 75% 10% (custom) +65 pts
music-6music 20% 0% (custom) +20 pts
music-7music 5% 10% (custom) -5 pts
quilting-8quilting 20% 0% (custom) +20 pts
quilting-9quilting 0% 25% (custom) -25 pts

The Shopify member surfaced more often in 6 of 9 pairs, the non-Shopify member in 3.

We are not calling that a platform effect, and neither should you. Two reasons. 9 pairs on one surface at one point in time is a count of what happened in these matchups, not evidence of a cause. And the two sides were not read equally deeply: Shopify catalogs are read through a documented endpoint and returned 120 products each, while the other members were read by sampling product pages and yielded 11–25. The shared panel splits its slots evenly between members, but the Shopify member's questions are drawn from a much deeper catalog, which plausibly makes them more specific and easier to answer. That is a property of how much we could read — the same access asymmetry this study is about — not a property of the store. The comparison needs re-running with both sides capped at equal depth before the split means anything.

Method

Two dimensions, deliberately separated by what they cost.

Agent-readiness is a crawl, so it runs across the whole corpus. For each store we parse robots.txt per named agent, sample product pages for JSON-LD Product markup, measure GTIN/MPN and brand presence across the sampled catalog, detect sitemap and product feed, and check the well-known path. Catalogs are read through documented endpoints and public pages at one to two requests per second, identifying ourselves honestly as Chorrus-Audit/0.1.

Discoverability asks assistants real buyer questions and records whether a store surfaces in the answer. Each matched pair answers one shared query panel built from the union of both catalogs, with every store's own name and brands banned from the query text, so neither store can win a question by being named in it.

Studied stores are pseudonymous throughout, in the write-up and in the dataset. Retailers that win an answer are named when we report answer results, because that is an observation about an answer rather than a verdict on a merchant.

Discovery is the subject: whether a store appears in an answer at all. Nothing here concerns how a purchase is completed.

Readiness by group

GroupStoresRefused our crawlerCatalogs read Median readinessManifest served
shopify400 36 79.0 35
closed80 2 75.5 6
custom390 7 87.0 0
unknown5122 0 2

unknown is a measurement outcome — the platform could not be fingerprinted because the store did not answer us — not a kind of shop.

Limits

  • Selection by access. The 45 readable catalogs are the stores that let an unrecognised crawler in. Every product-data figure describes that subset only, and it is not a random sample.
  • Our crawler is not an assistant. We fetch as Chorrus-Audit/0.1. A firewall that refuses us may well admit GPTBot — which those same robots.txt files permit. "We could not read it" is a fact about our access, never a claim about what an assistant sees.
  • Two different failures look alike. Of the 93 catalogs we could not read, 22 refused us and 71 answered but exposed no catalog we could enumerate — after following the sitemaps robots.txt declares as well as the conventional paths. We did not diagnose that group store by store, so treat it as unexplained rather than as a fault of the merchant.
  • Uneven groups. shopify 40, closed 8, custom 39, unknown 51. Comparisons across groups this unbalanced are directional at best.
  • Size proxy is catalog size. Pairing uses sampled catalog size because review and follower counts are not observable from a crawl. A big catalog and a big audience are different things.
  • Brands and retailers are mixed. Some stores sell only their own products; others carry many. An assistant naming a brand's own site is a different event from naming a multi-brand retailer.
  • n = 138, 5 categories, one surface, one point in time. These are descriptive counts, not a significance test, and assistant answers move day to day.

How does your store read?

The same checks, run on your catalog: crawler access, structured data, identifiers, and what assistants answer when buyers ask.

Run your free audit