Growthr
Resources
Book a Call
Research · Online retail

Can AI Read Your Online Store? We Scanned 100

We asked 100 online stores for a page the way an AI does. Half of them sent back nothing. The rest mostly cannot say when anything last changed.

The 100 domains we scanned
Blank without JavaScriptshare of sample
Handed back an empty page
Stores we could measure
50/ 97
Three domains rate-limited the scan

In September 2026 we asked 100 online stores for their homepage the way an AI does, and half of them handed back a page with nothing on it. The sample runs from Nike and Sephora down to single-product DTC brands, across apparel, beauty, home, food and drink, supplements, and the big-box retailers. We picked the names by hand, so this is a convenience sample rather than a random one, and no company here is named as failing anything.

Each store went through three passes. Our free scanner ran 22 checks on whether a machine can fetch and parse the site. A second pass read each sitemap. A third pulled one real article out of the sitemap, a blog post or a guide rather than a product page, and looked for a date on it.

Retail scored worse than any other industry we have run this on.

Half of these stores look empty to AI

Fifty of the 97 stores we could measure handed back a page with nothing on it. Open the same URL in Chrome and it looks fine, because the photos and the prices load a second later once your browser runs the store's code. Most AI crawlers never run that code, so the empty version is what they keep.

This is the heaviest check we run, and it is the one that decides whether anything else matters. A store that answers with an empty page is not competing badly in AI search. It is absent from it.

Blank without JavaScriptShare of the sample
Online stores52%
Crypto34%
Healthcare23%
Fintech24%
B2B software11%

The fix is not a rewrite. Most storefront platforms can render the important pages on the server and send finished HTML, and on a hosted platform it is a setting rather than a project. The pages worth doing first are the ones that answer a buying question: category pages, your best-selling products, and anything a person would ask an assistant about.

Almost nobody tells crawlers when anything changed

A sitemap is the list of pages you want found. Each entry can carry the date that page last changed, which is how a crawler decides what to come back and re-read instead of skipping.

Thirteen of the 100 stores publish those dates. That is the lowest of any industry we have measured, and it is a long way below the 61% we found among AI companies. Another 34 serve no sitemap at all.

SitemapsShare of 100
Serve a sitemap at the standard location66%
Carry dates saying what changed13%

So a crawler visiting most of these stores has no way to tell a product added this morning from one that has sat there for three years. It has to guess, and guessing usually means not coming back. If you add products or publish posts regularly, this is the cheapest thing on the list to fix, because the platform already knows the dates and simply is not publishing them.

One store in five turns away anything that is not a person

Nineteen of the 97 refused us entirely. Not the AI crawler specifically, everything: a plain request with an ordinary Chrome user-agent was challenged the same way. We re-tested these by hand from a second network before describing them, and some dropped the connection outright while others sent us round a redirect loop that never ends.

These are almost all large apparel and outdoor brands, and it is worth being careful about what it means. Their bot protection is doing what it was bought to do, which is stop scrapers and sneaker bots. Verified AI crawlers may still be allowed through by IP range in ways nobody outside the company can test. But it does mean the answer to "can an AI read our site" is sitting in a security console, not a marketing one.

Separately, eight stores served a browser normally and returned a 403 to an AI crawler, with nothing in their robots.txt mentioning it. None of the 100 documented a deliberate decision to block AI crawlers. That gap is the tell: a real policy gets written down, and these were not.

The one thing retail does better than everyone else

We expected stores to be the worst at dates too, and they were the best. Of the 52 stores where we found a real article, 62% carried a date a parser could read, the highest of any industry we have scanned. Only 56% showed the reader a date at all, which is the lowest.

That inversion is the platform doing it, not the marketer. Hosted store software emits the publication date into the page's structured data automatically, whether or not the theme prints it on screen. Everywhere else we looked, the reverse was true: sites show a date to people and hand the machine nothing. Crypto sites do this on 44% of their articles, healthcare on 38%.

If you run a store, this is the rare case where the default is on your side. Check that it is actually on, and then leave it alone.

Where our checks are opinions rather than standards

Two of our results look catastrophic and are not. Every store in the sample failed to link an llms.txt file, and 55% do not serve markdown to clients that ask for it. Both are Growthr bets on where agent tooling is going, and we have written before that no major AI engine has published a ranking benefit for llms.txt. Reading a 100% failure rate on that as an industry problem would be dishonest.

The checks that matter are the ones with evidence behind them: whether the crawler gets a page, whether that page has anything on it, whether it carries structure and a date, and whether other sites say the same things about you.

What we would fix first

  1. Send finished pages, not empty ones. If your category and product pages are blank until the browser fills them in, nothing else on this list will help. Start with the pages that answer a buying question.
  2. Ask your security vendor what it is turning away. Filter your edge logs for the GPTBot, ClaudeBot, PerplexityBot, and Google-Extended user-agents and look at what you returned. A fifth of the stores here are refusing far more than scrapers.
  3. Turn on sitemap dates. Your platform already knows when each page changed. Publishing that is usually a setting, and 87 of these 100 stores have not flipped it.
  4. Fill in the company block. Contact details, an address, and links to your real profiles are what let an engine resolve you to a company rather than a brand name it half-recognizes. Ninety-one percent of these stores are missing pieces of it.
  5. Give the reader the date too. Your structured data probably has it. Printing it on the page is a trust signal for the human, and costs one line in the template.

None of this decides whether an assistant recommends your product over a competitor's, which is governed mostly by what other sites say about you. It decides whether you are eligible to be considered at all.

Method and limits

We scanned in September 2026. The 100 domains were chosen by hand, so the sample is not random and the numbers describe these stores rather than all of retail. Each domain got a 22-check scan, a sitemap read, and one article pulled from that sitemap. The scanner ran 22 checks at the time of this scan; a 23rd, covering sitemap dates, was added afterwards, so a score you run today will not match one quoted here. An article was found for 52 of the 100, so every date percentage is out of 52 rather than 100. Three domains rate-limited the scan and are excluded from the scores, leaving 97. Every block was re-tested by hand from a second network before we described it.

The sample is named below; the results are not attributed to it. Every domain we scanned is listed, so anyone can reproduce this, but no company is identified as passing or failing any individual check. The point of the exercise is the pattern, and a per-company scoreboard would be a pile-on rather than research.

The 100 domains we scanned

allbirds.com, warbyparker.com, casper.com, purple.com, glossier.com, harrys.com, dollarshaveclub.com, getquip.com, bombas.com, mackweldon.com, everlane.com, rothys.com, awaytravel.com, brooklinen.com, parachutehome.com, burrow.com, article.com, thuma.co, madeincookware.com, fromourplace.com, carawayhome.com, materialkitchen.com, hydroflask.com, yeti.com, stanley1913.com, owalalife.com, liquiddeath.com, athleticbrewing.com, drinkolipop.com, drinkpoppi.com, drinkspindrift.com, magicspoon.com, rxbar.com, kodiakcakes.com, chomps.com, seed.com, ritual.com, drinkag1.com, drinklmnt.com, bulletproof.com, skims.com, aloyoga.com, vuoriclothing.com, gymshark.com, lululemon.com, outdoorvoices.com, tracksmith.com, on.com, hoka.com, brooksrunning.com, newbalance.com, nike.com, adidas.com, underarmour.com, patagonia.com, thenorthface.com, columbia.com, rei.com, backcountry.com, huckberry.com, revolve.com, ssense.com, farfetch.com, mrporter.com, asos.com, zappos.com, etsy.com, chewy.com, wayfair.com, target.com, walmart.com, costco.com, bestbuy.com, homedepot.com, lowes.com, sephora.com, ulta.com, iliabeauty.com, tatcha.com, drunkelephant.com, deciem.com, fentybeauty.com, rarebeauty.com, kyliecosmetics.com, olaplex.com, functionofbeauty.com, prose.com, curology.com, saie.com, meritbeauty.com, summerfridays.com, youthtothepeople.com, supergoop.com, beautycounter.com, nativecos.com, hellotushy.com, bluebottlecoffee.com, graza.co, deathwishcoffee.com, chamberlaincoffee.com

Frequently asked questions

Why can't AI read my Shopify store?

Most storefronts build the page in the visitor's browser rather than sending a finished page from the server. A person never notices, because their browser does the work in a fraction of a second. Most AI crawlers do not run that code, so they receive the empty shell that arrives first. In our scan of 100 online stores, 52% returned a page with no real content in it. The fix is server-side rendering for the pages that answer a buying question, which on a hosted platform is usually a setting rather than a rebuild.

Does my online store need a sitemap with lastmod dates?

It helps and it is nearly free. lastmod is an optional date on each sitemap entry saying when that page last changed, and crawlers use it to decide what deserves a re-fetch. Only 13 of the 100 stores we scanned publish those dates, the lowest of any industry we have measured. Your platform already tracks the dates, so this is usually a setting rather than development work. It only helps if it stays true, which means emitting it from the platform rather than hardcoding a date.

Is my bot protection blocking AI crawlers?

Quite possibly, and your robots.txt will not tell you. Nineteen of the 97 stores we could measure refused every automated client, not just AI crawlers, and eight more returned a 403 to an AI crawler while serving a browser normally. None of them documented the decision anywhere. The answer is in your CDN logs: filter for the GPTBot, ClaudeBot, PerplexityBot, and Google-Extended user-agents and look at the status codes you returned. That takes a minute and settles it.

Does a 403 to a scanner mean AI crawlers are blocked?

No. Bot-management products fingerprint the client through its TLS handshake and header order, not just its IP and user-agent, so command-line tools get challenged while browsers load the same page. The test that separates the cases is to try the same request from a second network and with a browser user-agent. If everything is challenged, that is client fingerprinting rather than an AI-crawler decision, and the site may still allowlist verified crawlers by IP range. Your CDN logs settle it.

How do I check whether AI can read my store?

Request your homepage with an AI crawler's user-agent string and confirm you get a 200 and real HTML. Then view source and check that your product names and prices exist in the raw HTML rather than arriving after JavaScript runs. Then check that your sitemap carries lastmod dates. Our free scanner runs these checks on your domain and shows you which ones fail.

Want this run on your site?

A scored audit of your SEO and AI search standing, a prioritized fix plan, and the implementation. Flat fee.

See SEO + GEO pricing →