Can AI Read Your AI Company's Site? We Scanned 100
The companies building this technology keep the tidiest sitemaps we have measured. Ten of them still return a 403 to GPTBot, and not one wrote that down.
Two companies scored above 90 and the best managed 91.
In September 2026 we scanned 100 AI companies to see how well the industry closest to this technology does at being read by it. They came out above average and nowhere near the top. The sample covers the labs, image and video generation, applied AI in law and healthcare and support, coding tools, data labelling, inference and compute, vector databases, evaluation tooling, speech, and robotics. We picked the names by hand, so this is a convenience sample rather than a random one, and no company here is named as failing anything.
Each domain went through three passes.
- Our free scanner ran 22 checks on whether a machine can fetch and parse the site.
- A second pass read each sitemap.
- A third pulled one real article out of the sitemap and looked for a date on it.
Median score 72.5, against 65 for online retail and 77 for business software. Two companies scored above 90 and the best managed 91.
They are the best in the series at the boring parts
AI companies keep better housekeeping than any industry we have scanned. Ninety-three percent serve a sitemap and 61% put dates in it saying what changed, both the highest we have measured. Sixty-one percent of their articles carry a date a parser can read, also the highest.
| Sitemaps carry dates saying what changed | Share |
|---|---|
| AI companies | 61% |
| Fintech | 55% |
| B2B software | 50% |
| Healthcare | 48% |
| Crypto | 39% |
| Online stores | 13% |
Most of this is the stack rather than the strategy. These are recent companies on modern frameworks, and sitemaps with real dates, canonical tags, and server-rendered pages are what those frameworks do without being asked. It is the same reason the young software companies in our B2B sample outscored the enterprises.
Ten of them turn AI crawlers away
Ten of the 100 refused a request carrying an AI crawler's name. Six challenged every automated client, not AI specifically, and four served a browser normally while returning a 403 to GPTBot and ClaudeBot with nothing in robots.txt about it. None of the 100 documented a deliberate decision to block.
It would be easy and wrong to write this up as hypocrisy. A company whose crawler reads the web is not the same team as the one that bought the bot protection, and a 403 to a crawler user-agent is almost always a default rule doing what defaults do. That is the point worth taking from it: if the companies building this technology are refusing its crawlers by accident, the odds that your own edge is doing the same without anyone deciding to are higher than you would like.
Filter your edge logs for the GPTBot, ClaudeBot, PerplexityBot, and Google-Extended user-agents and look at the status codes you returned. It takes a minute and it settles the question.
One in five is blank until the browser fills it in
Twenty-one of the 100 returned a homepage with nothing on it before JavaScript ran. That is better than crypto at 34% and online retail at 52%, and worse than business software at 11%.
The pattern is specific: it is mostly the product-led companies whose marketing site is built in the same framework as the app, with a hero that animates in and copy that arrives with it. It looks excellent and reads as empty. Docs pages are usually fine, which matters, because for most of these companies the docs are what a model actually quotes.
Half of them do not publish the file written for them
Forty-seven of the 100 serve no llms.txt, and 96 do not link one from their pages.
We should be careful about how much that means, because it is our own bet. No major AI engine has published a ranking benefit for llms.txt, and we have said so in writing. Treat it as cheap protocol-layer registration, not a visibility lever. Still, it is the file proposed specifically so that language models can read a site, and the companies building language models adopt it at about the same rate as everyone else. That says more about the file's traction than about the companies.
Ninety of them cannot tell an engine who they are
Ninety of the 100 are missing pieces of the structured block that states a company's contact details, address, and links to its verified profiles. Seventy-four percent are also missing at least one of an about, contact, or privacy page at a predictable URL.
For companies in a crowded market with similar names, this is a strange thing to skip. That block is how a site says, in a form a machine can check, that it is a specific company rather than a string that resembles four other startups. It costs an afternoon.
Where our checks are opinions rather than standards
The llms.txt numbers above and the 80% that do not serve markdown to clients asking for it are both Growthr bets on where agent tooling is going, not settled standards. Reading them as an industry failure would be dishonest.
The checks that matter are the ones with evidence behind them: whether the crawler gets a page, whether that page has anything on it, whether it carries structure and a date, and whether other sites say the same things about you.
What we would fix first
- Check what your edge returns to AI crawlers. Ten companies here refuse them and none of them wrote it down, which means most of them do not know.
- Send a finished homepage. If your marketing site is built in the app framework, the copy that makes your case is probably arriving after the crawler has left.
- Fill in the company block. Contact point, address, and the profile URLs that resolve you to a company rather than a name.
- Publish about, contact, and privacy at predictable URLs. Three in four are missing at least one.
- Add alt text. Fifty-five percent fail it, and it is worth doing for accessibility before anyone mentions AI.
None of this decides whether a model recommends you, which is governed mostly by what other sites say about you. It decides whether you are eligible to be read at all.
Method and limits
We scanned in September 2026. The 100 domains were chosen by hand, so the sample is not random and the numbers describe these companies rather than the industry. Each domain got a 22-check scan, a sitemap read, and one article pulled from that sitemap. The scanner ran 22 checks at the time of this scan; a 23rd, covering sitemap dates, was added afterwards, so a score you run today will not match one quoted here. An article was found for 72 of the 100, the highest of any industry we have scanned, so every date percentage is out of 72 rather than 100. No domain rate-limited the scan, so all 100 are scored. Every block was re-tested by hand from a second network before we described it.
The sample is named below; the results are not attributed to it. Every domain we scanned is listed, so anyone can reproduce this, but no company is identified as passing or failing any individual check. The point of the exercise is the pattern, and a per-company scoreboard would be a pile-on rather than research.
openai.com, anthropic.com, mistral.ai, cohere.com, ai21.com, stability.ai, midjourney.com, runwayml.com, lumalabs.ai, elevenlabs.io, suno.com, synthesia.io, heygen.com, descript.com, opus.pro, perplexity.ai, you.com, glean.com, hebbia.ai, harvey.ai, abridge.com, sierra.ai, decagon.com, cresta.com, observe.ai, gong.io, clari.com, jasper.ai, copy.ai, writer.com, grammarly.com, typeface.ai, pika.art, gamma.app, beautiful.ai, replit.com, cursor.com, windsurf.com, tabnine.com, sourcegraph.com, magic.dev, cognition.ai, poolside.ai, augmentcode.com, huggingface.co, wandb.ai, scale.com, labelbox.com, snorkel.ai, v7labs.com, roboflow.com, together.ai, fireworks.ai, replicate.com, modal.com, baseten.co, anyscale.com, runpod.io, lambdalabs.com, coreweave.com, crusoe.ai, groq.com, cerebras.ai, sambanova.ai, graphcore.ai, tenstorrent.com, langchain.com, llamaindex.ai, pinecone.io, weaviate.io, trychroma.com, qdrant.tech, zilliz.com, vespa.ai, arize.com, whylabs.ai, galileo.ai, humanloop.com, braintrust.dev, langfuse.com, deepgram.com, assemblyai.com, speechmatics.com, rev.com, otter.ai, fireflies.ai, granola.ai, lindy.ai, character.ai, inflection.ai, x.ai, reka.ai, contextual.ai, worldlabs.ai, skild.ai, figure.ai, 1x.tech, waymo.com, nuro.ai, wayve.ai
Frequently asked questions
Do AI companies block AI crawlers?
Ten of the 100 AI companies we scanned refused a request carrying an AI crawler's name. Six challenged every automated client rather than AI specifically, and four served a browser normally while returning a 403 to GPTBot and ClaudeBot with nothing in robots.txt about it. None documented a deliberate decision. This is almost always a default bot-management rule rather than a policy, which is the useful lesson: if it happens here, it can happen on your site without anyone choosing it.
Should my AI startup publish an llms.txt file?
It costs an hour and no major AI engine has published a ranking benefit for it, so treat it as cheap protocol-layer registration rather than a visibility lever. Forty-seven of the 100 AI companies we scanned do not have one. Crawlability, server-rendered content, structured data, and machine-readable dates all matter more and are all better evidenced. We wrote up the evidence in Does llms.txt actually work?
Why is my marketing site invisible to AI when my docs are fine?
Because they are usually built differently. Documentation is typically generated into static HTML, so a crawler gets the finished page. Marketing sites are often built in the same framework as the product, with copy that arrives after the browser runs the site's code, and most AI crawlers do not run it. Twenty-one of the 100 AI companies we scanned returned a homepage with no real content in it while their docs were readable.
Does a 403 to a scanner mean AI crawlers are blocked?
No. Bot-management products fingerprint the client through its TLS handshake and header order, not just its IP and user-agent, so command-line tools get challenged while browsers load the same page. The test that separates the cases is to try the same request from a second network and with a browser user-agent. If everything is challenged, that is client fingerprinting rather than an AI-crawler decision, and the site may still allowlist verified crawlers by IP range. Your CDN logs settle it.
How do I check whether AI can read my site?
Request your homepage with an AI crawler's user-agent string and confirm you get a 200 and real HTML. Then view source and check that your main claims exist in the raw HTML rather than arriving after JavaScript runs. Then check that your posts carry a date in a time element or in structured data. Our free scanner runs these checks on your domain and shows you which ones fail.
Want this run on your site?
A scored audit of your SEO and AI search standing, a prioritized fix plan, and the implementation. Flat fee.
See SEO + GEO pricing →© 2026 Growthr. All rights reserved. · llms.txt · About · Privacy · Terms