Growthr
Resources
Book a Call
Research · Healthcare

Can AI Read Your Healthcare Site? We Scanned 100

Healthcare builds its pages more carefully than most industries. Then a fifth of them sit behind a door that does not open for machines.

Refused an AI crawler98 measured
Healthcare sites, September 2026
21/ 98
The highest rate of any industry we have scanned
Why the crawler was refused21 sites
Everything automated is challenged, not just AI13
Serves a browser, refuses an AI crawler, says nothing about it6
Blocks AI crawlers on purpose and documents it2
Dates on the article page34 articles
Date a reader can see79%
Date a parser can read50%
Shown, nothing behind it38%
Uses a time element6%

In September 2026 we scanned 100 healthcare websites and one in five would not let an AI crawler in at all. The sample covers telehealth, behavioral health, virtual care, pharmacy, diagnostics, the software hospitals run on, the big insurers, and a dozen of the best-known hospital systems. We picked the names by hand, so this is a convenience sample rather than a random one, and no company here is named as failing anything.

Each domain went through three passes. Our free scanner ran 22 checks on whether a machine can fetch and parse the site. A second pass read each sitemap. A third pulled one real article out of the sitemap and looked for a date on it.

Healthcare's problem is not how the pages are built. It is who gets to see them.

One site in five turns the AI crawlers away

Twenty-one of the 98 sites we could measure refused a request carrying an AI crawler's name. That is the highest rate of any industry we have scanned, well above the 11 we found in business software. It splits into three cases that need completely different responses, and telling them apart is most of the work.

Why the crawler was refusedSites
Everything automated is challenged, not just AI13
Serves a browser, refuses an AI crawler, says nothing about it6
Blocks AI crawlers on purpose and documents it2

The 13 in the first row are mostly hospital systems and insurers, and their security is doing what it was bought to do. A plain request with an ordinary Chrome user-agent got challenged exactly the same way. We confirmed this by hand from a second network before describing any of them. These sites may still allow verified crawlers through by IP range in ways nobody outside the organization can test, and their own logs would answer it in a minute.

The six in the second row are the ones worth acting on. A browser gets the page. Change the name on the request to GPTBot or ClaudeBot and the same URL returns a 403. Nothing in their robots.txt mentions the decision, which is usually how you can tell it was made inside a security console rather than by anyone weighing what it costs in AI search.

The last two are the honest case, and both are medical content publishers whose business is people reading their articles. They named the crawlers in robots.txt and disallowed them. Someone decided that, wrote it down, and can reverse it in one line. Whatever you think of the choice, it is a policy rather than an accident.

If you run one of these sites, the useful question is not whether your robots.txt allows GPTBot. It is whether your edge does.

The pages themselves are fine

This is the surprise. Healthcare builds pages more conservatively than any industry we have looked at except business software, and it shows. Only 23% returned a homepage with nothing on it before JavaScript ran, against 52% of online stores. Half publish sitemap dates saying what changed, roughly the industry average.

So the content is mostly reachable, mostly parseable, and mostly server-rendered. Then a fifth of it sits behind a door that does not open.

Nearly nobody can tell an engine who they are

Ninety-five of the 98 are missing pieces of the block of structured data that states the organization's contact details, address, and links to its real profiles. That is the highest rate in the series, and it matters more here than elsewhere.

An engine answering a health question is trying to work out whether the source is a real, identifiable organization. That block is how a site says, in a form a machine can check, that it is a specific hospital at a specific address with a phone number and a verified profile somewhere else. Leaving it out does not make you look untrustworthy. It makes you look like a string of text.

Related, and cheap to fix: 80% are missing at least one of an about, contact, or privacy page at a predictable URL, and 79% publish no llms.txt.

Dates are shown to people and hidden from machines

We found and fetched a real article on 34 of the 100 sites. Of those, 79% printed a date a reader could see and only 50% carried one a parser could read. Thirty-eight percent showed a date to the human and handed the machine nothing.

On the article pageShare of 34
Shows a date a reader can see79%
Carries a date a parser can read50%
Shows a date with nothing machine-readable behind it38%
Uses a time element with a datetime attribute6%

Six percent is the lowest use of that one HTML element anywhere in the series. In an industry where "last reviewed" dates are a clinical convention and often a regulatory expectation, the date is nearly always there on the page. It is just written as ordinary text, which a parser cannot tell from any other sentence.

Fixing it takes minutes. Wrap the date you already print in <time datetime="2026-09-07">, and make the datePublished and dateModified in your structured data say the same thing the page says. On a templated site that is one change to one template.

Where our checks are opinions rather than standards

Two of our results look catastrophic and are not. Every site in the sample failed to link an llms.txt file, and 97% do not serve markdown to clients that ask for it. Both are Growthr bets on where agent tooling is going, and we have written before that no major AI engine has published a ranking benefit for llms.txt. Reading those as an industry failure would be dishonest.

The checks that matter are the ones with evidence behind them: whether the crawler gets a page, whether that page has anything on it, whether it carries structure and a date, and whether other sites say the same things about you.

What we would fix first

  1. Ask your security team what it is turning away. Filter your edge logs for the GPTBot, ClaudeBot, PerplexityBot, and Google-Extended user-agents and look at the status codes you returned. Six of the sites here are almost certainly refusing crawlers without anyone having decided to.
  2. Fill in the organization block. Contact point, address, and the profile URLs that let an engine resolve you to a named institution. Ninety-five of 98 are missing part of it.
  3. Mark up the review dates you already print. A time element and matching structured data, on the template rather than per article.
  4. Publish about, contact, and privacy at predictable URLs. Four in five are missing at least one, and they are the pages an engine checks to confirm an organization is real.
  5. Add alt text. Half the sample fails it. This one is worth doing for its own sake before anyone mentions AI.

None of this decides whether a model recommends you, which is governed mostly by what other sites say about you. It decides whether you are eligible to be read at all.

Method and limits

We scanned in September 2026. The 100 domains were chosen by hand, so the sample is not random and the numbers describe these organizations rather than the industry. Each domain got a 22-check scan, a sitemap read, and one article pulled from that sitemap. The scanner ran 22 checks at the time of this scan; a 23rd, covering sitemap dates, was added afterwards, so a score you run today will not match one quoted here. An article was found for 34 of the 100, so every date percentage is out of 34 rather than 100. Two domains rate-limited the scan and are excluded from the scores, leaving 98. Every block was re-tested by hand from a second network before we described it.

The sample is named below; the results are not attributed to it. Every domain we scanned is listed, so anyone can reproduce this, but no organization is identified as passing or failing any individual check. The point of the exercise is the pattern, and a per-company scoreboard would be a pile-on rather than research.

The 100 domains we scanned

teladoc.com, amwell.com, mdlive.com, zocdoc.com, onemedical.com, carbonhealth.com, oakstreethealth.com, villagemd.com, cityblock.com, devoted.com, cloverhealth.com, hioscar.com, alignmenthealthcare.com, headspace.com, calm.com, talkspace.com, betterhelp.com, lyrahealth.com, springhealth.com, modernhealth.com, hellobrightline.com, hingehealth.com, swordhealth.com, omadahealth.com, virtahealth.com, noom.com, weightwatchers.com, joincalibrate.com, joinfound.com, mavenclinic.com, progyny.com, kindbody.com, asktia.com, joinmidi.com, thirtymadison.com, nurx.com, hellowisp.com, capsule.com, alto.com, truepill.com, goodrx.com, blinkhealth.com, costplusdrugs.com, 23andme.com, color.com, invitae.com, tempus.com, grail.com, guardanthealth.com, exactsciences.com, natera.com, illumina.com, epic.com, athenahealth.com, veradigm.com, nextgen.com, drchrono.com, tebra.com, simplepractice.com, canvasmedical.com, elationhealth.com, gethealthie.com, sprucehealth.com, zushealth.com, redoxengine.com, healthgorilla.com, particlehealth.com, healthverity.com, komodohealth.com, sharecare.com, doximity.com, medscape.com, healthgrades.com, webmd.com, mayoclinic.org, clevelandclinic.org, hopkinsmedicine.org, mountsinai.org, nyulangone.org, cedars-sinai.org, massgeneralbrigham.org, kp.org, unitedhealthgroup.com, cigna.com, aetna.com, humana.com, elevancehealth.com, centene.com, molinahealthcare.com, khealth.com, curaihealth.com, buoyhealth.com, adahealth.com, infermedica.com, included.health, transcarent.com, accolade.com, carrumhealth.com, rula.com, headway.co

Frequently asked questions

Why would a hospital website block AI crawlers?

Usually nobody decided to. Bot-management products are bought to stop scrapers and credential stuffing, and their default rules challenge anything that is not a full browser, which catches AI crawlers as a side effect. In our scan of 100 healthcare sites, 13 challenged every automated client rather than AI specifically, and six refused an AI crawler while serving a browser normally with nothing in robots.txt about it. Only two documented a deliberate policy. The distinction matters because the first case is a security setting and the second is an accident.

Do publication dates affect AI search visibility?

AI systems weight recency, so a page that cannot prove when it was written competes against pages that can. The date has to be machine-readable to count: a time element with a datetime attribute, or datePublished and dateModified in structured data, matching whatever the page shows a reader. A date rendered as plain text is invisible to a parser. Healthcare shows a date on 79% of the articles we checked and makes only 50% of them machine-readable, the widest gap in our series after crypto.

What is Organization schema and why does it matter for healthcare?

It is a block of structured data naming the organization, its contact details, its address, and links to its verified profiles elsewhere. An engine answering a health question is trying to establish whether the source is a real, identifiable institution, and that block is how a site states it in a form a machine can check. Ninety-five of the 98 healthcare sites we measured are missing part of it, the highest rate of any industry in our series.

Does a 403 to a scanner mean AI crawlers are blocked?

No. Bot-management products fingerprint the client through its TLS handshake and header order, not just its IP and user-agent, so command-line tools get challenged while browsers load the same page. The test that separates the cases is to try the same request from a second network and with a browser user-agent. If everything is challenged, that is client fingerprinting rather than an AI-crawler decision, and the site may still allowlist verified crawlers by IP range. Your CDN logs settle it.

How do I check whether AI can read my healthcare site?

Request your homepage with an AI crawler's user-agent string and confirm you get a 200 and real HTML. Then view source and check that your main claims exist in the raw HTML rather than arriving after JavaScript runs. Then check that your articles carry a date in a time element or in structured data. Our free scanner runs these checks on your domain and shows you which ones fail.

Want this run on your site?

A scored audit of your SEO and AI search standing, a prioritized fix plan, and the implementation. Flat fee.

See SEO + GEO pricing →