Can AI Read Your SaaS Site? We Scanned 140
Business software is the healthiest industry we have scanned. Then we ran the same checks over 40 startups, and the startups scored higher.
In September 2026 we scanned 100 B2B software companies, then scanned 40 much smaller ones to test an assumption, and the small companies won. The main sample covers work management, developer tools, data and analytics, infrastructure, customer support, HR and payroll, security, and content platforms. We picked the names by hand, so this is a convenience sample rather than a random one, and no company here is named as failing anything.
Each domain went through three passes. Our free scanner ran 22 checks on whether a machine can fetch and parse the site. A second pass read each sitemap. A third pulled one real article out of the sitemap and looked for a date on it.
Business software is the healthiest industry we have measured. Median 77, against 72 for fintech and 65 for online retail. Four percent scored above 90, which is four percent more than fintech, healthcare, or retail managed.
The startups beat the incumbents
We expected the opposite. Big companies have SEO teams, budgets, and agencies, and small ones have a founder doing marketing at midnight. So we scanned 40 seed-to-Series-B and independent software companies with the same 22 checks.
| Cohort | Median score |
|---|---|
| Small and young (40 companies) | 78.5 |
| Large and established (99 companies) | 77 |
The gap widens at the top: 10% of the small companies scored above 90, against 4% of the large ones.
The reason is not effort, it is age. A company that built its site in the last three years is on a framework that emits server-rendered pages, canonical tags, sitemaps with dates, and valid structured data as defaults, because that is what the framework does when you do nothing. A company that built its site in 2016 is on a content management system that predates all of it, wrapped in years of tag managers and redirects, behind a security appliance somebody configured once.
Which is worth saying plainly: if your site scores badly, that is mostly a fact about when it was built, not about how much you care. It also means the fix is usually replacing a layer rather than persuading anyone.
Only 11 in 100 are blank without JavaScript
Eleven of the 99 sites we could measure returned a homepage with nothing on it before JavaScript ran. That is the best rate in the series by a distance: crypto is 34%, online stores 52%.
Software companies sell to people who read documentation, so their sites are built to be read, and the docs and blog are usually static HTML. This is the check that matters most, it carries the heaviest weight we assign, and this industry mostly passes it.
They publish more than anyone and date it less
Here is the one place business software does worse than everybody. We found and fetched a real article on 62 of the 100 sites, and only 65% of those printed a date a reader could see. That is the lowest of any industry except online retail, and retail has the excuse of being shops rather than publishers.
| On the article page | Share of 62 |
|---|---|
| Shows a date a reader can see | 65% |
| Carries a date a parser can read | 44% |
| Shows a date with nothing machine-readable behind it | 26% |
Uses a time element with a datetime attribute | 13% |
Undated content is a deliberate choice in a lot of software marketing. It keeps a guide from looking stale and saves the work of maintaining it. The trade is that a model weighing two guides has no way to prefer yours, and "no date" does not read as timeless to a machine, it reads as unknown.
If a page is genuinely evergreen, give it a dateModified and update it when you touch it. That is the version of undated that still competes.
Eleven refuse the AI crawlers
Eleven of the 99 turned away a request carrying an AI crawler's name. Six challenged every automated client rather than AI specifically, and five served a browser normally while returning a 403 to GPTBot and ClaudeBot with nothing in robots.txt about it. None of the 100 documented a deliberate policy.
Challenged everything
- Sites
- 6 of 99
- Browser
- Also challenged
- Reading
- Client fingerprinting
Refused AI only
- Sites
- 5 of 99
- Browser
- Served normally
- robots.txt
- Says nothing
Several of the five sell developer infrastructure, which is the detail worth sitting with. These are companies whose customers integrate with them programmatically, refusing a documented crawler by accident. If it can happen there, the odds your own edge is doing it are not small.
Where our checks are opinions rather than standards
Ninety-four percent do not link an llms.txt file and 73% do not serve markdown to clients asking for it. Both are Growthr bets on where agent tooling is going, and we have written before that no major AI engine has published a ranking benefit for llms.txt. Reading those as failures would be dishonest.
The checks that matter are the ones with evidence behind them: whether the crawler gets a page, whether that page has anything on it, whether it carries structure and a date, and whether other sites say the same things about you.
What we would fix first
- Date your content. The clearest gap in this industry. A visible date and matching structured data, on the template rather than per post.
- Check what your edge returns to AI crawlers. Filter your logs for the GPTBot, ClaudeBot, PerplexityBot, and Google-Extended user-agents and look at the status codes.
- Fill in the company block. Eighty-five percent are missing contact details, address, or the profile links that resolve you to a company.
- Publish about, contact, and privacy at predictable URLs. Fifty-eight percent are missing at least one.
- Add alt text. Forty-nine percent fail it, and it is worth doing for accessibility before anyone mentions AI.
None of this decides whether a model recommends you, which is governed mostly by what other sites say about you. It decides whether you are eligible to be read at all.
Method and limits
We scanned in September 2026. The 100 main domains and 40 comparison domains were chosen by hand, so neither sample is random and the numbers describe these companies rather than the industry. Each domain got a 22-check scan, a sitemap read, and one article pulled from that sitemap. The scanner ran 22 checks at the time of this scan; a 23rd, covering sitemap dates, was added afterwards, so a score you run today will not match one quoted here. An article was found for 62 of the 100, so every date percentage is out of 62 rather than 100. One domain rate-limited the scan and is excluded from the scores, leaving 99. The 40-company comparison cohort got the same 22-check scan and no sitemap or date probe. Every block was re-tested by hand from a second network before we described it. Three of these domains also appear in our fintech sample, because they are both, so the two studies are not independent samples.
The sample is named below; the results are not attributed to it. Every domain in the main sample is listed, so anyone can reproduce this, but no company is identified as passing or failing any individual check. The point of the exercise is the pattern, and a per-company scoreboard would be a pile-on rather than research.
The 100 domains we scanned
salesforce.com, hubspot.com, zendesk.com, atlassian.com, slack.com, notion.so, airtable.com, asana.com, monday.com, clickup.com, smartsheet.com, figma.com, miro.com, loom.com, dropbox.com, box.com, docusign.com, zoom.us, calendly.com, gitlab.com, github.com, jetbrains.com, circleci.com, harness.io, launchdarkly.com, datadoghq.com, newrelic.com, sentry.io, pagerduty.com, grafana.com, splunk.com, elastic.co, sumologic.com, snowflake.com, databricks.com, mongodb.com, cockroachlabs.com, redis.io, confluent.io, fivetran.com, getdbt.com, airbyte.com, segment.com, amplitude.com, mixpanel.com, heap.io, looker.com, tableau.com, sigmacomputing.com, hex.tech, retool.com, vercel.com, netlify.com, render.com, digitalocean.com, cloudflare.com, fastly.com, twilio.com, sendgrid.com, postmarkapp.com, mailchimp.com, klaviyo.com, braze.com, iterable.com, customer.io, intercom.com, front.com, helpscout.com, freshworks.com, servicenow.com, workday.com, gusto.com, rippling.com, deel.com, remote.com, lattice.com, cultureamp.com, greenhouse.io, lever.co, ashbyhq.com, checkr.com, okta.com, auth0.com, 1password.com, crowdstrike.com, snyk.io, vanta.com, drata.com, secureframe.com, tailscale.com, postman.com, algolia.com, contentful.com, sanity.io, storyblok.com, webflow.com, zapier.com, make.com, pendo.io, userpilot.com
Frequently asked questions
Do bigger companies have better SEO than startups?
Not in our scan. We ran the same 22 checks over 99 large, established software companies and 40 seed-to-Series-B and independent ones. The small companies had a median score of 78.5 against 77, and 10% of them scored above 90 against 4% of the large ones. The reason appears to be the age of the stack rather than the effort: a site built in the last three years gets server-rendered pages, sitemaps with dates, and valid structured data as framework defaults.
Should blog posts have visible dates?
Yes, and a machine-readable one alongside it. Undated content is common in software marketing because it keeps a guide from looking stale, but a model weighing two guides has no way to prefer yours, and an absent date reads as unknown rather than timeless. Only 65% of the software articles we checked showed a reader a date, the lowest of any industry except retail. If a page really is evergreen, give it a dateModified and update it whenever you touch the page.
Does a 403 to a scanner mean AI crawlers are blocked?
No. Bot-management products fingerprint the client through its TLS handshake and header order, not just its IP and user-agent, so command-line tools get challenged while browsers load the same page. The test that separates the cases is to try the same request from a second network and with a browser user-agent. If everything is challenged, that is client fingerprinting rather than an AI-crawler decision, and the site may still allowlist verified crawlers by IP range. Your CDN logs settle it.
What is the single highest-impact fix for AI search visibility?
Serving real HTML. The heaviest check we run asks whether the homepage carries content before JavaScript runs, because a page that is empty until it hydrates is a page most retrieval pipelines never see. Business software passes this more than any industry we have scanned, at 11% failing, so for most software companies the next most valuable fix is dating their content and completing the organization block.
How do I check whether AI can read my SaaS site?
Request your homepage with an AI crawler's user-agent string and confirm you get a 200 and real HTML. Then view source and check that your main claims exist in the raw HTML rather than arriving after JavaScript runs. Then check that your posts carry a date in a time element or in structured data. Our free scanner runs these checks on your domain and shows you which ones fail.
Want this run on your site?
A scored audit of your SEO and AI search standing, a prioritized fix plan, and the implementation. Flat fee.
See SEO + GEO pricing →© 2026 Growthr. All rights reserved. · llms.txt · About · Privacy · Terms