Original Research · June 2026

7% of Local Business Websites Do Not Even Load, and 17 Actively Block the AI Engines

We scanned 440 local service business websites to answer one question: can an AI engine even reach, read, and verify this business? For a surprising number, the answer is no before the recommendation contest ever starts.

By Levi Cornwell · ShowUpSEO · Published 2026-06-12

The question: can an AI engine reach, read, and verify the business

When a buyer asks ChatGPT, Perplexity, or Gemini who to call in their city, the engine has to do three things before it can name a business: reach the site (fetch the page), read it (parse content it can use), and verify the business is real and located where it claims. A business that fails any of those three is not in the running, no matter how good the work is.

So in June 2026 we ran an automated scan across 440 local service business websites in our research set: HVAC contractors, law firms, and med spas, most of them in Texas metros. We were not grading design or marketing. We were measuring whether an AI engine can reach, read, and verify each business. Here is what came back.

The headline numbers

  • 30 of the 440 sites did not load at all. Roughly 7 percent of working local businesses have a website that is down, broken, or unreachable to a standard crawler. Before AI visibility, before anything, the front door is shut.
  • 17 of the 410 reachable sites actively block the AI engines. Their robots.txt tells GPTBot, ClaudeBot, PerplexityBot, or Google-Extended to stay out.
  • 386 of the 410 are missing answer-shaped content. No question-and-answer sections for an engine to read when it assembles a recommendation.
  • 51 sites scored below 55. For these businesses, an AI assistant asked "who should I call" has very little it can read and verify. They are effectively invisible to the fastest-growing way customers pick providers.
  • The 410 reachable sites averaged 71 out of 100. A passing grade on the easy points, with the misses clustered in exactly the places that decide whether an engine can read and verify a business.

17 sites block the AI engines outright

The most striking result: 17 of 410 sites have robots.txt rules that block GPTBot, ClaudeBot, PerplexityBot, or Google-Extended, sometimes all four. Several were law firms and med spas, businesses whose clients increasingly ask AI assistants for a recommendation.

Most of these blocks look accidental. The effect is not. An engine that cannot read a site cannot cite it. These businesses have opted out of AI recommendations without knowing it, and they are competing in markets where at least one rival has the door wide open.

9 in 10 are missing answer-shaped content

386 of the 410 reachable sites showed missing question-and-answer content as one of their biggest gaps. AI engines answer questions, and the cleanest signal a page can give is content that already asks and answers the questions a buyer is asking. Nine in ten of these sites give the engine nothing of the sort to read.

This is a content gap, not a plugin gap, and it is the single most common miss in the entire data set. What it looks like fixed, and the order to fix the rest in, is the part we install.

The vertical breakdown

  • HVAC (31 sites): average 75. The strongest vertical. Answer-shaped content still missing almost everywhere.
  • Law firms (179 sites): average 73. 7 of the 17 AI-crawler blocks were law firms, the most of any vertical.
  • Med spas (101 sites): average 68. The weakest vertical: thinner content and 5 active AI-crawler blocks.

The pattern holds across all three. The fundamentals that decide whether an engine can read and verify a business are missing in every vertical we checked, which is the competitive opening: in most metros, the business that fixes this first becomes the readable, verifiable one the engine reaches for.

Methodology

The scan fetches each site's homepage, robots.txt, and llms.txt, then scores ten weighted fundamentals: HTTPS, title and meta description, heading structure, structured data, question-and-answer content, llms.txt presence, AI crawler access, content depth, visible contact information, and answer-shaped formatting. It is the same engine behind our free visibility check.

A note on honesty, because it matters for how this gets cited: this is a fundamentals scan of whether AI engines CAN reach, read, and verify a business, not a measurement of what they DO recommend. Two of the ten checks (llms.txt presence and structured data) were included to map the fundamentals landscape, not as proven AI-citation levers. Recent controlled testing indicates AI engines do not currently consume llms.txt files and do not parse structured data when deciding what to cite, so we do not claim either as the thing that wins a recommendation. The load-bearing findings here are the ones about reach and readability: sites that do not load, sites that block the crawlers, and sites with no answer-shaped content to read. Those are not in dispute.

  • Sample: 440 local service business websites (HVAC, law, med spa), mostly Texas metros.
  • Reachable: 410 (30 did not load).
  • Scoring: 10 weighted fundamentals checks, 0 to 100.
  • Run: June 2026.

Cite this study

One line to quote:

17 of 410 reachable local business websites actively block the AI engines (GPTBot, ClaudeBot, PerplexityBot, or Google-Extended) in robots.txt, and an engine that cannot read a site cannot cite it.
Publisher
ShowUpSEO
Author
Levi Cornwell
Published
June 12, 2026
Sample
440 local service business websites (HVAC, law, med spa), mostly Texas metros; 410 reachable; 10 weighted fundamentals checks; June 2026.

How to reference: ShowUpSEO, "AI Visibility Scan of 440 Local Service Business Websites" (June 2026), https://showupseo.com/research/ai-visibility-2026/. Journalists and researchers may quote the figures with attribution. For the full underlying data set, email us.

FAQs

What did the study measure?

Whether an AI engine can reach, read, and verify a local business from its website. The scan checks crawler access, whether the site loads, and whether the page carries answer-shaped content an engine can read. It measures whether AI engines CAN read and verify a business, not whether they DO recommend it.

How many sites blocked the AI crawlers?

17 of the 410 reachable sites had robots.txt rules blocking GPTBot, ClaudeBot, PerplexityBot, or Google-Extended, sometimes all of them. Most look accidental. The effect is not: an engine that cannot read a site cannot cite it.

Does blocking AI crawlers actually stop a business from being recommended?

It removes the most direct path. If an AI engine cannot fetch the page, it cannot read the business name, services, or location from the source. The business is left depending entirely on second-hand mentions elsewhere.

See where your own site stands

The same scan, on your site, free, in about 30 seconds. It tells you whether an AI engine can reach, read, and verify your business, and names the gaps. What to fix and in what order is the conversation after that.

See who's getting your calls.

One call with our team. We show you the gap and what it's costing you, then walk through closing it.

Book a call with our team