When businesses think about AI visibility, it’s usually about content and optimisation. And rightly so. But before investing time and money in either, there’s something worth checking: can AI search services actually access the site? It's like a shop with the lights on but the door locked.

That doesn’t mean every AI crawler should be allowed through. Some collect content for model training, while others help AI search tools find and surface information. Before changing anything, it’s worth checking who you’re letting in, who you’re keeping out and why.

How are AI training and AI search different?

Generative Engine Optimisation (GEO) is about helping your business appear clearly and accurately in AI-generated answers. Like SEO, that starts with useful content and a website that search services can access. But AI access isn’t all or nothing. Think of it more like a guest list.

OpenAI separates GPTBot from OAI-SearchBot. GPTBot collects content that may be used to train OpenAI’s models, while OAI-SearchBot helps surface websites in ChatGPT search. You can block one while allowing the other.

Anthropic makes a similar distinction between ClaudeBot and Claude-SearchBot, separating potential training use from search.

Google handles this differently. Googlebot supports Google Search, including AI Overviews and AI Mode. Blocking Google-Extended, a separate control for certain Gemini uses, does not remove your site from Google Search or its AI features.

The practical point: decide which uses you want to permit, then check the relevant provider’s controls.

1. Check your robots.txt file

Your robots.txt file gives crawlers instructions about which parts of your website they can access. Think of it as reception: it tells different visitors where they can and can’t go.

To see yours, visit yourdomain.com/robots.txt.

Here’s a simple example:

User-agent: *

Allow: /

User-agent: GPTBot

Disallow: /

User-agent: OAI-SearchBot

Allow: /

User-agent: * means the rule applies to all crawlers unless a crawler has its own specific instructions. Allow: / means they are allowed to crawl the whole site.

In this example, most crawlers are allowed in. GPTBot is specifically blocked, while OAI-SearchBot is specifically allowed. In practical terms, OpenAI’s training crawler is blocked, but its search crawler can still access the site.

Disallow: / means the crawler named above it should not crawl any part of the site. So here, it applies only to GPTBot.

If Disallow: / appeared under User-agent: *, it would apply much more broadly and could prevent many crawlers from accessing the site.

Before making changes, check with your web team. Your robots.txt file may contain additional rules for a reason.

Remember, robots.txt is an instruction, not a security control. It should never be used to protect private information.

2. Check your firewall and CDN settings

A crawler can be allowed by robots.txt and still be blocked before it reaches your site. If robots.txt is reception, your firewall is the security guard at the door.

Firewalls and content delivery networks (CDNs), such as Cloudflare, help protect websites from unwanted traffic. Their settings can also affect legitimate search crawlers, including through defaults selected when a site is set up.

On 1 July 2025, Cloudflare announced default blocking of AI crawlers, with new domains asked whether to permit access during setup. That provides a concrete reason a restriction could exist without the marketing team having requested it.

Those defaults have continued to evolve. Cloudflare subsequently announced revised defaults for 15 September 2026, allowing Search by default while blocking Training and Agent categories on pages displaying ads for newly onboarded domains. Cloudflare also introduced settings allowing site owners to disallow AI training while remaining discoverable in search.

Since it keeps changing, ask whoever manages your hosting or security to check the current settings and whether intended search crawlers are being blocked or given a security challenge they cannot complete. Cloudflare’s AI Crawl Control supports checks on individual crawlers, while request logs help establish what is actually happening.

3. Check page-level indexing settings

A page can be accessible but still tell search services not to index it. Basically, it can get in but not store any of the information to be searchable later on.

This can happen through a noindex instruction in the page or its HTTP response headers. Search services need to crawl the page to discover that instruction.

On an important public page, view the page source and search for noindex. Your web team can confirm whether the instruction applies and check the response headers too.

Check more than your homepage. Different pages and files can have different settings, and indexing controls are not handled identically by every service.

4. Check whether crawlers can read your content

Reaching a page does not mean a crawler can read everything a visitor sees.

Many AI crawlers do not execute JavaScript, the code that can load content into a page after it opens. Vercel and MERJ’s crawler research found this limitation across several major AI crawlers. Capabilities vary, so content that depends entirely on this process can be missed.

Booking engines and interactive tools deserve particular attention. If room details or service information only appear after someone enters dates, completes a form or signs in, a crawler may never reach them. Our guide to how LLMs reach your website explores these barriers.

PDFs need a separate check. They are not automatically invisible: Google can index PDFs, and some AI services can read them. However, access restrictions, scanned pages and complex layouts can make information harder to retrieve reliably.

Ask your web team whether key information is available in the page’s initial HTML, without scripts or user interaction. Keep useful descriptions on public web pages, even when a booking engine or downloadable brochure provides further detail.

Kooba’s AI discoverability troubleshooter

Try a few questions your customers might actually ask, using an AI tool with web search enabled. Search for your service and location, a problem you solve, or information someone needs before choosing a business like yours.

Then inspect the answers and their sources. Use this table to guide your investigation. The causes are possibilities, not diagnoses, and one answer is not a definitive visibility score.

  • Symptom:
    Your business doesn’t appear at all
    • Likely causes:
      Access or indexing restrictions, unreadable content, limited relevant information, or variation between answers.
      • First check:
        Run the four checks above on your main service pages. Test several relevant questions before drawing conclusions.
  • Symptom:
    Competitors appear instead
    • Likely causes:
      More specific service information, stronger supporting evidence, or relevant coverage elsewhere.
      • First check:
        Open the cited sources and identify useful information or evidence your site lacks.
  • Symptom:
    Your business appears, but details are vague or wrong
    • Likely causes:
      Outdated pages, conflicting listings, unclear descriptions, or an AI error
      • First check:
        Compare the answer with your website and cited sources. Correct inaccurate information where you have control.
  • Symptom:
    AI mentions your business without linking to your website
    • Likely causes:
      The answer may draw on directories, reviews or other coverage. Citation choices also vary.
      • First check:
        Inspect the sources provided. A missing link alone does not show that your site is blocked.
  • Symptom:
    You appear for your brand name but not broader service questions
    • Likely causes:
      Your brand is identifiable, but relevant services, specialisms or customer needs have limited coverage.
      • First check:
        Review whether your service pages and case studies answer those broader questions with useful detail.

What should you check after access is confirmed?

This is Kooba’s department. Once search services can reach and read your content, the next step is to review how clearly it explains your business and helps potential customers choose you.

We bring strategy, content, design and development into the same conversation. We can review your service pages against customer questions, examine where your business appears in AI answers and identify gaps in the information those answers draw on. Where access is an issue, we work with your hosting or security team to investigate the settings involved.

That might mean strengthening a service page, adding relevant case studies, correcting inconsistent information or adapting your CMS so your team can maintain important details more easily.

As we explore in our guide to improving AI visibility, useful content and a sound technical setup work together. The aim is to give search services clear information to reference and potential customers useful reasons to take the next step.


FAQs

Will allowing AI crawlers guarantee that my business appears?

No. Access makes content available for consideration; it does not guarantee a mention, citation or recommendation. Relevance, available sources and the question being asked all influence the answer.

How quickly will changes appear in AI answers?

There is no single timetable. A service may need to revisit your pages or refresh its search sources. Record what changed and when, then review a consistent set of questions over time.

When should we repeat these checks?

Revisit them after a website launch, migration, hosting change or significant security update. Alongside those checks, keep important business information current and review visibility across relevant questions.

Who should own AI discoverability within a business?

Marketing should define the audiences, services and questions that matter. Your web, development and hosting teams can check access and implement technical changes. A shared review helps connect website settings with the business’s visibility goals.