GEO & Technical AI SEO

Your Biggest GEO Problem Might Be Your Firewall

Tandeep Sangra
September 6, 2026
8 min read
TL;DR: Your website can be fully live, fast, and well-written — and still be invisible to AI, because the block isn't happening at the content layer. It's happening at the CDN/WAF layer sitting in front of it. Cloudflare now blocks AI crawlers by default for most new and free-plan sites, and as of September 2026 that default was extended even further. A 403 error tells you access was denied — it doesn't tell you who denied it. Before you diagnose a GEO content problem, rule out a firewall problem first.
~20%
of all internet traffic runs through Cloudflare's infrastructure — meaning this default setting affects a huge share of the web, often without site owners realizing it
Jul 2025
Cloudflare became the first major infrastructure provider to block AI crawlers by default for new domains and free-plan customers
Sep 2026
The default expanded further — "mixed-use" crawlers blending search, agent, and training functions are now blocked by default on any ad-hosting page

The Website Is Live. But Can AI Actually Reach It?

Every AI visibility conversation eventually gets to the same question: why isn't my brand showing up in ChatGPT, Perplexity, or Google AI Mode? Most of the time, people go straight to the content layer — better schema, better entity signals, better structured answers.

But there's a layer underneath all of that which almost nobody checks first: can the AI crawler even reach the page at all?

A site can be fast, well-structured, and full of exactly the kind of content an AI system would want to cite — and still be functionally invisible, because something between the crawler and the content is quietly saying no.

A Real Pattern We Keep Running Into

Here's a scenario that comes up more often than you'd expect. A team is actively working on their AEO/GEO strategy, and their previous hosting setup kept getting in the way — blocking AI crawlers inconsistently, making it hard to even test whether their optimisation work was landing. So they move the domain behind Cloudflare specifically to get more granular control over crawler access, rather than leaving that decision to the hosting provider by default.

Reasonable move. Except when they test the site afterward, they find something unexpected: the homepage triggers a Cloudflare verification challenge, while individual product pages are perfectly accessible and readable to the same crawler.

That's the whole problem in miniature. The site isn't uniformly blocked or uniformly open — it's inconsistently gated, one layer removed from anything the content or SEO team actually controls.

Why a 403 Doesn't Mean What Most People Assume

When a crawler test comes back with a 403 or a challenge page, the instinct is to read that as "this website blocks AI." That's often the wrong conclusion.

A 403 tells you access was denied. It does not tell you who denied it. It could be:

The distinction matters enormously, because each of those has a completely different fix — and if you diagnose it as a content or schema problem when it's actually a firewall setting, you can spend weeks optimising the wrong layer.

Cloudflare Blocks AI Crawlers by Default — And Most Site Owners Don't Know It

This is the part worth sitting with. In a July 2025 announcement, Cloudflare became the first major internet infrastructure provider to block AI crawlers from accessing content without permission by default. Roughly a fifth of all internet traffic runs through Cloudflare's network, so this wasn't a niche setting — it flipped the default for a huge share of the web in one move.

And the story didn't stop there. As of September 15, 2026, Cloudflare tightened this further: crawlers that blend search indexing, AI-agent retrieval, and model training into one "mixed-use" identity are now blocked by default on any page that hosts ads — unless the site owner explicitly overrides it.

If you set up your Cloudflare account any time recently, or if you're on a free plan, there's a real chance AI crawlers are being blocked right now without you ever having made that choice consciously. It's not malicious, and it's not even unreasonable — Cloudflare built this in response to years of AI companies scraping content with no attribution or compensation. But from a GEO standpoint, it means the setting most likely to be silently working against you is one you may never have touched.

The Real Goal Isn't "Allow All" or "Block All"

The instinctive fix, once you learn this, is to just flip everything open — allow every AI bot, disable every WAF rule that might catch one. That's the wrong instinct too.

The goal isn't blanket permission. It's controlled AI accessibility — knowing precisely which crawlers can reach your content, what they actually receive when they do, and whether that content can be retrieved and cited the way you intend.

Three questions capture this well:

Sometimes the GEO problem was never your content. It's the security layer sitting quietly between your content and the crawler trying to read it.

A Practical Checklist: Auditing Your Own AI Accessibility

  • →
    Cloudflare AI bot / WAF settings
    Check Security > Bots in your dashboard for AI-specific rules, and review your WAF custom rules for anything blocking by user-agent.
  • →
    robots.txt
    Confirm you're not explicitly disallowing GPTBot, ClaudeBot, PerplexityBot, or other AI crawlers you actually want indexing you.
  • →
    HTTP status codes
    Test key pages directly — a 200 with full content is very different from a 403 or a redirect to a challenge page.
  • →
    What the crawler actually receives
    A status code alone isn't enough — check whether the response body is your real HTML or a CAPTCHA/verification page.
  • →
    More than just the homepage
    Test your highest-value pages individually. As the pattern above shows, a homepage block and full page-level access can coexist on the same site.
The Core Insight

An invisible-to-AI website isn't always a content problem. Sometimes it's an infrastructure setting nobody remembered to check.

Your CDN and WAF exist to protect you from real threats — but the same tools built to keep bad bots out can just as easily keep the AI systems you want citing you locked out too, often as a default you never explicitly chose. Checking this layer takes minutes. Assuming it's fine and optimising the wrong thing for weeks costs a lot more.

Is your website actually AI-accessible?

Run a free AI Visibility Audit and find out whether your homepage, your key pages, and your technical setup are letting AI systems in — or quietly shutting them out.

Run Your Free AI Visibility Audit →

Frequently Asked Questions

Yes. Since July 2025, Cloudflare blocks AI crawlers by default for new domains and free-plan customers. As of September 15, 2026, this extends further: "mixed-use" crawlers that blend search, AI-agent, and training functions are blocked by default on any page hosting ads, unless the site owner manually overrides it.
Not necessarily. A 403 tells you access was denied, not who denied it or why. It could be your CMS, your hosting provider, or — increasingly common — your CDN/WAF layer like Cloudflare making that decision on your behalf, often without your knowledge.
Log into your Cloudflare dashboard and check Security > Bots for AI-bot-specific rules, and check your WAF custom rules for anything blocking by user-agent. Also test manually: fetch your homepage and a few key pages with a tool that sends a GPTBot or similar user-agent string and see what status code and content comes back.
Not automatically. The goal isn't "allow everything" — it's controlled accessibility. Decide deliberately which crawlers you want reaching your content (the major AI search platforms you want citing you) versus bots you have no reason to allow, rather than defaulting to either extreme.
This usually points to a rule or challenge specifically triggered on the root domain — sometimes from a security setting, a bot-fight-mode default, or a caching rule that behaves differently on the homepage than on individual URLs. It's a strong sign the block is happening at the CDN/WAF layer, not the content itself.
Yes, directly. If a crawler can't retrieve your page's actual content — because it hits a challenge screen instead of your HTML — there's nothing for the AI system to read, summarise, or cite from that page, regardless of how good the content itself is.
No — Cloudflare is simply the largest and most commonly discussed example because of how much of the web it sits in front of. Any CDN, WAF, or bot-management layer can behave the same way, so the checklist (status codes, actual response content, page-by-page testing) applies regardless of which provider you use.