Quick answer
Yes, a firewall or CDN setting meant to stop AI bots can also block search engines. Cloudflare treats crawlers that do both search and AI training, including Googlebot, Bingbot and Applebot, by their strictest purpose, so blocking training can block them too. New defaults for new Cloudflare domains took effect on 15 September 2026. If you want to be found, allow search crawlers, handle training bots with precise robots.txt rules, and check Search Console for crawl errors after any change.
Key takeaways
- Cloudflare sorts bots into Search, Training and Agent. Crawlers that do both search and training are treated by the stricter rule.
- Blocking AI training at the firewall can therefore block Googlebot, Bingbot and Applebot.
- Real sites have lost weeks of traffic, Google Ads and Merchant Center listings this way.
- Blocking Google-Extended in robots.txt does not affect Google Search, according to Google. Broad firewall rules are the risk.
- After any bot or firewall change, test with Search Console URL Inspection and watch crawl stats.
Most business owners never look at their firewall settings. Someone else set them up, often an IT provider or a developer who has since moved on.
That is exactly how this problem sneaks in.
Over the past year, services like Cloudflare made it easy to block AI bots with a switch. Many site owners turned it on. Some chose stricter settings to cut server load from aggressive crawlers.
In some cases Google got blocked along with everything else. Traffic fell, and nobody connected the drop to the firewall for weeks.
How we got here
On 1 July 2025, Cloudflare changed its default to block AI crawlers on new domains unless they pay creators for content. Its chief executive called it Content Independence Day.
Cloudflare CEO Matthew Prince on the July 2025 change. View the post on X.
A year later, Cloudflare went further and began sorting crawlers by what they do. That is where the risk for search visibility comes from.
What changed at Cloudflare in 2026?
On 1 July 2026, Cloudflare announced new AI traffic options for all customers. Its changelog describes three categories of bot behaviour:
- Search: crawlers that index content for search results.
- Agent: bots acting in real time for a user, such as AI assistants and browser agents.
- Training: crawlers collecting content to train AI models.
For each one, site owners can block on all pages, block only on pages that show ads, or not block at all.
The 15 September defaults
From 15 September 2026, new domains block Training and Agent bots on ad-supported pages, while Search bots stay allowed. Existing customers were told they could opt out of the new defaults before that date.
The catch: crawlers with two jobs
A crawler that performs both Search and Training will be blocked if a site blocks Training.
Search Engine Journal, summarising Cloudflare’s rules
Cloudflare names Googlebot, Applebot and Bingbot as examples of these mixed-purpose crawlers. Block training, and you may block the search engines people use to find you.

Has this actually happened to real websites?
Yes, several times, and the symptoms rarely point at the firewall.
The sitemap that returned 403
In August 2026, Search Engine Journal reported a case from the r/TechSEO community. With the AI training block turned on, Googlebot and Bingbot got 403 errors when fetching the sitemap. Turning the block off restored access straight away.
Two weeks of lost traffic
Please be careful with CloudFlare configuration. I just had to fix a website who’s traffic tanked for 2 weeks.
Jonathan Bird, SEO consultant, via Search Engine Roundtable
In that case, an IT provider had set Cloudflare to block all bots. Google Ads broke and Merchant Center listings disappeared as well as organic traffic.
The drop that looked like an algorithm update
SEO consultant Brodie Clark described a marketplace where bot rules added at the firewall to protect server performance caused what he called an SEO disaster. As he put it, it may look like a core or spam update, but it was not.
And sometimes it is not Cloudflare at all
In a Cloudflare community thread from July 2026, a site owner saw verified Bingbot blocked from their sitemap by the AI training rule. The final cause turned out to be a user-agent rule in their own server’s .htaccess file. Check your origin server rules too, not just your CDN.
What are the warning signs?
Look for two or more of these together:
- A sudden, sharp drop in impressions across the whole site, not just some pages.
- Crawl errors or 403 responses in Search Console, especially on the sitemap.
- The URL Inspection tool failing to fetch pages that load fine in your browser.
- Google Ads disapprovals for landing pages that are clearly online.
- Merchant Center products disappearing without a stated reason.
- In Cloudflare’s security analytics, requests from Googlebot or Bingbot being blocked by a managed rule.
If you see these, check the firewall before assuming an algorithm update. We made the same point in our note on the August spam update: rule out site changes before blaming Google.
Does blocking AI training bots hurt your SEO?
Blocking a pure training bot, done precisely, does not affect Google rankings. The method is what matters.
Google-Extended is safe to block
Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.
Google Search Central, Google’s common crawlers
Google-Extended controls whether your content is used for Gemini training and grounding. A robots.txt rule for it is precise. A broad firewall rule based on bot categories can catch Googlebot itself.
OpenAI’s bots do different jobs
According to OpenAI, GPTBot collects content for training, while OAI-SearchBot powers ChatGPT’s search results. You can block the first and still appear in ChatGPT answers. Block the second and your pages stop appearing in ChatGPT search.
Check that a bot is really who it says
Anyone can fake a user agent. Google explains how to verify that a visitor is really Googlebot using reverse DNS or its published IP ranges. Good firewall rules use verification, not names alone.
What is a safe setup for most businesses?
- Allow the search crawlers: Googlebot, Bingbot, Applebot, DuckDuckBot and OAI-SearchBot.
- Decide separately about training bots such as GPTBot and Google-Extended, and use robots.txt rules for precise control.
- In Cloudflare, review the AI crawler settings, Bot Fight Mode and any custom rules. Make sure none of them block the search crawlers above.
- Check your origin server too: .htaccess, security plugins and hosting firewalls can block bots independently.
- If you need rate limits for server load, apply them to specific abusive bots, not whole categories.
- After any change, run Search Console’s URL Inspection on your homepage and a key service page, and watch crawl stats for a week.
Our technical SEO services include this check in every audit, and our web design and development team sets it up correctly on new builds.

Who should own these settings?
Whoever changes them should understand search. Often that is not the case.
IT providers think about security and server load. Marketing teams think about visibility. They do not always talk.
Agree one simple rule inside your business: no firewall, CDN or bot setting changes without telling whoever manages your SEO. It costs nothing and prevents the most expensive kind of mistake.
Common crawlers and what blocking them does
| Crawler | Main purpose | If you block it |
|---|---|---|
| Googlebot | Google Search index | You disappear from Google Search and its AI features |
| Google-Extended | Gemini training and grounding | No effect on Google Search, according to Google |
| Bingbot | Bing index, also used by other services | You lose Bing and the tools built on it |
| OAI-SearchBot | ChatGPT search results | Your pages stop appearing in ChatGPT search answers |
| GPTBot | OpenAI model training | Your content is not used for training |
| Applebot | Siri and Spotlight suggestions | You lose visibility in Apple's search features |
How does this connect to AI visibility?
Directly. AI tools can only cite what they can reach.
If your firewall blocks OAI-SearchBot, ChatGPT search cannot use your pages. If it blocks Googlebot, you are out of AI Overviews too, because Google’s AI features rely on its normal index.
Since August, ChatGPT has also been searching named websites directly far more often, as we explained in our Reddit citations field note. A blocked site misses that opportunity completely.
You can check whether Google shows you in its AI features with Search Console’s Generative AI performance report.
Read more: Can you pay to show up in ChatGPT?
What should you do this week?
- Log in to your CDN or firewall, or ask whoever manages it, and list every bot-related setting that is switched on.
- Run URL Inspection in Search Console on your homepage and one key service page.
- Open your robots.txt and confirm Googlebot and OAI-SearchBot are not disallowed.
- If anything fails, fix it today. Every day of blocked crawling is a day of lost visibility.
If you would like someone to check it for you, Queens Digital is a digital marketing agency in Nepal and crawl access is the first thing we review in our SEO services in Nepal.
Explore More
Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.
To help ensure your site appears in search results, we recommend allowing OAI-SearchBot in your site’s robots.txt file.
Reviewed by
K.K. Bhatta
K.K. Bhatta is a general digital marketer at Queens Digital Agency: hands-on across SEO, paid media, and whatever platform changes next. A lifelong learner rather than a one-topic specialist, he tests ideas on live campaigns before writing about them, and stays skeptical of anything that has not actually been tried.
Every claim in this article was checked against a live source before publishing.
Yes, if set that way. Cloudflare treats crawlers that do both search and AI training, including Googlebot, by their strictest purpose, so blocking AI training can also block Googlebot.
New domains block Training and Agent bots on ad-supported pages by default, while Search bots stay allowed. Crawlers that do both search and training are affected by the training block.
Blocking Google-Extended through robots.txt does not affect Google Search, according to Google. The risk comes from broad firewall rules that also catch Googlebot.
Look for a sudden site-wide drop in impressions, 403 or crawl errors in Search Console, and URL Inspection failing to fetch pages that load normally.
Allow Googlebot, Bingbot and Applebot for search, and OAI-SearchBot for ChatGPT search. Training bots like GPTBot and Google-Extended can be handled separately in robots.txt.
Sources and references
Every claim above comes from the sources below, verified live before use. Accessed 25 September 2026.
- Cloudflare, “Your site, your rules: new AI traffic options for all customers”
- Cloudflare changelog, New options to manage AI traffic, 1 July 2026
- Cloudflare, “Content Independence Day: no AI crawl without compensation”, 1 July 2025
- Matthew Prince on X, 1 July 2025
- Search Engine Journal, “Cloudflare’s AI Crawler Rules Can Block Googlebot”
- Search Engine Journal, report on Cloudflare AI bot blocking and Googlebot
- Search Engine Roundtable, “Misconfiguring Cloudflare Can Hurt Your SEO Badly”
- Cloudflare Community, verified Bingbot receives 403 for sitemap
- Google Search Central, Google’s common crawlers
- Google Search Central, Verifying Googlebot and other Google crawlers
- OpenAI, Overview of OpenAI crawlers

