If you built your site on Wix, Squarespace, Shopify, Webflow or WordPress.com, you probably never wrote a robots.txt file. Your platform wrote one for you. So it is fair to ask whether that file is quietly turning ChatGPT, Perplexity and Claude away, and plenty of articles claim one platform or another does exactly that.
We checked. On 8 October 2026 we read the default robots.txt that each platform actually serves, alongside each platform's own help pages. The short answer: none of the major website builders blocks AI crawlers out of the box. Every one of them lets the AI bots in until somebody changes a setting.
That is not the end of the story, though, because the real risk was never the default. It is the one switch each platform gives you for AI bots, and what that switch actually sweeps up when someone flips it. Those switches are not built the same way. Some block only the bots that collect training data. Others also block the bots that decide whether an AI engine can quote you today. And since September 2026 there is a second layer that can block AI crawlers, and even Google, without anything showing in your robots.txt at all.
This guide goes platform by platform: what the default file says, where the AI setting lives, and the catch hiding in each one.
The short answer
| Platform | Blocks AI crawlers by default? | The AI setting | The catch |
|---|---|---|---|
| Shopify | No | No switch. Edit the robots.txt.liquid theme template | Old copy-pasted "block AI" snippets live on in the template |
| Wix | No | Robots.txt Editor, edited by hand | Wix's own example blocks a ChatGPT bot that only visits when a person asks |
| Squarespace | No | Settings > Crawlers, one checkbox | The checkbox also blocks DuckDuckGo's AI answer bot |
| Webflow | No | Traffic control toggles in Site settings | "AI bots" includes Perplexity's search crawler |
| WordPress.com | No | "Prevent third-party sharing" checkbox | Also blocks Perplexity's search crawler |
| WordPress (self-hosted) | No | Plugins, or a file in your site's root folder | Whatever a plugin or an old file says wins |
Every row says "No" for the default. The rest of this guide is about the last column.
Why "by default" is the wrong question
A robots.txt file can only say no. Every crawler is allowed everywhere unless a rule names it, or names everyone, and tells it to stay out. So a default file that does not mention AI bots is a default that welcomes them. Our guide to allowing AI crawlers covers how those rules are read in detail.
What matters more is that AI bots do three different jobs, and blocking each one costs you something different:
- Training crawlers (GPTBot, ClaudeBot, CCBot and others) collect pages to train future AI models. Blocking them has no effect on whether an AI engine can quote you today. Google is the odd one out:
Google-Extendedis not a crawler but a name you can block, and it covers both Gemini training and Gemini using your pages to back up its answers, so blocking it is a Gemini decision, not just a training one. - Search crawlers (OAI-SearchBot for ChatGPT search, PerplexityBot, Claude-SearchBot) build the index that AI engines quote from. Block one and you drop out of that engine's answers.
- Fetchers that act for a person (ChatGPT-User, Claude-User, Perplexity-User) open your page because someone asked an assistant about it. Block one and the assistant comes back empty-handed to someone who was asking about you.
Our AI crawler library has a page for each of these bots, with what it does and what blocking it changes. The point for this guide is simple: a platform switch labelled "block AI" can mean "block training only," or it can mean "block everything with AI in its name." The label never tells you which. Only the bot list behind it does.
Shopify
Default: open. The default file on Shopify's own demo store gives every crawler the same rules (User-agent: *), keeps them out of the admin, cart, checkout and account pages, and leaves every product, collection, page and blog post open. No AI bot is named, so none is blocked.
The file also does something no other platform's does. It opens with a block of comments written to AI agents rather than to search engines. They point to an agents.md file and a shopping endpoint built for AI assistants, and they say, in plain terms, that checkouts are for humans and an agent must get a person's approval before paying. Shopify is not trying to keep AI out. It is laying out a route for AI shopping assistants to come in, and setting the rules for them.
The setting: there is no AI switch. To change the file you add a robots.txt.liquid template to your theme, which replaces the file Shopify generates. Shopify's developer documentation recommends building on its own default rules inside that template rather than writing a file from scratch, and warns that the output of a custom template does not always match the file Shopify would have produced.
The catch: guides from 2023 and 2024 handed out "block all AI bots" snippets for this template, at a time when blocking looked like the cautious choice. That template does not update itself. Shopify can change its default file as often as it likes, but a store with a custom template keeps serving whatever was pasted in years ago. Apps that edit the file add another way for old rules to creep in. If you have read that "Shopify blocks AI bots by default," check your own store before believing it: the default we read blocks none, but your theme might.
Wix
Default: open. Wix generates the file automatically and marks it as such at the bottom. The default lets every crawler in, blocks just one bot by name (PetalBot, which belongs to Huawei's search engine), slows down two SEO tool crawlers, and names no AI bot at all.
The setting: there is no AI switch. You edit the file by hand in the Robots.txt Editor (in the dashboard, under SEO & GEO, then Tools and settings). Page-level control lives in each page's Advanced SEO settings.
The catch: Wix's own help article on blocking AI crawlers gives a ready-made block for four bot names: CCBot, GPTBot, ChatGPT-User and BingAI. Paste it in and two things happen that you probably did not intend.
First, ChatGPT-User is not a crawler that collects your content. It is the fetcher that opens your page when a ChatGPT user asks about it. OpenAI says it is not used to decide whether content appears in ChatGPT search, and that robots.txt rules may not apply to it at all, because a person started the visit. So that line either does nothing or turns away people who were asking about you, and it does nothing about training, which is GPTBot's job.
Second, BingAI does not appear on Microsoft's own list of Bing crawlers. Bing runs its search and its Copilot answers off the same crawler, Bingbot, and has no separate AI-only bot. A rule for a name no crawler uses matches nothing, so the line does nothing.
If you want to stop AI training on Wix, block the training crawlers by their real names and leave the search and fetcher bots alone.
Squarespace
Default: open, and Squarespace says so directly. Its crawler settings article says: "We default to having the box unchecked."
There is a quirk that catches people out. The default Squarespace file already lists 26 AI bot names near the top: GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider and more. Read it quickly and it looks like they are all blocked. They are not. In a robots.txt file, consecutive User-agent lines form one group that shares the rules underneath, and in the default file those 26 names sit in the same group as User-agent: *. They get exactly what every other crawler gets: everything except a few system folders. The names are listed so the checkbox has somewhere to switch.
The setting: one checkbox, "Block known artificial intelligence crawlers," under Settings, then Crawlers. Squarespace says checking it "updates your robots.txt file to tell the following bots not to crawl your site," and gives the same 26 names. You cannot edit the file in any other way.
The catch: the list is mostly training crawlers, which is fine if training is what you want to stop. But it also includes DuckAssistBot, and that one is not a training crawler. DuckDuckGo describes it as the bot that "crawls pages in real-time for our AI-assisted answers, which prominently cite their sources," and says its data "is not used in any way to train AI models." Tick the box to stop training and you also drop out of DuckDuckGo's AI answers. The list includes Google-Extended too, which also keeps your pages out of Gemini's answers.
The list also has gaps in the other direction. It does not include OAI-SearchBot, PerplexityBot, Claude-SearchBot or any of the fetchers that act for a person, so the main AI search engines can still reach you with the box ticked. Because Squarespace does not let you edit the file, it is all 26 or none of them.
Webflow
Default: open. Webflow's indexing help article puts it plainly: "Traffic from search engine crawlers and AI bots is allowed by default."
The setting: Webflow has the most complete controls of any builder here, all under Site settings, then SEO, then Indexing. There is a free-text robots.txt box, a Content-Signal option (more on that below), and two "Traffic control" toggles that write rules for whole groups of bots at once: one for search engine crawlers, one for AI bots.
The catch: the AI bots toggle covers a list Webflow publishes, and it mixes the three jobs together. Alongside training crawlers like GPTBot, ClaudeBot and CCBot, it includes PerplexityBot (Perplexity's search crawler) and ChatGPT-User and Claude-User (the fetchers that act for a person). Meanwhile OAI-SearchBot, ChatGPT's search crawler, sits in the search engine group. So switching AI bots off removes you from Perplexity's answers but not from ChatGPT search, which is probably not what anyone choosing that toggle had in mind. If you only want to stop training, write the rules yourself in the robots.txt box instead.
Two more Webflow details are worth knowing:
- The
webflow.ioaddress. Every Webflow site also lives at awebflow.ioaddress. Turning "Staging indexing" off publishes arobots.txton that address only, telling every crawler to stay away. That is the right setting once your site runs on its own domain. It becomes a problem only if thewebflow.ioaddress is the one you have been sharing. - Content-Signal. Webflow can add a Content-Signal line, such as
ai-train=no, search=yes, ai-input=no, which states how your content may be used rather than whether it may be visited. Webflow notes that it is based on a proposed standard, not an accepted one, so treat it as a statement of your wishes rather than a lock.
WordPress.com
Default: open. The hosted WordPress.com sites we checked that had not changed the setting name no AI bots in their file.
The setting: a checkbox called "Prevent third-party sharing," which WordPress.com's visibility settings article places under Settings, then Reading, in the Site Visibility section. The article says it "adds known AI bots to the 'disallow' list in your site's robots.txt file," and also keeps your content out of the data WordPress.com shares with outside partners.
The catch: on sites with the box ticked, the list blocks about fifteen bots outright. Most are training crawlers, but one is PerplexityBot, Perplexity's search crawler. Ticking a box that reads as a data-sharing choice also takes you out of Perplexity's answers, and, because Google-Extended is on the list, out of Gemini's. The checkbox does not block ChatGPT's or Claude's search crawlers.
One more detail: when WordPress.com introduced the setting in February 2024 (then called "Prevent third-party data sharing"), it switched it on automatically for sites that had already asked search engines not to index them. If your site was set to discourage search engines at any point back then, check this box even if you never touched it.
WordPress you host yourself
Default: open. WordPress generates its own robots.txt when no real file exists, and that generated version names no AI bots.
The setting: there isn't one in WordPress itself. Control comes from SEO plugins (most have a robots.txt editor), from security plugins, from your host, or from a real robots.txt file placed in your site's root folder, which replaces the generated one entirely.
The catch: with several tools able to write the file, the version actually served may not be the one you last edited, and a real file in the root folder overrides everything the plugins generate. Open the live file to see which version won.
The layer in front of your site: Cloudflare
Everything above is about robots.txt. But many sites, on every platform, sit behind a service that handles traffic before it reaches the site at all. The biggest is Cloudflare, and it changed its defaults twice in the last fifteen months. If you, your developer or your agency connected your domain to a Cloudflare account, these settings apply to you no matter which builder you use.
July 2025: Cloudflare started asking every new domain at sign-up whether to allow AI crawlers, and made blocking the starting position. Blocks set here never appear in your robots.txt. The crawler simply gets refused.
September 2026: Cloudflare replaced its single "Block AI bots" switch with three categories, matching the three jobs above: Search, Agent (bots acting for a person in real time) and Training. Its bot settings documentation says that for new domains from 15 September 2026, "bots classified as Training or as Agent will be blocked on pages that display ads, and Search will remain allowed."
The second part of that change matters more. Cloudflare classes some crawlers as doing two jobs at once, search and training, and names Googlebot, Bingbot and Applebot among them. From 15 September, Cloudflare treats a crawler that does two jobs according to its strictest one, so a site that blocks Training now blocks those crawlers as well. In plain terms: a site owner who switched on "block AI training" a year ago to keep their writing out of AI models may now be blocking Google's main search crawler too.
Cloudflare added an option called "Disallow AI Training" for exactly this case. Instead of refusing those crawlers, it writes a no-training preference into your robots.txt and keeps letting them in for search. Cloudflare notes that Microsoft expects to support that preference only in early 2027, so for now it sends no no-training signal to Bing.
If your site uses Cloudflare, open its security settings and look at the AI crawler section. If Training is blocked, decide whether you meant to block Googlebot and Bingbot as well. Almost nobody did.
What each switch actually blocks
Here is the same information arranged by what you lose, assuming each platform's AI switch is turned on:
| Switch turned on | Training crawlers | AI search crawlers | Fetchers acting for a person |
|---|---|---|---|
| Squarespace "Block known AI crawlers" | Blocked (most) | DuckAssistBot and Gemini (via Google-Extended) blocked; ChatGPT, Perplexity and Claude search not | Not blocked |
| Webflow AI bots toggle off | Blocked (most) | PerplexityBot blocked; ChatGPT search not | ChatGPT-User and Claude-User blocked |
| WordPress.com "Prevent third-party sharing" | Blocked (most) | PerplexityBot and Gemini (via Google-Extended) blocked; ChatGPT and Claude search not | Not blocked |
| Wix help article's example | GPTBot and CCBot only | Not blocked | ChatGPT-User named (may not be honoured) |
| Cloudflare, Training blocked (since September 2026) | Blocked | Search allowed, but Googlebot, Bingbot and Applebot blocked | Agents blocked on ad pages for new domains |
| Shopify, WordPress you host yourself | Whatever your template, plugin or file says | Same | Same |
No two switches agree. If your goal is "stay quotable in AI answers but keep my content out of training," none of the one-click switches does exactly that. Writing your own rules does, on the platforms that let you.
How to check what your site is really telling AI crawlers
- Read the file that is actually served. Open
yourdomain.com/robots.txtin a browser. Not the settings screen, not a copy in a folder: the live file. - Find each named group. Look for
User-agent:lines naming AI bots and read the rule directly underneath.Disallow: /means fully blocked. A name sitting in the same group asUser-agent: *gets the general rules, as on Squarespace's default. - Check the search crawlers and the fetchers by name. The ones that matter for being quoted are OAI-SearchBot, PerplexityBot and Claude-SearchBot, plus ChatGPT-User, Claude-User and Perplexity-User. If any of them is blocked, find which setting put it there.
- Check the layer in front. If your domain runs through Cloudflare, open its AI crawler settings. A block there never shows in
robots.txt. - Confirm with real visits. Your server or hosting logs show whether the AI crawlers are actually arriving and what answer they get. Our guide to reading server logs shows what to search for. On hosted builders that don't give you logs, Cloudflare's own bot analytics can show the same thing.
Common mistakes
Trusting a blog post's claim about your platform's default. Defaults change, and several widely shared articles describe settings that are no longer true. Your live robots.txt is the only source that counts for your site.
Reading a list of bot names as a list of blocks. Squarespace's default file names 26 AI bots and blocks none of them. A name only means something next to the rule underneath it.
Copying a platform's example rule without checking the names. Wix's example names a person-triggered fetcher and a bot that doesn't exist. Check every name against the bot operator's own documentation before pasting it.
Treating "block AI" as one decision. Blocking training and staying quotable are separate choices, and most switches bundle them. Our piece on whether to block AI crawlers walks through the trade-off for each kind of bot.
Forgetting a choice made years ago. A Shopify template from 2023, a WordPress.com setting switched on automatically in 2024, a Cloudflare toggle from 2025 that now catches Googlebot: the settings most likely to hurt you are the ones nobody remembers making.
Frequently Asked Questions
Is your website ready for AI?
Check now. Free readiness score in under a minute, no signup, no card.