Should Your Law Firm Block GPTBot? A Cautious Answer for 2025
Blocking GPTBot keeps a firm's pages out of OpenAI's training data. It does not affect ChatGPT's search answers, which use a different crawler, OAI-SearchBot. Most firms that blocked 'AI bots' in 2023 blocked both without meaning to. The bot list, what each one feeds, and the robots.txt we recommend for a law firm that wants to be found.
Short answer: Blocking GPTBot is a defensible choice; it keeps the firm's content out of the datasets OpenAI trains on. It is a separate choice from being found. ChatGPT's search answers are built from pages fetched by OAI-SearchBot, and the answers cite the pages they use. A firm that blocks GPTBot and allows OAI-SearchBot gets the protection without losing the citations. A firm that blocked everything with "GPT" in the name in 2023, as many did, is invisible to ChatGPT search today and should fix its robots.txt this week.
Which bots are there?
The ones that matter for a law firm site in early 2025:
| Bot | Operator | Feeds | Recommendation |
|---|---|---|---|
GPTBot | OpenAI | Model training | Your choice; blocking costs nothing in search |
OAI-SearchBot | OpenAI | ChatGPT search results and citations | Allow |
ChatGPT-User | OpenAI | Fetches a page when a user asks about it | Allow |
Googlebot | Search and AI Overviews | Allow | |
Google-Extended | Gemini training; does not affect Search or AI Overviews | Your choice | |
PerplexityBot | Perplexity | Perplexity answers | Allow |
ClaudeBot | Anthropic | Model training | Your choice |
Bingbot | Microsoft | Bing and Copilot | Allow |
Applebot | Apple | Siri, Spotlight, Apple Intelligence | Allow |
CCBot | Common Crawl | Public dataset used by many models | Your choice |
Bytespider | ByteDance | Training; ignores robots.txt in practice | Block at the firewall if it is a problem |
Google publishes its crawler list, OpenAI documents its three bots, and the rest are documented by their operators. Bot names change; check the operator's page before adding a rule.
What does blocking GPTBot achieve?
It expresses a preference OpenAI says it honors: pages disallowed for GPTBot are not used in future training runs. It does not remove anything already trained on, and it does not affect whether ChatGPT can find and cite the page today, because that is a different crawler.
For a law firm the content at stake is practice pages, attorney profiles and articles. There is a reasonable argument that a firm prefers its explanatory content not be used to train a model that answers legal questions for free. There is an equally reasonable argument that it does not matter. We do not take a position; we ask the firm and set the rule accordingly.
What does blocking OAI-SearchBot achieve?
Invisibility in a product that, as of this spring, is available to every ChatGPT user and is being used to research lawyers. A firm's pages cannot be cited if they cannot be fetched. There is no training benefit; OpenAI says search-fetched content is not used for training. There is no reason a firm that wants clients would block it.
What did the 2023 blocks do?
In the second half of 2023 many hosting providers, security plugins and agencies added blanket rules like:
User-agent: GPTBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /
and some added wildcard rules matching any user agent containing "GPT" or "AI". Those rules predate OAI-SearchBot, but the wildcard versions catch it, and the ChatGPT-User block stops the model from reading a page a user asks it to look at. We find one of these on roughly a third of the Los Angeles firm sites we audit, usually with nobody at the firm aware it is there.
What robots.txt do we recommend?
For a firm that wants to be found and has no view on training:
User-agent: *
Allow: /
Disallow: /wp-admin/
Disallow: /thank-you/
Sitemap: https://www.examplefirm.com/sitemap.xml
For a firm that wants to be found and wants to opt out of training:
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: *
Allow: /
Disallow: /wp-admin/
Sitemap: https://www.examplefirm.com/sitemap.xml
OAI-SearchBot, ChatGPT-User, PerplexityBot, Bingbot, Googlebot and Applebot fall under the * group and are allowed. Robots rules apply per group: a bot uses the most specific group that names it, so the training bots get their own Disallow and everything else gets the default.
How do you check the current file?
Open https://www.yourfirm.com/robots.txt in a browser. Read every User-agent line. If any group would match a search crawler and disallow it, fix it. Then confirm in Search Console that Googlebot is unaffected. Our AI crawler checker reads a robots.txt and reports what each of the bots above is allowed to do.
What does not work?
- Blocking bots to "protect content" from competitors. Competitors read the site in a browser.
- A
noaimeta tag as the only measure. Not widely honored; use robots.txt. - Blocking everything and hoping ranking is unaffected. Googlebot is often caught in the wildcard.
The crawler pass is the first item in the free visibility check, because a blocked crawler makes every other finding irrelevant. The reasoning behind allowing search crawlers is in the AI search visibility guide, and the technical work is part of the law firm SEO service.
Check your firm
See whether ChatGPT, Perplexity, Gemini and Google AI Overviews name your firm for your practice area and city.
Free AI visibility check