Isonn

Should Your Law Firm Block GPTBot? A Cautious Answer for 2025

Blocking GPTBot keeps a firm's pages out of OpenAI's training data. It does not affect ChatGPT's search answers, which use a different crawler, OAI-SearchBot. Most firms that blocked 'AI bots' in 2023 blocked both without meaning to. The bot list, what each one feeds, and the robots.txt we recommend for a law firm that wants to be found.

By Dennis ArakelyanPublished April 22, 20255 min read

Short answer: Blocking GPTBot is a defensible choice; it keeps the firm's content out of the datasets OpenAI trains on. It is a separate choice from being found. ChatGPT's search answers are built from pages fetched by OAI-SearchBot, and the answers cite the pages they use. A firm that blocks GPTBot and allows OAI-SearchBot gets the protection without losing the citations. A firm that blocked everything with "GPT" in the name in 2023, as many did, is invisible to ChatGPT search today and should fix its robots.txt this week.

Which bots are there?

The ones that matter for a law firm site in early 2025:

BotOperatorFeedsRecommendation
GPTBotOpenAIModel trainingYour choice; blocking costs nothing in search
OAI-SearchBotOpenAIChatGPT search results and citationsAllow
ChatGPT-UserOpenAIFetches a page when a user asks about itAllow
GooglebotGoogleSearch and AI OverviewsAllow
Google-ExtendedGoogleGemini training; does not affect Search or AI OverviewsYour choice
PerplexityBotPerplexityPerplexity answersAllow
ClaudeBotAnthropicModel trainingYour choice
BingbotMicrosoftBing and CopilotAllow
ApplebotAppleSiri, Spotlight, Apple IntelligenceAllow
CCBotCommon CrawlPublic dataset used by many modelsYour choice
BytespiderByteDanceTraining; ignores robots.txt in practiceBlock at the firewall if it is a problem

Google publishes its crawler list, OpenAI documents its three bots, and the rest are documented by their operators. Bot names change; check the operator's page before adding a rule.

What does blocking GPTBot achieve?

It expresses a preference OpenAI says it honors: pages disallowed for GPTBot are not used in future training runs. It does not remove anything already trained on, and it does not affect whether ChatGPT can find and cite the page today, because that is a different crawler.

For a law firm the content at stake is practice pages, attorney profiles and articles. There is a reasonable argument that a firm prefers its explanatory content not be used to train a model that answers legal questions for free. There is an equally reasonable argument that it does not matter. We do not take a position; we ask the firm and set the rule accordingly.

What does blocking OAI-SearchBot achieve?

Invisibility in a product that, as of this spring, is available to every ChatGPT user and is being used to research lawyers. A firm's pages cannot be cited if they cannot be fetched. There is no training benefit; OpenAI says search-fetched content is not used for training. There is no reason a firm that wants clients would block it.

What did the 2023 blocks do?

In the second half of 2023 many hosting providers, security plugins and agencies added blanket rules like:

User-agent: GPTBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /

and some added wildcard rules matching any user agent containing "GPT" or "AI". Those rules predate OAI-SearchBot, but the wildcard versions catch it, and the ChatGPT-User block stops the model from reading a page a user asks it to look at. We find one of these on roughly a third of the Los Angeles firm sites we audit, usually with nobody at the firm aware it is there.

What robots.txt do we recommend?

For a firm that wants to be found and has no view on training:

User-agent: *
Allow: /
Disallow: /wp-admin/
Disallow: /thank-you/

Sitemap: https://www.examplefirm.com/sitemap.xml

For a firm that wants to be found and wants to opt out of training:

User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: *
Allow: /
Disallow: /wp-admin/

Sitemap: https://www.examplefirm.com/sitemap.xml

OAI-SearchBot, ChatGPT-User, PerplexityBot, Bingbot, Googlebot and Applebot fall under the * group and are allowed. Robots rules apply per group: a bot uses the most specific group that names it, so the training bots get their own Disallow and everything else gets the default.

How do you check the current file?

Open https://www.yourfirm.com/robots.txt in a browser. Read every User-agent line. If any group would match a search crawler and disallow it, fix it. Then confirm in Search Console that Googlebot is unaffected. Our AI crawler checker reads a robots.txt and reports what each of the bots above is allowed to do.

What does not work?

  • Blocking bots to "protect content" from competitors. Competitors read the site in a browser.
  • A noai meta tag as the only measure. Not widely honored; use robots.txt.
  • Blocking everything and hoping ranking is unaffected. Googlebot is often caught in the wildcard.

The crawler pass is the first item in the free visibility check, because a blocked crawler makes every other finding irrelevant. The reasoning behind allowing search crawlers is in the AI search visibility guide, and the technical work is part of the law firm SEO service.

ShareLinkedInEmail

Check your firm

See whether ChatGPT, Perplexity, Gemini and Google AI Overviews name your firm for your practice area and city.

Free AI visibility check

Free citation audit

Find out who AI recommends instead of you.

Written report plus a thirty-minute walkthrough. No pitch until you ask for one.

Is AI citing your firm?Free check, five business days