Which AI Crawlers Should a Law Firm Allow? A robots.txt Guide for 2026
Allow the retrieval bots that fetch pages to answer live questions (OAI-SearchBot, ChatGPT-User, PerplexityBot, Claude-SearchBot, Google-Extended) and you stay citable. Blocking them removes your firm from AI answers, not from training.
Short answer: Allow every crawler that fetches pages to answer a live question: OAI-SearchBot and ChatGPT-User (OpenAI), PerplexityBot and Perplexity-User, Claude-SearchBot and Claude-User (Anthropic), Google-Extended and the regular Googlebot, Bingbot and Applebot. Blocking them removes your firm from AI answers. If you want to opt out of model training only, block GPTBot, ClaudeBot and Applebot-Extended and keep the rest open.
Why does a law firm's robots.txt matter for AI answers?
Because AI engines cannot name a firm they cannot read. When someone asks ChatGPT for a family lawyer in Glendale, the engine runs a search, fetches a handful of pages in real time, and writes an answer from what it fetched. If your robots.txt turns that fetch away, the engine quotes someone else.
In August 2026, 41.9% of surveyed consumers said they would use ChatGPT to research which lawyer to hire (iLawyer Marketing, 1,110 respondents). Roughly a quarter of legal searches on Google now open with an AI Overview (SEO Discovery, 2026). A blanket block on "AI bots", which many hosting panels and security plugins now offer as a one-click option, opts a firm out of both.
What is the difference between a training crawler and a retrieval crawler?
A training crawler collects pages to build the next version of a model. A retrieval crawler fetches a page because a person asked a question right now. Blocking the first has no effect on whether you are cited. Blocking the second guarantees you are not.
The providers publish separate user agents for each job, which is what makes a selective policy possible.
| Provider | Training crawler | Retrieval / live-answer bots | What blocking retrieval does |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot (search index), ChatGPT-User (fetch on request) | Your pages disappear from ChatGPT search results and answers |
| Anthropic | ClaudeBot | Claude-SearchBot, Claude-User | Claude cannot cite you |
| Perplexity | (none separate) | PerplexityBot, Perplexity-User | Perplexity, the most citation-heavy engine, skips you |
| Google-Extended (Gemini training + grounding) | Googlebot serves AI Overviews and AI Mode | Blocking Googlebot removes you from Google entirely; blocking Google-Extended limits Gemini | |
| Apple | Applebot-Extended | Applebot | Siri and Spotlight answers lose you |
| Microsoft | (Bingbot covers both) | Bingbot | Copilot and Bing answers lose you |
| Meta | meta-externalagent | meta-externalagent | Meta AI answers lose you |
| DuckDuckGo | (none) | DuckAssistBot | DuckAssist answers lose you |
Google's AI Overviews and AI Mode are served by the same Googlebot that indexes the site. There is no way to stay in Google search and leave Google's AI answers; the "Google-Extended" token only affects Gemini training and grounding.
Which bots does Isonn allow on client sites?
All of them, with the training crawlers left open too. Our reasoning: a small firm's website is not a valuable training corpus, and the tiny principled gain from blocking GPTBot is not worth the risk that a future OpenAI product routes retrieval through it. The robots.txt we ship on isonn.com itself allows these user agents explicitly:
User-agent: *
Allow: /
Disallow: /api/
Disallow: /*?*utm_
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
User-agent: Google-Extended
Allow: /
User-agent: Applebot
Allow: /
User-agent: Applebot-Extended
Allow: /
User-agent: Bingbot
Allow: /
User-agent: DuckAssistBot
Allow: /
User-agent: meta-externalagent
Allow: /
Sitemap: https://isonn.com/sitemap.xml
Explicit allows are redundant when the wildcard already allows everything, but they survive the day someone adds a Disallow under User-agent: * without thinking about AI engines. They also document intent for the next developer.
What should a firm do if it wants to opt out of training but stay citable?
Block the training agents only. This is the policy we recommend to firms whose malpractice carrier or managing partner insists on it:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: meta-externalagent
Disallow: /
Leave OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Googlebot, Google-Extended, Bingbot and Applebot open. The firm stays in every live answer and contributes nothing to model weights.
What else silently blocks AI crawlers?
robots.txt is the visible layer. Three other layers block more firms than robots.txt does, and none of them show up in a rankings report:
- CDN bot management. Cloudflare's "Block AI bots" toggle, enabled by default on some plans since 2025, returns a 403 to every agent in the list above, including the retrieval ones. Check the Security settings, not just robots.txt.
- WordPress security plugins. Several popular plugins ship a "block AI scrapers" rule that matches on the string "GPT" and blocks OAI-SearchBot as collateral.
- JavaScript-only rendering. Most retrieval bots do not execute JavaScript. A site whose practice pages render client-side sends an empty shell to the engine. Server-rendered HTML with the answer in the first 30% of the page is what gets quoted; about 55% of cited passages come from that top portion.
How do you check whether AI bots can reach your site?
- Fetch
https://yourfirm.com/robots.txtand read it as a bot would: the most specificUser-agentblock wins, and aDisallow: /under*applies to every agent not listed separately. - Request a practice page with a retrieval user agent from the command line and confirm you get a 200 and real HTML, not a challenge page:
curl -A "OAI-SearchBot/1.0" -I https://yourfirm.com/family-law/
curl -A "PerplexityBot" -I https://yourfirm.com/family-law/
- Read your server logs for the last thirty days and count hits from
OAI-SearchBot,ChatGPT-UserandPerplexityBot. Zero hits on a site that has been live for months usually means a block somewhere upstream. - Ask the engines directly. Run your own client questions through ChatGPT and Perplexity and see whether your pages ever appear as sources. The free AI visibility check does this across four engines with a written report.
Does allowing crawlers create any legal exposure for a law firm?
Not beyond what publishing the page already created. A public practice page is public. Bar advertising rules (California Rules of Professional Conduct 7.1 to 7.3) govern what the page says, not who reads it. The one practical caution: keep client-confidential material, intake portals and unpublished drafts behind authentication, not behind robots.txt, which is a request rather than a lock.
What to do this week
- Read your robots.txt and your CDN's bot settings; remove any blanket AI block.
- If training opt-out is required, block GPTBot, ClaudeBot, Applebot-Extended and meta-externalagent only.
- Confirm practice pages return server-rendered HTML to
OAI-SearchBotandPerplexityBot. - Then fix what the engines see once they arrive: schema and entity signals, answer-first pages and review velocity. Crawl access is the door; AI search visibility is what happens after it opens.
Check your firm
See whether ChatGPT, Perplexity, Gemini and Google AI Overviews name your firm for your practice area and city.
Free AI visibility check