Which Sources Do AI Engines Cite for Legal Queries? What the Published Studies Say
The studies agree on the shape: a small set of domains earns most citations, earned mentions matter more than links, most cited passages sit near the top of a page, and the engines differ sharply in how often they name a brand at all. What Ahrefs, Soar, 5W, Digital Applied and others found, what it implies for a Los Angeles law firm, and what our own Los Angeles logs add.
Short answer: The published research on AI citations, across engines and industries, converges on five findings. Citations concentrate on a small set of domains. Earned mentions on third-party sites correlate with being named far more than backlinks do. Most quoted passages come from the first third of a page. Being on page one of Google is neither necessary nor sufficient. And ChatGPT names brands rarely while Perplexity names them freely. For legal queries in Los Angeles, our own logs match all five, with the trusted domains being the courts, the publishers, the directories, the local press and the bar associations. The implication for a firm is consistent: be present in those sources, and put the answer at the top of the page.
What have the studies found?
| Study | Finding | What it means for a firm |
|---|---|---|
| Ahrefs, 75,000 brands | Correlation with AI Overview brand mentions: YouTube presence 0.74, unlinked web mentions 0.66, backlinks 0.22 | Being talked about matters more than being linked to |
| 5W / Leapd | Roughly 68% of citations come from about 14 domains | The trusted set is small; presence in it is the work |
| Soar | Earned media accounts for 82–89% of citations; 38% of AI Overview citations come from the organic top ten | Third-party pages dominate; rankings help but are not the mechanism |
| Digital Applied | 29.8% of domains cited in AI Overviews are not on page one | A page can be cited without ranking, if it answers |
| Multiple passage analyses | Around 55% of cited passages sit in the first 30% of the page | The 40-to-60-word answer at the top is what gets quoted |
| 5W | ChatGPT cites brands in ~0.6% of answers; Perplexity ~13% | Engine choice changes the baseline entirely |
| 137,000-site study | 97% of llms.txt files received zero bot traffic | The file is harmless and does nothing |
| iLawyer Marketing, Aug 2026 (n=1,110) | 41.9% of consumers would use ChatGPT to research a lawyer, up from 9% in 2023; Google 71.9%; heaviest AI adopters are 45–60 | The estate, elder and business demographic is the one asking AI |
Figures as published by each source; methods differ and none is legal-specific except the last.
What do the Los Angeles logs add?
Our fixed query list, run monthly since 2024 across five surfaces, matches the shape. In words, because the sample is ours and not a study:
- The trusted set for legal questions is the California Courts self-help site, Nolo and similar publishers, Avvo, Justia and FindLaw, the local press, the bar associations, and Reddit (for Perplexity). Firm pages are cited when they are the clearest source for one specific question.
- Courthouse and process pages are the firm pages cited most often, followed by cost pages with figures. Overview and "about" pages are not cited.
- Gemini names firms from Maps; Perplexity from directories and Yelp; ChatGPT rarely, from directories; the Google surfaces from the organic and map results.
- The cited passage is one to three sentences, near the top of a section, with a number or rule in it.
- Run-to-run variance is large: about a third of repeat runs change the named firms. Any single-run claim is noise.
What does this imply, practically?
- Presence in the trusted set. Complete, consistent directory profiles; local press and bar mentions earned honestly; a Yelp page; a defensible Reddit footprint. This is the "earned mentions" finding applied.
- Pages that answer one question at the top. The passage-position finding applied. Courthouse, process and cost pages, with the answer first.
- Reviews with velocity and text. The Gemini and map-pack mechanism.
- Entity consistency. So the engine is sure which firm it is naming.
- Measurement by engine. Because a firm can be doing well on Perplexity and Gemini and absent from ChatGPT, and the fix differs.
- No time on tricks. llms.txt, hidden text, "AI-optimized" plugins. None appears in any finding.
What do the studies not settle?
- Whether any of the correlations is causal. Ahrefs and the others say so themselves.
- Legal-specific citation behavior at scale. Nobody has published a law-firm-only dataset of any size; ours is a Los Angeles sample. We will publish our own quarterly index once the run is large enough to support percentages, and we will publish the method with it.
- How the ad products change the answer surface. Too new.
What does not work?
- Reading a 0.74 correlation as "make YouTube videos". It is a signal that being discussed elsewhere matters; a firm without a reason to be on YouTube should not manufacture one.
- Reading the 14-domain finding as "get on those 14". For legal, the trusted set is different and local. Find it in the source-share tab of the firm's own run.
- Treating any of these numbers as a target. They describe the environment, not the firm.
The methodology page documents how we run the Los Angeles list; the AI search visibility guide walks through each engine; and the free visibility check runs the list for a firm's practice and cities and reports who is cited, which is the only version of these findings that applies to one firm.
Check your firm
See whether ChatGPT, Perplexity, Gemini and Google AI Overviews name your firm for your practice area and city.
Free AI visibility check