ChatGPT chooses which sources to cite in 2026 through two gates: a retrieval gate, where it rewrites your prompt into fan out queries and pulls candidates from Bing’s index plus OpenAI’s own OAI-SearchBot index, and a selection gate, where it scores each candidate’s title, URL, and snippet against those fan out queries before opening the page and lifting a passage. Ahrefs’ April 2026 study of 1.4 million prompts found about 50 percent of retrieved URLs get cited, 88 percent of cited URLs arrive through the general search channel, and cited titles score 0.656 in similarity to ChatGPT’s fan out queries versus 0.484 for non cited pages. Seer Interactive measured the other half: 87 percent of ChatGPT search citations match Bing’s top 10. Page A beats page B when it ranks in Bing for the sub question ChatGPT wrote and its title and opening passage answer that sub question in plain words.
The pipeline overview is in how ChatGPT search works and the per answer counts are in how many sources ChatGPT cites. This post answers the question Semrush, Profound, SE Ranking, BrightEdge, and 5W Public Relations keep measuring: why one page wins the slot.
What are the two gates a page passes through before ChatGPT cites it?
A page must be retrieved first, then selected, and the two stages run on different data. Retrieval decides whether ChatGPT ever sees your URL. Selection decides whether it opens the page and lifts a passage. Research by Dan Petrovic at Dejan Marketing documented that each retrieved result reaches the model as a bundle: title, short snippet, URL, and an ID number. ChatGPT reads that bundle, not your page, when it decides what to open.
Across Ahrefs’ 1.4 million prompts, ChatGPT pulled an average of 16.57 cited and 16.58 non cited URLs per prompt, and the channel a URL arrives through matters most. Ahrefs found five internal retrieval channels, labeled by a field called ref_type, with very different citation rates: search at 88.46 percent, news at 12.01 percent, Reddit at 1.93 percent, YouTube at 0.51 percent, and academia at 0.40 percent. Enter only through a supplementary feed and you are context, not a citation.
How do fan out queries decide which pages get retrieved?
ChatGPT retrieves against the sub questions it writes, not the prompt the user typed, so a page wins retrieval by matching a fan out query it never saw. “Best personal injury lawyer in Charleston” becomes rewrites like how to evaluate an injury attorney, Charleston firms with jury verdicts, and contingency fee ranges in South Carolina. Each goes to Bing and OpenAI’s own index in parallel.
Ahrefs measured how tightly title relevance tracks citation. Cosine similarity between the prompt and a cited page’s title averaged 0.602, against 0.484 for non cited pages; measured against the best matching fan out query, cited titles rose to 0.656. Slugs mattered too: results with natural language URL slugs were cited 89.78 percent versus 81.11 percent for opaque ones. Ahrefs’ Brand Radar exposes those fan out queries for any prompt, and that list is the real keyword set.
Want to know which of your pages ChatGPT is already citing, and for which fan out questions? Get your free AI visibility audit and we will map every URL on your site that is winning or losing a citation today.
Which signals decide page A over page B once both are retrieved?
Six signals explain most of the selection gap, and they stack.
1. Retrieval channel and Bing rank
The strongest single predictor is whether the page sits in Bing’s top 10 for the fan out query. Seer Interactive’s analysis of more than 500 ChatGPT search citations found 87 percent matched Bing’s top 10 organic results, while only 56 percent matched Google’s. Ahrefs found the reverse overlap is tiny: only 6.82 percent of ChatGPT cited pages appear in Google’s top 10 for the equivalent query.
2. Title and URL match to the fan out query
ChatGPT scores the title and URL against its sub questions before it opens anything, so a clever headline loses to a plain one that restates the sub question. “What a Charleston personal injury lawyer costs in 2026” beats “Justice starts here” because the model never reads the second page.
3. Domain authority and referring domains
SE Ranking’s study of 129,000 domains and 216,524 pages across 20 niches found referring domain count at the top of the list of ChatGPT citation drivers, with high traffic domains earning about three times more citations than low traffic ones. This is the tiebreaker: when two pages match the same fan out query, the domain more independent sites link to wins.
4. Freshness, with a twist
Ahrefs’ study of 17 million citations found ChatGPT cited URLs 458 days newer on average than Google’s organic results, the strongest freshness preference of any engine tested. But inside a single prompt’s retrieval set, the median cited page was about 500 days old, and the freshest pages were discarded most often. Freshness gets you into the pool; relevance and authority get you picked. In the news channel, cited pages had a median age near 200 days against 300 for non cited.
5. Extractable structure
Once ChatGPT decides to open a page, it abandons the snippet and reads the full document, per David McSweeney’s research at QueryBurst. From there, structure decides which passage gets lifted. SE Ranking found pages with sections of 120 to 180 words between headings earned about 70 percent more ChatGPT citations than pages with sections under 50 words. Question form headings, a direct answer under each, and tables for comparisons give the model a clean block to cite.
6. Corroboration and entity clarity
The model prefers claims it can see repeated elsewhere. A firm whose name, practice areas, and numbers appear consistently across its site, Avvo, LinkedIn, a trade outlet, and a press placement reads as a resolved entity. A firm whose only evidence is its own homepage reads as an unverified claim, and earned media is what moves this signal.
Why do Wikipedia and Reddit get so much weight, and why did it change in September 2025?
Wikipedia and Reddit are still the two most cited domains in ChatGPT, but their share collapsed in September 2025. Semrush’s 13 week study of 230,000 prompts and more than 100 million citations found ChatGPT cited Reddit in close to 60 percent of responses in early August 2025, then around 10 percent by mid September. Wikipedia fell from roughly 55 percent of responses to under 20 percent. Neither domain moved much on Google AI Mode or Perplexity, which points at an OpenAI side change; Semrush’s Sergei Rogulin attributed it to an effort to stop over citing a few domains.
After the reset, 5W Public Relations’ January to February 2026 research found Wikipedia at 13.15 percent and Reddit at 11.97 percent of all US ChatGPT citations, over a quarter combined. Profound’s 680 million citation dataset puts Wikipedia at 7.8 percent of total citations and 47.9 percent of top 10 source share. The Ahrefs ref_type data adds the uncomfortable detail: Reddit made up 67.8 percent of all non cited URLs, cited at only 1.93 percent through its dedicated channel.
Do OpenAI’s publisher licensing deals change which sources get cited?
They change the news channel, not the general search channel, and the news channel cites only 12 percent of what it retrieves. As of June 2026, OpenAI held roughly 20 verified news publisher partnerships covering more than 160 outlets, per LLM Pulse’s deal tracker, plus a platform data deal with Reddit signed in May 2024. The publishers include Axel Springer (Politico, Bild), the Associated Press, News Corp (Wall Street Journal, Barron’s, MarketWatch), the Financial Times, Condé Nast (Wired, The New Yorker, Vogue), Vox Media, and The Atlantic. News Corp’s deal is worth up to $250 million over five years; Axel Springer’s runs about $13 million a year.
The deals guarantee attribution and crawl access. They do not override the selection signals above. 5W Public Relations found the Wall Street Journal, New York Times, and Bloomberg absent from the top 20 ChatGPT cited domains in early 2026, despite two of the three holding licenses, because paywalls and structure limit what is extractable. Semrush’s post September 2025 data shows the biggest ChatGPT gainers were PRNewswire, Forbes, and Medium, and LinkedIn moved from eleventh to fifth most cited domain, appearing in 14.3 percent of ChatGPT search responses. Open, syndicated content beat licensed but locked content, which for a law firm or surgical practice favors press placements the engines can read over a bylined piece behind a paywall.
How does ChatGPT decide citation order and which passage it lifts?
Citation order follows the order in which claims appear in the generated answer, not a ranked list of sources. The model drafts the response, attaches a source to each claim it grounded, and numbers them as they occur. The page that supplies the opening definition gets citation one. No step awards position one to the most authoritative domain, so an early citation means owning the sub question the model answers first, usually the definition or the direct yes or no.
Passage selection works the same way at page level: after opening the document, the model looks for the block that states the claim it needs, and citation studies across 2025 and 2026 consistently find the first third of a page supplies most lifted passages. A page with the answer in its first 100 words gets cited for the answer; a page with 600 words of context first gets cited for nothing. Semrush’s January to June 2026 tracking of 1,094 buyer categories, covering 600,000 citations and 50,000 brands, found only 15.2 percent have a clear brand owner, so the first slot in most niches is still open.
What gets a page dropped, and do robots.txt and llms.txt matter?
Pages get dropped for access failures before they are ever judged on quality, and robots.txt is where most of those failures start. OpenAI runs three agents with separate jobs. OAI-SearchBot builds the index ChatGPT search retrieves from; block it and you exit the retrieval gate. ChatGPT-User fetches a page live when a conversation calls for it. GPTBot collects training data and has no effect on search citations. Sites that blocked “all AI bots” in a 2024 or 2025 cleanup often blocked OAI-SearchBot by accident, and CDN level bot protection does the same thing silently. Ahrefs’ analysis of roughly 140 million websites found GPTBot the most blocked AI crawler, and many of those rules catch the search agent too. Our guide to whether ChatGPT can see your website walks through the check.
llms.txt does not help. No major provider, OpenAI included, has committed to reading the file in production, and OpenAI’s publisher guidance points to allowing OAI-SearchBot instead. One 2026 analysis of 515 million LLM bot requests found the share touching /llms.txt negligible, and a separate audit found 97 percent of existing llms.txt files received zero requests in a month.
Beyond access, the common drop reasons are all upstream of quality: not indexed in Bing, client side rendering that leaves the crawler an empty shell, a title that resembles no fan out query, a new page with no referring domains, and claims a stronger domain already covers in the same retrieval set.
Frequently asked questions
Does ChatGPT use Google to choose sources?
No. ChatGPT search retrieves from Bing’s index and OpenAI’s own OAI-SearchBot index, with no Google integration. Seer Interactive found 87 percent of ChatGPT search citations match Bing’s top 10 while only 56 percent match Google’s, and Ahrefs found only 6.82 percent of ChatGPT cited pages appear in Google’s top 10 for the equivalent query.
Does domain authority matter for ChatGPT citations?
Yes, as a tiebreaker once relevance is established. SE Ranking’s study of 129,000 domains found referring domain count at the top of the list of citation drivers, with high traffic domains earning about three times more ChatGPT citations than low traffic ones. But Ahrefs’ 1.4 million prompt data shows title relevance to fan out queries decides retrieval first; authority only helps a page that already matches the sub question.
Does fresh content get cited more by ChatGPT?
Across the web, yes: Ahrefs found ChatGPT cited URLs 458 days newer on average than Google’s organic results. Inside a single answer’s retrieval set, the median cited page was about 500 days old and the freshest candidates were discarded most often. Freshness earns retrieval; relevance and authority earn the citation. The news channel is the exception, where recency breaks ties.
Why does ChatGPT cite Wikipedia and Reddit so much?
Both are open, heavily linked, and structured for extraction, and Reddit has held a licensing deal with OpenAI since May 2024. 5W Public Relations found Wikipedia at 13.15 percent and Reddit at 11.97 percent of US ChatGPT citations in early 2026. Their share fell hard in September 2025, when Semrush recorded Reddit dropping from about 60 percent of responses to around 10 percent, but they remain the two most cited domains.
Do OpenAI’s licensing deals with publishers guarantee citations?
No. They guarantee attribution and crawl access, and licensed outlets do better in the news channel, but Ahrefs found that channel cites only 12.01 percent of what it retrieves. 5W Public Relations found the Wall Street Journal, New York Times, and Bloomberg absent from the top 20 ChatGPT cited domains in early 2026 despite licensing. PRNewswire, Forbes, Medium, and LinkedIn gained the most share after September 2025.
Does llms.txt help ChatGPT find and cite my pages?
No measurable effect. OpenAI has not committed to reading llms.txt in production and its publisher guidance points to allowing OAI-SearchBot. A 2026 analysis of 515 million LLM bot requests found a negligible share touching /llms.txt, and a separate audit found 97 percent of existing files got zero requests in a month. Spend the effort on Bing indexation, question form titles, and referring domains.
The bottom line
ChatGPT does not rank sources; it matches them, then trusts them. A page wins the citation when its title and URL echo a sub question the model wrote, when Bing already ranks it for that sub question, when independent domains vouch for it, and when the answer sits in the first block of text. Domain size, licensing deals, and llms.txt files do not override any of that, which is why Semrush found 84.8 percent of buyer categories still had no clear owner in ChatGPT halfway through 2026. The firms taking those slots know which of their pages clear both gates.
If you are not sure which of your pages ChatGPT is already citing and which are being retrieved and discarded, request your free AI visibility audit. We run your priority queries, list every URL that earns a citation, and show you the fan out questions where a competitor’s page is taking the slot instead.
Tagged