How ChatGPT Search Finds and Cites Sources
In short
ChatGPT search rewrites a question into web searches, retrieves pages through third-party search providers (OpenAI names Bing) and partner content, and cites the pages it uses. Sites must allow OAI-SearchBot to be shown in its answers.
OpenAI launched ChatGPT search in October 2024. In the launch post, OpenAI said the feature "leverages third-party search providers, as well as content provided directly by our partners". Bing is the provider OpenAI names in its help documentation. OpenAI does not publish a ranking formula, so anything beyond that is inference.
What OpenAI does document is crawler access. It runs three separate agents, each with its own robots.txt setting:
- OAI-SearchBot. Used to surface websites in ChatGPT's search features. OpenAI states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.
- GPTBot. Crawls content that may be used to train OpenAI's foundation models. Disallowing it opts you out of training, not out of search.
- ChatGPT-User. Used for certain user-initiated actions, such as visiting a page a user asks about. OpenAI notes that robots.txt rules may not apply to these requests.
OpenAI says it can take about 24 hours for its systems to pick up a robots.txt change. Crawler policy is usually the first thing an AI search optimization consultant checks before any content work begins.
What We Observed in 618 ChatGPT Search Queries
In short
In an Alice Labs scan of 200 runs in September 2026, the model sent 618 search queries. Most queries from Swedish questions were in English, it targeted official sources with site: operators, ran about 3 searches per answer, and favored pages with the question in the title and a short answer early.
Most advice on ChatGPT citation is guesswork, because OpenAI does not publish how sources are chosen. So we measured what the model does before it cites anything: the search queries it sends.
Method
- When: 23-25 September 2026.
- Setup: OpenAI Responses API, model
chat-latest, web search required on every run, user country set to Sweden. - Prompts: 40 buyer-style questions (half Swedish, half English) across six business needs, each run 5 times, for 200 runs in total.
- What we logged: the 618 search queries the model itself issued, the pages it fetched and the sources it cited.
What we saw
| Observation | What the data showed |
|---|---|
| Language of queries from Swedish questions | 187 of 299 were in English, often with "Sweden" added |
| site: operators against official sources | imy.se 33, learn.microsoft.com 29, fortnox.se 19, digg.se 17, openai.com 12 |
| Searches per answer | About 3 on average |
| Queries containing a year | 11% |
| Pages cited more often | Pages with the question's wording in the title and a short answer early, such as dedicated "GEO agency" and "AI for leadership teams" pages |
| Source mix in answers | Usually a vendor paired with an official source |
| Brand searches | The model searched known brands' domains directly, e.g. site:knowit.se and site:afry.com |
What it means in practice
- Match queries in more than one language. If the model turns a Swedish question into an English search, a page that only uses Swedish terms is invisible to that search. Use the English terms buyers and the model use, and frame English pages for the local market.
- Cite official sources. The model looks for regulators and vendor documentation on its own. Pages that link to those same sources fit the answer it is building.
- Put the question in the title. One page per high-value question beats one broad page that covers ten.
- Answer first. A short, direct answer near the top is easy to lift and attribute. Detail can follow.
- Build entity strength. The model searched well-known brands by domain. If it does not know your brand, it will not search for you by name, so you depend entirely on matching generic queries.
Entity Clarity: Make Your Brand Searchable by Name
In short
Entity clarity means making it obvious what your page defines, who wrote it and which organization publishes it. In our scan the model searched known brands' domains directly, so a strong entity gets searched for by name.
Language models work with entities: people, products, concepts and organizations. When the model already knows a brand, it can search that brand's domain directly, which is what we saw with site:knowit.se and site:afry.com. A lesser-known brand has to win on generic queries instead.
To establish entity clarity, do three things:
- Open every key page with a definition. State what the entity is in 40-50 words, in the first paragraph. Make it self-contained and extractable.
- Use consistent naming. If your product is called "Acme Analytics," don't alternate between "our platform," "the tool," and "Acme." Consistency helps the model link mentions to the same entity.
- Cross-link entity signals. Your Schema.org Organization markup, your About page, your LinkedIn profiles, and your Wikidata entry (if applicable) should all describe the same entity the same way.
Schema.org Markup: Useful for Clarity, Not a Shortcut
In short
Structured data helps machines classify your content, author and dates. OpenAI does not document schema as a ChatGPT ranking factor, and Google says no special markup is needed for its AI features. Use Article, FAQPage, HowTo, Organization and Person in JSON-LD where they fit.
Structured data will not get a weak page cited. What it does is remove ambiguity about what a page is, who wrote it and when it was updated. These are the types worth having:
- Article / NewsArticle. Declares your content type, headline, author, datePublished, and dateModified. This is the baseline.
- FAQPage. Marks up question-answer pairs. Useful because FAQ answers are short and self-contained, which is the format the model quotes.
- HowTo. Step-by-step instructions with name and text for procedural content.
- Organization + Person. Establishes who publishes and who writes, with sameAs links to official profiles.
Use JSON-LD and validate before deploying. Malformed schema adds noise instead of clarity.
Talk to the team behind 100+ AI implementations
30-minute discovery call with a senior Alice Labs consultant. No slide deck, no sales pitch — just a scoping conversation.
Book a Discovery CallWriting Citation-Rich Content That LLMs Prefer
In short
Research by Aggarwal et al. (2024) found that adding citations, statistics and quotations improved visibility in generative engines by up to 40%. In our scan the model also paired vendors with official sources, so linking to those sources fits how answers are built.
The GEO paper (Aggarwal et al., 2024, arXiv:2311.09735) tested nine optimization strategies across generative engines. Three produced the largest visibility gains:
- Cite sources inline. Name the study, author, or organization behind every claim. A claim with a named source is easier to trust and quote than a vague one.
- Include specific statistics. Numbers with named sources give the model something concrete to extract.
- Add quotations from credible sources. Quoting a regulator, standard or expert adds authority the model can attribute.
Our own scan points the same way. The model searched regulators and vendor documentation directly and usually paired a vendor with an official source. A page that already links to that official source makes the pairing easy.
SparkToro's 2024 zero-click study found that roughly 60% of Google searches ended without a click to any external website. Being the source that gets cited, even without a click, is part of visibility now.
Freshness and llms.txt: One Matters, One Is Optional
In short
11% of the model's search queries in our scan contained a year, so visible dates matter for time-sensitive topics. llms.txt is a proposed standard that neither OpenAI nor Google says it uses; it is cheap but optional.
Freshness showed up directly in the data: 11% of the search queries the model sent contained a year. For those queries, an undated or visibly old page is a weak match. Every key page should have:
- A visible "Last updated" date on the page itself.
datePublishedanddateModifiedin Article schema.- A regular review cycle, so statistics, links and examples stay correct. Only change the date when the content actually changes.
The llms.txt file, proposed by Jeremy Howard (Answer.AI) in September 2024, sits at your domain root and gives LLMs a curated overview of your site. Follow the format at llmstxt.org: H1 site name, blockquote summary, H2 sections with bulleted links. OpenAI has not said ChatGPT search uses it, and Google states that no AI text files are needed for its AI features.
Is ChatGPT citing your competitors but not you?
We run a 30-prompt ChatGPT citation audit across your top keywords to show exactly where your brand is — and isn't — getting cited today.
Request ChatGPT citation auditHow to Measure Whether ChatGPT Is Citing You
In short
Measure ChatGPT citation with repeated prompt audits and analytics filtered on utm_source=chatgpt.com. Bing's AI Performance report shows Copilot and partner citations, not ChatGPT.
OpenAI offers no search-console-style report for ChatGPT search. Measurement today is a three-part stack:
- Repeated prompt audits. Write the questions your buyers ask and run each one several times with web search on. Log the cited sources, your position and the competitors next to you. If you use the API, also log the search queries the model sends; that is where the insights in this article came from.
- Analytics referrals. OpenAI's publisher FAQ says ChatGPT referral links include
utm_source=chatgpt.com. Filter on it in GA4 under Acquisition > Traffic Acquisition. - Bing AI Performance. Bing Webmaster Tools shows how often Microsoft Copilot and partner experiences cite your pages, and the grounding queries behind those citations. It is a useful signal for AI search in general, but it is not a ChatGPT report.
Set a baseline before making changes, then rerun the same prompts after each change. Single runs vary, so compare across repeated runs.
12-Point Checklist for Getting Cited by ChatGPT
In short
The checklist combines what OpenAI documents (OAI-SearchBot access, third-party search providers, referral tracking) with what we observed in 618 search queries (language, official sources, question-shaped titles, answer-first pages, entity strength).
Items 1 and 7 are about access and come from official documentation. Items 2 to 6 and 8 come from our own scan. The rest are standard practice that makes the others easier to verify.
- OAI-SearchBot allowed. Without it, OpenAI says you will not be shown in ChatGPT search answers.
- Question in the title. Use the words buyers and the model use, one question per page.
- Answer in the first paragraph. 40-60 words that stand on their own.
- Bilingual matching. Include English terms on local-language pages and local framing on English pages.
- Official sources linked. The regulator, standard or vendor documentation your topic depends on.
- Entity strength. Consistent brand name, Organization schema with sameAs, a real About page.
- Bing indexing. Verified in Bing Webmaster Tools, sitemap submitted, updates pushed with IndexNow.
- Visible dates. For the share of queries that include a year.
- Clean JSON-LD. Article, FAQPage, HowTo, Organization and Person, validated.
- Named author and reviewer. Role, profile link and topical expertise.
- Third-party mentions. Credible pages in your category that name you.
- Measurement loop. Repeated prompt audits plus utm_source=chatgpt.com in analytics.
For a deeper implementation walkthrough, see our LLM citation optimization guide, or engage Alice Labs directly for AI SEO consulting. Swedish readers can read about our GEO and AI search service in Swedish. Teams staffing the citation loop internally can compare software from our directory of ai content optimization tools.
ChatGPT Search vs Training Data: How They Differ
In short
ChatGPT search retrieves live pages through search providers and OAI-SearchBot and shows linked sources. Training data is collected ahead of time (GPTBot is OpenAI's training crawler) and is not cited. Search responds to changes you make now; training data does not.
ChatGPT can answer from two very different substrates, and they respond to different levers.
| Dimension | ChatGPT search | Training data |
|---|---|---|
| Source | Third-party search providers (Bing named), OAI-SearchBot, partner content | Content collected before training, including by GPTBot, plus licensed content |
| Freshness | Live searches at answer time | Fixed until the next model is trained |
| Citations shown | Yes, linked sources | No |
| Crawler control | OAI-SearchBot in robots.txt | GPTBot in robots.txt |
| What you can influence | Access, titles, answers, language, sources, dates, indexing | Long-term brand presence across the open web |
Spend most of your effort on search, because it responds to changes you make now. Training data still matters indirectly: a brand the model already knows is one it can search for by name, which is what we saw with well-known consultancies in our scan.
Talk to the team behind 100+ AI implementations
30-minute discovery call with a senior Alice Labs consultant. No slide deck, no sales pitch — just a scoping conversation.
Book a Discovery CallPublisher Partners vs Organic Citation
In short
OpenAI launched ChatGPT search with content partners including the Associated Press, Axel Springer, Condé Nast, Financial Times, News Corp and Reuters. OpenAI does not document how partner content is weighted against organic results, and B2B vendors can compete on organic pages.
At launch in October 2024, OpenAI named partners including the Associated Press, Axel Springer, Condé Nast, Dotdash Meredith, Financial Times, GEDI, Hearst, Le Monde, News Corp, Prisa (El País), Reuters, The Atlantic, Time and Vox Media. These are licensing and content agreements, and OpenAI says ChatGPT search uses content provided directly by partners alongside third-party search.
OpenAI does not publish how much weight partner content gets. What we can say from our own scan is that for B2B buyer questions, the model built answers from vendor pages and official sources such as regulators and vendor documentation, not from news archives.
Actionable read: if you sell into news-adjacent categories, a partner outlet mentioning you may matter. If you sell B2B or specialist services, focus on the organic checklist above.
About the Authors & Reviewers

CEO & Co-Founder, Alice Labs
CEO & Co-Founder at Alice Labs. Author of 7 research reports on AI adoption, governance and labor markets cited across EU, OECD and US benchmarks.
- 8+ years in AI strategy & implementation
- Top-5 AI Speaker, Sweden (Mindley 2025)
- 100+ enterprise AI engagements

Co-Founder, Alice Labs
Co-Founder at Alice Labs. Builds AI automation, agent workflows and integration systems that hold up in real business operations.
- AI automation & agent systems lead
- Workflow design across 100+ deployments
- Specialist in RAG, integrations & APIs
Frequently Asked Questions
How do you get cited by ChatGPT?
Allow OAI-SearchBot in robots.txt, get indexed in Bing, and write one page per buyer question with the question in the title and a short answer first. Link official sources, match the English queries the model sends, and build a consistent brand entity.
What does ChatGPT search for when it answers a question?
In our September 2026 scan the model sent 618 search queries across 200 runs, about 3 per answer. It often used site: operators against official sources such as imy.se, learn.microsoft.com and openai.com, searched known brands' domains directly, and 11% of its queries contained a year.
Does ChatGPT search in English when asked in Swedish?
Often, yes. In our scan, 187 of 299 search queries generated from Swedish questions were in English, frequently with "Sweden" added. Swedish pages should include the English terms buyers use, and English pages should mention Sweden. The scan used the API, not the ChatGPT app.
How does ChatGPT pick sources?
OpenAI says ChatGPT search uses third-party search providers, with Bing named, plus content from partners. It does not publish ranking factors. In our scan the model favored pages with the question in the title and a short early answer, and usually paired a vendor with an official source.
Does ChatGPT crawl my site?
OpenAI runs OAI-SearchBot for ChatGPT search, GPTBot for training and ChatGPT-User for certain user-initiated visits. OAI-SearchBot and GPTBot follow their own robots.txt rules; OpenAI notes robots.txt may not apply to ChatGPT-User. Changes take about 24 hours to take effect.
Do I need to be in the Bing index to be cited by ChatGPT?
OpenAI names Bing among the third-party search providers behind ChatGPT search, so Bing indexing is a sensible baseline. OpenAI does not say it is strictly required. Verify your site in Bing Webmaster Tools, submit a sitemap and use IndexNow for new URLs.
What is the best schema.org markup for ChatGPT?
Use JSON-LD for Article, FAQPage, HowTo where relevant, and Organization plus Person with sameAs links. OpenAI does not document schema as a ranking factor, so treat it as clarity for machines rather than a shortcut to citations.
How long until ChatGPT cites a new page?
There is no official timeline. OpenAI says robots.txt changes take about 24 hours to apply, and the page must be findable through search first. Rerun the same prompts after publishing and after each change to see when it starts being cited.
Do backlinks help with ChatGPT citations?
Indirectly. ChatGPT search relies on search providers, and links and mentions help pages rank and get found there. Mentions on credible third-party pages in your category also give the model more places to find and confirm your brand.
What does the OAI-SearchBot user agent do?
OAI-SearchBot is the OpenAI crawler used to surface websites in ChatGPT search. OpenAI states that sites opted out of it will not be shown in ChatGPT search answers. It is separate from GPTBot, which collects training data, so you can allow one and block the other.
Do OpenAI publisher deals matter for organic citation?
OpenAI works with content partners such as the Associated Press, Axel Springer, Financial Times, News Corp and Reuters, but does not document how partner content is weighted. For B2B buyer questions in our scan, answers were built from vendor pages and official sources.
Can I opt out of ChatGPT crawling?
Yes. Disallow OAI-SearchBot to stay out of ChatGPT search answers and GPTBot to opt out of training. OpenAI notes that robots.txt rules may not apply to ChatGPT-User, which handles user-initiated visits. Blocking OAI-SearchBot means no ChatGPT search citations.
What Is LLMO? Large Language Model Optimization (2026)
Next in AI Search & LLMOAI Search Engine Market Share 2026: ChatGPT, Perplexity, Google & Microsoft
Further reading
- OpenAI: Introducing ChatGPT search (Oct 2024)· openai.com
- OpenAI crawlers documentation (OAI-SearchBot, GPTBot, ChatGPT-User)· developers.openai.com
- OpenAI Help: Publishers and Developers FAQ· help.openai.com
- Google Search Central: AI features and your website· developers.google.com
- Bing Webmaster Tools: AI Performance· bing.com
- GEO: Generative Engine Optimization (Aggarwal et al., 2024)· arxiv.org
Related services
Related reading
What Is LLMO? Large Language Model Optimization Explained
Glossary definition of LLMO — the overarching discipline behind ChatGPT citation optimization.
7 min comparisonGEO vs SEO: What's the Difference?
Side-by-side comparison of generative engine optimization and search engine optimization.
8 min pillarAI Search Optimization: Complete Guide for 2026
Full playbook covering ChatGPT, Perplexity, Claude, and Google AI Overviews.
14 minSources
- Alice Labs LLM scan: 618 model-issued search queries, OpenAI Responses API, country SE (23-25 Sep 2026)(accessed 2026-09-29)
- OpenAI: Introducing ChatGPT search(accessed 2026-09-29)
- OpenAI crawlers documentation (OAI-SearchBot, GPTBot, ChatGPT-User)(accessed 2026-09-29)
- OpenAI Help Center: Publishers and Developers FAQ(accessed 2026-09-29)
- Google Search Central: AI features and your website(accessed 2026-09-29)
- Bing Webmaster Tools: AI Performance(accessed 2026-09-29)
- Aggarwal et al.: GEO: Generative Engine Optimization (arXiv:2311.09735)(accessed 2026-09-29)
- llms.txt proposal (Answer.AI)(accessed 2026-09-29)
Next scheduled review: