How Does Claude Find Sources on the Web?
In short
Claude uses two tools. Web search runs queries and returns result pages. Web fetch reads the full content of a specific URL. Claude decides when to use them based on the prompt, and can search several times in one answer.
In claude.ai, users turn on web search from the tools menu. On Team and Enterprise plans an owner must first enable it for the organization. According to Anthropic's help center, when web search is on, Claude can also retrieve content directly from web pages when given a specific URL.
Developers get the same capabilities in the Claude API as two server tools:
- Web search tool. Claude decides when to search, the API runs the searches, and this can repeat several times in one request. Anthropic's docs say Claude searches for current or changing information, specific organizations, people or products, and explicit requests to look something up. Developers can restrict results with
allowed_domainsorblocked_domainsand localize them withuser_location. - Web fetch tool. Retrieves the full text of a web page or PDF. For security it can only fetch URLs that already appeared in the conversation, for example in the user's message or in earlier search results. It does not render JavaScript, and a URL blocked by robots.txt returns a
url_not_allowederror.
The practical consequence: unless a user pastes your URL, Claude usually has to find your page through web search before it can read it in full.
ClaudeBot, Claude-User and Claude-SearchBot: What robots.txt Controls
In short
Anthropic documents three agents. ClaudeBot collects training data, Claude-User fetches pages when users ask Claude questions, and Claude-SearchBot improves search result quality. All honor robots.txt, so each can be allowed or blocked separately.
Anthropic's help center describes the agents and what blocking each one means:
ClaudeBotcollects web content that may contribute to training. Blocking it means your future content should be excluded from training data.Claude-Useraccesses websites when individuals ask Claude questions. Blocking it may reduce your visibility for user-directed web search.Claude-SearchBotnavigates the web to improve search result quality. Blocking it may reduce your visibility and accuracy in user search results.
Anthropic also supports the non-standard Crawl-delay directive, and the rules apply per subdomain. Older guides mention an anthropic-ai user agent; it is not listed in Anthropic's current documentation.
A configuration that opts out of training but stays citable:
User-agent: ClaudeBot Disallow: / User-agent: Claude-User Allow: / User-agent: Claude-SearchBot Allow: /
Talk to the team behind 100+ AI implementations
30-minute discovery call with a senior Alice Labs consultant. No slide deck, no sales pitch — just a scoping conversation.
Book a Discovery CallHow Claude Shows Citations
In short
In claude.ai every web search response includes citations with the source URL and the text Claude relied on. In the API, web search citations are always on and include the URL, title and up to 150 characters of cited text. Web fetch citations are optional.
Anthropic's help center states that every web search response includes citations so users can verify sources, and that citations include the specific text Claude used along with the source URLs.
In the API, each web search citation is a web_search_result_location with url, title and cited_text (up to 150 characters). Search results also carry a page_age field. Developers who show API output directly to end users must include citations to the original source. For web fetch, citations are disabled by default and must be turned on by the developer.
What this means for you: the unit that gets cited is a short passage attached to your URL and page title. A clear title and a passage that makes sense on its own are the parts a reader sees.
Is Claude citing your competitors but not you?
We run your buyers' real questions through Claude with web search on and show which sources it cites, and why yours is or isn't among them.
Request a Claude citation checkWhat We Observed in an OpenAI Scan (Hypothesis for Claude)
In short
In a scan of OpenAI's model with web search, most search phrases from Swedish questions were English, the model used site: operators against official sources, and pages with the question in the title and a short early answer were cited more often. We have not measured Claude.
In September 2026 Alice Labs ran Swedish and English buyer questions through OpenAI's API with web search required and logged every search phrase the model sent. This was OpenAI, not Claude. Three observations:
- Swedish questions were searched in English. 187 of 299 search phrases generated from Swedish questions were in English.
- The model went straight to official sources. It used
site:operators against authorities and vendor documentation, most often imy.se (33 times) and learn.microsoft.com (29 times). - Question-shaped pages won. Pages with the question's wording in the title and a short answer early on were cited more often.
Why we think this likely applies to Claude: Claude also writes its own search queries, can search several times per answer, and cites short passages next to a URL and title. The same mechanics reward pages whose title matches the query and whose opening answers it. This is a hypothesis. We have not run the scan against Claude, and Anthropic does not document its ranking.
How to Check Whether Claude Cites You
In short
Anthropic offers no citation report for site owners. Check access in server logs for Claude-User and Claude-SearchBot, then run your buyers' real questions in claude.ai with web search on and log which URLs are cited.
Two checks cover most of what you can know today:
- Access. Filter server logs for
Claude-UserandClaude-SearchBot. Check that they get 200 responses with the real content. Compare IP addresses with Anthropic's published list at claude.com/crawling/bots.json. - Citation. Run a fixed set of buyer questions in claude.ai with web search on. Log cited URLs and which competitors appear. Repeat after each change to the page.
We run this loop in our AI search work for Swedish companies (in Swedish: AI-sök). If your team uses Claude daily, our Claude training covers web search and source checking.
Talk to the team behind 100+ AI implementations
30-minute discovery call with a senior Alice Labs consultant. No slide deck, no sales pitch — just a scoping conversation.
Book a Discovery CallAbout the Authors & Reviewers

CEO & Co-Founder, Alice Labs
CEO & Co-Founder at Alice Labs. Author of 7 research reports on AI adoption, governance and labor markets cited across EU, OECD and US benchmarks.
- 8+ years in AI strategy & implementation
- Top-5 AI Speaker, Sweden (Mindley 2025)
- 100+ enterprise AI engagements

Co-Founder, Alice Labs
Co-Founder at Alice Labs. Builds AI automation, agent workflows and integration systems that hold up in real business operations.
- AI automation & agent systems lead
- Workflow design across 100+ deployments
- Specialist in RAG, integrations & APIs
Frequently Asked Questions
How does Claude find sources?
Claude uses a web search tool to find pages and a web fetch tool to read a specific URL. In claude.ai both are available when web search is turned on. Claude decides when to search based on the question and can search several times in one answer.
Does blocking ClaudeBot stop Claude from citing my site?
Not necessarily. Anthropic describes ClaudeBot as collecting training data. The agents tied to search are Claude-User and Claude-SearchBot, and Anthropic says blocking those may reduce your visibility in Claude's search answers.
Which search engine does Claude use?
Anthropic does not say in its claude.ai or API documentation which index powers web search or how results are ranked. Its help center says image search is powered by Bing, and its Claude for Government web search connector calls the Brave Search API.
Can Claude read pages built with JavaScript?
Anthropic's API documentation says the web fetch tool does not support websites dynamically rendered with JavaScript. Serve the main content as HTML so it is readable without a browser.
How are citations shown in Claude?
In claude.ai, every web search response includes citations with the source URL and the text Claude used. In the API, web search citations are always on and include the URL, title and up to 150 characters of cited text.
Can I pay to be cited by Claude?
Anthropic does not document any paid placement in Claude's web search answers. The documented levers for site owners are access through robots.txt and readable content.
Does llms.txt or Schema.org markup help with Claude?
Anthropic does not document either as a signal for Claude's web search or citations. They may help other systems, but do not rely on them for Claude. Readable HTML and allowed user agents are the documented requirements.
llms.txt Guide 2026: How to Create & Optimize for AI Crawlers
Next in AI Search & LLMOWhat Is GEO? Generative Engine Optimization Explained
Further reading
- Anthropic — Does Anthropic crawl data from the web?· support.claude.com
- Claude Platform Docs — Web search tool· platform.claude.com
- Claude Platform Docs — Web fetch tool· platform.claude.com
Related services
Related reading
How to Get Cited by ChatGPT
Companion guide for ChatGPT search and OpenAI's crawlers.
15 min howtoHow to Get Cited by Perplexity AI
Tactics for Perplexity's crawler and answer engine.
10 min pillarAI Search Optimization: Complete Guide for 2026
Pillar guide covering ChatGPT, Perplexity, Claude and Google AI features.
14 minSources
- Claude Help Center — Does Anthropic crawl data from the web, and how can site owners block the crawler?(accessed 2026-09-29)
- Claude Help Center — Enable and use web search(accessed 2026-09-29)
- Claude Platform Docs — Web search tool(accessed 2026-09-29)
- Claude Platform Docs — Web fetch tool(accessed 2026-09-29)
- Claude Help Center — MCP: Web Search (Claude for Government)(accessed 2026-09-29)
Next scheduled review: