AI Search & LLMOHow-to GuideFreshLast reviewed: · 8d ago

    How Does Claude Find and Cite Sources, and How Do You Get Cited?

    TL;DR

    Quick Answer
    Cited by AI
    Claude cites pages it finds with web search or reads with web fetch. Allow Claude-User and Claude-SearchBot, serve text without JavaScript, answer early.

    Claude cites the web through two tools: web search and web fetch. Anthropic documents three user agents, how citations are returned and what blocks a fetch. This guide sticks to what is documented and marks the rest as hypothesis.

    Getting cited by Claude means your page appears as a linked source when Claude, the AI assistant built by Anthropic, answers with web search or web fetch. Claude reaches pages through Anthropic's Claude-User and Claude-SearchBot agents, which honor robots.txt, and every web search answer includes citations to source URLs.

    Difficulty

    Intermediate

    Tools

    Your robots.txt and server access logs, curl or another text-only HTTP client, Claude.ai with web search turned on…

    Before you start

    • Access to robots.txt, server logs and page templates
    • A list of the questions your buyers actually ask

    What you'll have at the end

    Pages that Anthropic's agents are allowed to reach, that return readable HTML, and that answer the question early, with a repeatable prompt log to see whether Claude cites them.

    Linus Ingemarsson - Author at Alice Labs
    Written by
    Eric Lundberg - Reviewer at Alice Labs
    Reviewed by
    Published ·Updated
    9 min read

    7-step process

    0/7 complete
    1. Step 1: Allow Claude-User and Claude-SearchBot in robots.txt

      Open yourdomain.com/robots.txt and check every rule for Claude-User, Claude-SearchBot and wildcard user agents. Anthropic says blocking Claude-User may reduce visibility for user-directed web search and blocking Claude-SearchBot may reduce visibility in search results. ClaudeBot controls training only, so you can block it separately if you want.

    2. Step 2: Serve the main content as HTML text

      Anthropic's web fetch tool does not support pages rendered with JavaScript and only handles text, HTML and PDF. Make sure the answer, headings and dates are in the HTML response itself. Test with curl or a text-only fetch and confirm the key paragraphs are present.

    3. Step 3: Put the question in the title and the answer first

      Use the wording people actually ask as the page title and H1, then answer in one or two sentences directly below. API citations quote at most 150 characters, so a short, self-contained answer is the easiest passage to cite. This is our hypothesis, not an Anthropic rule.

    4. Step 4: Cover the English phrasing, not only your local language

      In our OpenAI scan most search phrases generated from Swedish questions were in English. If you serve a non-English market, make sure the page also matches the English terms a model is likely to search for, for example by adding English terminology or an English version.

    5. Step 5: Cite and link the official sources your topic depends on

      Link the regulator, standard or vendor documentation behind each claim. Models tend to search official sources directly, so a page that agrees with and links to them is easier to use alongside them. Keep dates visible, since API search results carry a page_age field.

    6. Step 6: Verify access in your server logs

      Search your access logs for the Claude-User and Claude-SearchBot user agents and check the response codes. Anthropic publishes its crawler IP addresses at claude.com/crawling/bots.json, and warns that IP blocking may not reliably opt you out, so use robots.txt for control.

    7. Step 7: Test real prompts in claude.ai with web search on

      Write the questions your buyers ask, run them in claude.ai with web search turned on, and log which URLs are cited. Repeat after every change. There is no Anthropic tool that reports citations to site owners, so this manual log is your baseline.

    Key Takeaways

    • Claude reaches the live web through two tools: web search (find pages) and web fetch (read a specific URL). In claude.ai both are available when web search is turned on.
    • Anthropic documents three agents: ClaudeBot (training data), Claude-User (fetches pages when users ask Claude questions) and Claude-SearchBot (improves search result quality). All honor robots.txt.
    • Blocking ClaudeBot opts you out of training. Blocking Claude-User or Claude-SearchBot is what Anthropic says may reduce your visibility in Claude's search answers.
    • Web search answers always include citations. In the API each citation carries the URL, title and up to 150 characters of cited text.
    • The web fetch tool does not render JavaScript and can only fetch URLs already in the conversation, so being found by search usually comes first.
    • Anthropic does not document how its web search ranks results. Anything beyond the documented mechanics is hypothesis, including ours.
    01 / 05Step

    How Does Claude Find Sources on the Web?

    In short

    Claude uses two tools. Web search runs queries and returns result pages. Web fetch reads the full content of a specific URL. Claude decides when to use them based on the prompt, and can search several times in one answer.

    In claude.ai, users turn on web search from the tools menu. On Team and Enterprise plans an owner must first enable it for the organization. According to Anthropic's help center, when web search is on, Claude can also retrieve content directly from web pages when given a specific URL.

    Developers get the same capabilities in the Claude API as two server tools:

    • Web search tool. Claude decides when to search, the API runs the searches, and this can repeat several times in one request. Anthropic's docs say Claude searches for current or changing information, specific organizations, people or products, and explicit requests to look something up. Developers can restrict results with allowed_domains or blocked_domains and localize them with user_location.
    • Web fetch tool. Retrieves the full text of a web page or PDF. For security it can only fetch URLs that already appeared in the conversation, for example in the user's message or in earlier search results. It does not render JavaScript, and a URL blocked by robots.txt returns a url_not_allowed error.

    The practical consequence: unless a user pastes your URL, Claude usually has to find your page through web search before it can read it in full.

    02 / 05Step

    ClaudeBot, Claude-User and Claude-SearchBot: What robots.txt Controls

    In short

    Anthropic documents three agents. ClaudeBot collects training data, Claude-User fetches pages when users ask Claude questions, and Claude-SearchBot improves search result quality. All honor robots.txt, so each can be allowed or blocked separately.

    Anthropic's help center describes the agents and what blocking each one means:

    • ClaudeBot collects web content that may contribute to training. Blocking it means your future content should be excluded from training data.
    • Claude-User accesses websites when individuals ask Claude questions. Blocking it may reduce your visibility for user-directed web search.
    • Claude-SearchBot navigates the web to improve search result quality. Blocking it may reduce your visibility and accuracy in user search results.

    Anthropic also supports the non-standard Crawl-delay directive, and the rules apply per subdomain. Older guides mention an anthropic-ai user agent; it is not listed in Anthropic's current documentation.

    A configuration that opts out of training but stays citable:

    User-agent: ClaudeBot
    Disallow: /
    
    User-agent: Claude-User
    Allow: /
    
    User-agent: Claude-SearchBot
    Allow: /
    Linus IngemarssonEric Lundberg
    Alice Labs practitioner team

    Talk to the team behind 100+ AI implementations

    30-minute discovery call with a senior Alice Labs consultant. No slide deck, no sales pitch — just a scoping conversation.

    Book a Discovery Call
    03 / 05Step

    How Claude Shows Citations

    In short

    In claude.ai every web search response includes citations with the source URL and the text Claude relied on. In the API, web search citations are always on and include the URL, title and up to 150 characters of cited text. Web fetch citations are optional.

    Anthropic's help center states that every web search response includes citations so users can verify sources, and that citations include the specific text Claude used along with the source URLs.

    In the API, each web search citation is a web_search_result_location with url, title and cited_text (up to 150 characters). Search results also carry a page_age field. Developers who show API output directly to end users must include citations to the original source. For web fetch, citations are disabled by default and must be turned on by the developer.

    What this means for you: the unit that gets cited is a short passage attached to your URL and page title. A clear title and a passage that makes sense on its own are the parts a reader sees.

    Linus IngemarssonEric Lundberg
    Alice Labs practitioner team

    Is Claude citing your competitors but not you?

    We run your buyers' real questions through Claude with web search on and show which sources it cites, and why yours is or isn't among them.

    Request a Claude citation check
    04 / 05Step

    What We Observed in an OpenAI Scan (Hypothesis for Claude)

    In short

    In a scan of OpenAI's model with web search, most search phrases from Swedish questions were English, the model used site: operators against official sources, and pages with the question in the title and a short early answer were cited more often. We have not measured Claude.

    In September 2026 Alice Labs ran Swedish and English buyer questions through OpenAI's API with web search required and logged every search phrase the model sent. This was OpenAI, not Claude. Three observations:

    1. Swedish questions were searched in English. 187 of 299 search phrases generated from Swedish questions were in English.
    2. The model went straight to official sources. It used site: operators against authorities and vendor documentation, most often imy.se (33 times) and learn.microsoft.com (29 times).
    3. Question-shaped pages won. Pages with the question's wording in the title and a short answer early on were cited more often.

    Why we think this likely applies to Claude: Claude also writes its own search queries, can search several times per answer, and cites short passages next to a URL and title. The same mechanics reward pages whose title matches the query and whose opening answers it. This is a hypothesis. We have not run the scan against Claude, and Anthropic does not document its ranking.

    05 / 05Step

    How to Check Whether Claude Cites You

    In short

    Anthropic offers no citation report for site owners. Check access in server logs for Claude-User and Claude-SearchBot, then run your buyers' real questions in claude.ai with web search on and log which URLs are cited.

    Two checks cover most of what you can know today:

    1. Access. Filter server logs for Claude-User and Claude-SearchBot. Check that they get 200 responses with the real content. Compare IP addresses with Anthropic's published list at claude.com/crawling/bots.json.
    2. Citation. Run a fixed set of buyer questions in claude.ai with web search on. Log cited URLs and which competitors appear. Repeat after each change to the page.

    We run this loop in our AI search work for Swedish companies (in Swedish: AI-sök). If your team uses Claude daily, our Claude training covers web search and source checking.

    Linus IngemarssonEric Lundberg
    Alice Labs practitioner team

    Talk to the team behind 100+ AI implementations

    30-minute discovery call with a senior Alice Labs consultant. No slide deck, no sales pitch — just a scoping conversation.

    Book a Discovery Call

    About the Authors & Reviewers

    Published ·Updated
    Written by
    Linus Ingemarsson - CEO & Co-Founder, Alice Labs at Alice Labs
    Linus Ingemarsson

    CEO & Co-Founder, Alice Labs

    CEO & Co-Founder at Alice Labs. Author of 7 research reports on AI adoption, governance and labor markets cited across EU, OECD and US benchmarks.

    • 8+ years in AI strategy & implementation
    • Top-5 AI Speaker, Sweden (Mindley 2025)
    • 100+ enterprise AI engagements
    Reviewed by
    Eric Lundberg - Co-Founder, Alice Labs at Alice Labs
    Eric Lundberg

    Co-Founder, Alice Labs

    Co-Founder at Alice Labs. Builds AI automation, agent workflows and integration systems that hold up in real business operations.

    • AI automation & agent systems lead
    • Workflow design across 100+ deployments
    • Specialist in RAG, integrations & APIs
    Published · Updated
    Reviewed for technical accuracy, methodology and source integrity.·All claims trace to public sources cited in-line.

    Frequently Asked Questions

    How does Claude find sources?

    Claude uses a web search tool to find pages and a web fetch tool to read a specific URL. In claude.ai both are available when web search is turned on. Claude decides when to search based on the question and can search several times in one answer.

    Does blocking ClaudeBot stop Claude from citing my site?

    Not necessarily. Anthropic describes ClaudeBot as collecting training data. The agents tied to search are Claude-User and Claude-SearchBot, and Anthropic says blocking those may reduce your visibility in Claude's search answers.

    Which search engine does Claude use?

    Anthropic does not say in its claude.ai or API documentation which index powers web search or how results are ranked. Its help center says image search is powered by Bing, and its Claude for Government web search connector calls the Brave Search API.

    Can Claude read pages built with JavaScript?

    Anthropic's API documentation says the web fetch tool does not support websites dynamically rendered with JavaScript. Serve the main content as HTML so it is readable without a browser.

    How are citations shown in Claude?

    In claude.ai, every web search response includes citations with the source URL and the text Claude used. In the API, web search citations are always on and include the URL, title and up to 150 characters of cited text.

    Can I pay to be cited by Claude?

    Anthropic does not document any paid placement in Claude's web search answers. The documented levers for site owners are access through robots.txt and readable content.

    Does llms.txt or Schema.org markup help with Claude?

    Anthropic does not document either as a signal for Claude's web search or citations. They may help other systems, but do not rely on them for Claude. Readable HTML and allowed user agents are the documented requirements.

    Previous in AI Search & LLMO

    llms.txt Guide 2026: How to Create & Optimize for AI Crawlers

    Next in AI Search & LLMO

    What Is GEO? Generative Engine Optimization Explained

    Further reading

    Related services

    Related reading

    Sources

    1. Claude Help Center — Does Anthropic crawl data from the web, and how can site owners block the crawler?(accessed 2026-09-29)
    2. Claude Help Center — Enable and use web search(accessed 2026-09-29)
    3. Claude Platform Docs — Web search tool(accessed 2026-09-29)
    4. Claude Platform Docs — Web fetch tool(accessed 2026-09-29)
    5. Claude Help Center — MCP: Web Search (Claude for Government)(accessed 2026-09-29)

    Next scheduled review:

    Linus IngemarssonEric Lundberg
    Alice Labs practitioner team

    Talk to the team behind 100+ AI implementations

    30-minute discovery call with a senior Alice Labs consultant. No slide deck, no sales pitch — just a scoping conversation.

    Book a Discovery Call
    Share

    Get in Touch!

    The lab usually responds within 24 hours.

    Need help with AI?Get in touch