AI Search & LLMOHow-to GuideFreshLast reviewed: · 60d ago

    The GEO Audit Checklist: 12 Categories for AI Search Readiness

    TL;DR

    Quick Answer
    Cited by AI
    A GEO audit reviews 12 categories: schema markup, entity signals, llms.txt, robots.txt, citation-rich content, FAQ structure, author schema, freshness, off-site citations, brand monitoring, knowledge panel, and quick-answer formatting.

    Before you optimize for ChatGPT, Perplexity, Claude, and Google AI Overviews, run a structured audit. This checklist covers the 12 categories that determine whether AI engines can crawl, understand, and cite your site today.

    A GEO audit (generative engine optimization audit) is a structured review of how prepared a website is to be cited by AI search engines like ChatGPT, Perplexity, Claude, and Google AI Overviews. It evaluates 12 categories spanning on-page schema, entity signals, crawler access, citation patterns, and off-site authority — producing a prioritized backlog of fixes.

    Time

    1-2 days for audit, 4-12 weeks for fixes

    Difficulty

    Intermediate

    Tools

    Schema.org validator + Google Rich Results Test, Screaming Frog (or similar crawler), Ahrefs / Mention (off-site citation discovery)…

    Before you start

    • Access to your site's CMS, robots.txt, and analytics
    • Familiarity with JSON-LD and Schema.org
    • Ability to run a full-site crawl

    What you'll have at the end

    A scored 12-category audit report with a prioritized fix backlog, owner assignments, and a 90-day implementation roadmap to make your site GEO-ready.

    Linus Ingemarsson - Author at Alice Labs
    Written by
    Eric Lundberg - Reviewer at Alice Labs
    Reviewed by
    Published ·Updated
    12 min read

    9-step process

    0/9 complete
    1. Step 1: Inventory current schema markup

      Crawl your site and extract every JSON-LD block. Catalog which pages have Article, FAQPage, HowTo, Organization, and Person schema. Validate each block with the Schema.org validator and Google Rich Results Test. Flag pages with missing, malformed, or inconsistent markup. This is the foundation — without valid schema, downstream audits are unreliable.

    2. Step 2: Audit entity signals (sameAs, NAP consistency)

      Verify your Schema.org Organization markup includes a sameAs property linking to Wikipedia, Wikidata, Crunchbase, and LinkedIn where applicable. Check that name, address, and phone (NAP) are identical across your site, Google Business Profile, social profiles, and external directories. Inconsistent entity signals confuse LLM retrieval systems.

    3. Step 3: Audit llms.txt presence and quality

      Check whether yourdomain.com/llms.txt exists and follows the Answer.AI specification (llmstxt.org). A complete file has an H1 site name, blockquote summary, H2 sections with curated link lists, and ideally a longer llms-full.txt variant. Missing or incomplete files are one of the most common audit findings.

    4. Step 4: Audit robots.txt and AI crawler access

      Pull your robots.txt and verify which AI user agents are allowed or blocked. Check OAI-SearchBot (ChatGPT), GPTBot (OpenAI training), PerplexityBot (Perplexity), ClaudeBot (Anthropic), and Google-Extended (Google AI). Decide deliberately: allowlist for citation visibility, blocklist for training-data control. Accidental blocks are common on sites that copied robots.txt from older templates.

    5. Step 5: Audit on-page LLMO elements

      For each priority page, check three things: a 40-50 word entity definition in the first paragraph, a quick-answer box under 160 characters near the top, and citation-rich body content with named sources, dates, and specific statistics. The Aggarwal et al. (2024) study found these patterns drive up to 40% lift in generative-engine visibility.

    6. Step 6: Audit FAQ structure and author schema

      Verify pillar pages have FAQ sections with FAQPage schema and 5-8 self-contained question-answer pairs. Check that every article has Person schema for the author with credentials, plus a reviewer Person schema where applicable. Missing author markup is a common E-E-A-T weakness that hurts citation probability.

    7. Step 7: Audit freshness signals

      Confirm every page emits both datePublished and dateModified in Article schema, and displays a visible 'Last updated' date on-page. Identify stale pages (no updates in 12+ months) on high-value URLs. Freshness is heavily weighted by AI search retrieval — undated content is deprioritized.

    8. Step 8: Audit off-site citation flow

      Map inbound mentions and citations from Tier-1 sources: industry publications, Wikipedia references, podcast transcripts, authoritative directories, and academic citations. Tools like Ahrefs and Mention help. Audit how often your brand is mentioned in LLM responses across ChatGPT, Perplexity, and Claude using Otterly.ai or Profound. Off-site authority is the slowest-moving but highest-leverage audit category.

    9. Step 9: Deliver a scorecard with prioritized fixes

      Score each of the 12 categories on a 0-3 scale (missing, partial, complete, best-in-class). Rank issues by impact and effort. Output a prioritized backlog: quick wins first (llms.txt, robots.txt, schema fixes), then medium-effort upgrades (entity signals, FAQ rollouts), then long-horizon initiatives (off-site citations, knowledge panel). Assign owners and review cadence.

    Key Takeaways

    • AI search engines (ChatGPT, Perplexity, Claude, Google AI Overviews) all rely on structured data, entity signals, and citation-worthy content — but each engine weights signals differently.
    • The Aggarwal et al. (2024) GEO paper found citation-rich content lifts visibility in generative engines by up to 40%.
    • An audit covers 12 categories: schema, entity signals, llms.txt, robots.txt, on-page LLMO elements, FAQ + author schema, freshness, off-site citations, brand monitoring, and knowledge panel presence.
    • Roughly 60% of Google searches now end without a click (SparkToro 2024), making citation-share the new visibility metric.
    • Most enterprise sites pass on schema basics but fail on llms.txt, author markup, and off-site citation flow.
    • The audit output is a prioritized backlog: high-impact fixes first, with effort estimates and owner assignments.
    01 / 05Step

    Why Audit Before Optimizing

    In short

    Optimizing without auditing wastes effort on the wrong fixes. An audit reveals which of the 12 GEO categories are weakest, so the team works on highest-leverage problems first instead of guessing.

    Most teams jump straight to writing FAQ blocks or adding schema. They skip the audit step.

    The result is predictable. Effort goes into categories that were already fine, while the real bottleneck — often robots.txt, missing entity signals, or stale freshness data — stays unsolved.

    A GEO audit forces a structured baseline. Every category gets scored. Every gap gets ranked by impact versus effort. The team starts with whatever moves the most citations per hour of work invested.

    There is also a strategic reason. AI search is a citation game, not a click game. The SparkToro 2024 study found roughly 60% of Google searches end without a click. If your brand is not the source being cited, you are invisible — and an audit is the fastest way to find out where you stand.

    02 / 05Step

    The 12 Audit Categories

    In short

    Every GEO audit must cover 12 categories: schema, entity signals, llms.txt, robots.txt, citation-rich content, quick-answer boxes, FAQ structure, author schema, freshness, inbound citations, brand monitoring, and knowledge panel presence.

    The 12 categories below come from triangulating three sources. The Aggarwal et al. (2024) GEO paper. Schema.org documentation. And recurring patterns we observe across 100+ Nordic enterprise implementations.

    Each category is independently testable. Each maps to a specific signal that AI retrieval systems consume.

    12-Category GEO Audit Checklist — what to check and which tools to use
    # Category What to check Tools
    1 On-page schema markup Article, FAQPage, HowTo, Organization, Person — JSON-LD, validates clean Schema.org validator, Google Rich Results Test
    2 Entity signals Organization sameAs links to Wikipedia/Wikidata/Crunchbase, NAP consistent everywhere Manual review, Brand SERP audit
    3 llms.txt presence File exists at /llms.txt, follows Answer.AI spec, includes curated key pages Manual fetch, llmstxt.org spec check
    4 robots.txt + AI crawlers OAI-SearchBot, GPTBot, PerplexityBot, ClaudeBot, Google-Extended — deliberate allow or block Manual robots.txt review
    5 Citation-rich content Named sources, specific statistics, dates on every claim, authoritative tone Manual content review, sample 10-20 priority pages
    6 Quick-answer box <160 char self-contained answer at top of every pillar page Manual review
    7 FAQ structure 5-8 question-answer pairs per pillar page with FAQPage schema Schema validator, manual content review
    8 Author + reviewer schema Person schema with credentials, role, sameAs to LinkedIn for author and reviewer Schema validator, manual review
    9 Freshness signals datePublished + dateModified in schema, visible 'Last updated' on-page Crawler audit, schema validator
    10 Inbound citations Mentions in Tier-1 publications, Wikipedia, podcast transcripts, authoritative directories Ahrefs, Mention, manual research
    11 LLM brand monitoring How often the brand appears in ChatGPT, Perplexity, Claude responses for target prompts Otterly.ai, Profound, manual prompt audits
    12 Knowledge Panel presence Google Knowledge Panel for the brand exists and contains accurate sameAs/founder data Branded SERP review
    03 / 05Step

    How Alice Labs Runs a GEO Readiness Audit

    In short

    The Alice Labs GEO Readiness Audit is a proprietary 30-point checklist that expands the 12 categories into specific, testable subchecks — delivered in roughly 1-2 days with a scorecard and prioritized backlog.

    The Alice Labs audit framework breaks each of the 12 categories into 2-3 concrete subchecks. That gives a 30-point checklist that two consultants can complete in a working day for a typical enterprise site.

    Phase one is data collection. We crawl the site, pull every JSON-LD block, fetch robots.txt and llms.txt, and run a 20-prompt audit across ChatGPT, Perplexity, and Claude.

    Phase two is scoring. Each of the 30 subchecks gets a 0-3 score with evidence attached. We benchmark against the Alice Labs LLMO Citation Benchmark — our internal dataset covering 100 SaaS brands across the Nordic region.

    Phase three is the backlog. We rank every gap by impact (citation potential) versus effort (engineering hours). The output is a sequenced 90-day roadmap with owners.

    The framework is deliberately bias-corrected. We cap how much weight any single category can carry, because in our experience teams over-invest in schema and under-invest in entity signals and off-site citations.

    Skip the DIY: get the Alice Labs GEO Readiness Audit

    Our consultants run the 30-point audit, benchmark you against 100 SaaS peers, and hand over a sequenced fix backlog with owners — typically in 1-2 working days.

    Talk to Alice Labs
    04 / 05Step

    Common Audit Findings on Enterprise Sites

    In short

    Most enterprise sites pass on basic Article schema but fail on llms.txt, author markup, entity sameAs links, and off-site citation flow. These are the four highest-leverage gaps to fix first.

    We see a recurring pattern across enterprise audits. The findings cluster into four buckets.

    Bucket one: missing llms.txt. Most enterprise sites have no llms.txt file. The Answer.AI standard launched September 2024 and adoption is still early. Publishing one takes under an hour.

    Bucket two: weak entity signals. Schema.org Organization markup often exists but lacks the sameAs property linking to Wikipedia, Wikidata, Crunchbase, and LinkedIn. NAP data is inconsistent across the site, Google Business Profile, and directory listings.

    Bucket three: incomplete author schema. Articles list a byline but do not emit Person schema with credentials or reviewer markup. This weakens E-E-A-T — Google's quality framework covering Experience, Expertise, Authoritativeness, and Trustworthiness.

    Bucket four: no off-site citation flow. Brands rarely show up in Tier-1 publications, Wikipedia references, or authoritative directories. Off-site citations are the slowest-moving signal, but they compound — and they are heavily weighted by ChatGPT, Perplexity, and Claude.

    05 / 05Step

    Building the Prioritized Fix Backlog

    In short

    A prioritized backlog ranks every audit gap by impact versus effort, sequences quick wins first, and assigns owners with review cadences — so the team executes the right fixes in the right order.

    The audit scorecard is the input. The fix backlog is the output. Without sequencing, even a perfect audit becomes a wishlist that nobody executes.

    We rank fixes on a 2x2 matrix: impact (low-high citation potential) and effort (low-high engineering hours).

    • Quick wins (high impact, low effort). Publish llms.txt. Fix robots.txt for AI crawlers. Add sameAs to Organization schema. Add datePublished/dateModified everywhere. Ship in week 1.
    • Medium projects (high impact, medium effort). Roll out FAQPage schema across pillar pages. Add Person schema with reviewer markup. Add quick-answer boxes and entity definitions to top 20 pages. Ship in weeks 2-6.
    • Long horizon (high impact, high effort). Build off-site citation flow. Earn Wikipedia references. Pursue Knowledge Panel claim and accuracy. Ship over months 3-12.
    • Skip or defer (low impact). Most schema cosmetic tweaks, micro-format experiments, and exotic markup types. Revisit only after the high-impact backlog is clear.

    Assign one owner per fix. Set a 30-day review cadence. Re-run the audit at 90 days to confirm the score actually improved. Without re-measurement, the audit becomes a one-off artifact instead of a feedback loop.

    About the Authors & Reviewers

    Published ·Updated
    Written by
    Linus Ingemarsson - Co-Founder, Alice Labs at Alice Labs
    Linus Ingemarsson

    Co-Founder, Alice Labs

    Co-Founder at Alice Labs. Author of 7 research reports on AI adoption, governance and labor markets cited across EU, OECD and US benchmarks.

    • 8+ years in AI strategy & implementation
    • Top-5 AI Speaker, Sweden (Mindley 2025)
    • 100+ enterprise AI engagements
    Reviewed by
    Eric Lundberg - Co-Founder, Alice Labs at Alice Labs
    Eric Lundberg

    Co-Founder, Alice Labs

    Co-Founder at Alice Labs. Builds AI automation, agent workflows and integration systems that hold up in real business operations.

    • AI automation & agent systems lead
    • Workflow design across 100+ deployments
    • Specialist in RAG, integrations & APIs
    Published · Updated
    Reviewed for technical accuracy, methodology and source integrity.·All claims trace to public sources cited in-line.

    Frequently Asked Questions

    What is a GEO audit?

    A GEO audit is a structured review of how prepared a website is to be cited by AI search engines like ChatGPT, Perplexity, Claude, and Google AI Overviews. It evaluates 12 categories spanning schema markup, entity signals, crawler access, content quality, and off-site authority — producing a scored report and prioritized fix backlog.

    How long does a GEO audit take?

    The audit itself takes 1-2 working days for a typical enterprise site when run by experienced consultants. Implementing the resulting fix backlog typically takes 4-12 weeks, depending on which categories are weakest and how many priority pages need updating.

    How is a GEO audit different from a traditional SEO audit?

    A traditional SEO audit focuses on rankings, backlinks, technical performance, and on-page targeting. A GEO audit focuses on whether AI systems can crawl, understand, and cite your content. They overlap on schema and freshness, but a GEO audit adds llms.txt, AI-crawler robots.txt rules, entity sameAs signals, citation-rich content patterns, and LLM brand monitoring.

    Do I need an llms.txt file to pass a GEO audit?

    No major LLM provider has publicly confirmed they consume llms.txt during inference, so it is not a hard requirement. However, the file is free to publish, takes under an hour, and signals AI-readiness. Most thorough GEO audits flag a missing llms.txt as a low-effort gap worth closing.

    Which AI crawlers should I allow in robots.txt?

    For citation visibility you typically want to allow OAI-SearchBot (ChatGPT search), PerplexityBot (Perplexity), ClaudeBot (Anthropic), and Google-Extended (Google AI Overviews). GPTBot is OpenAI's training crawler and is a separate decision based on whether you want your content used for model training. The audit should make every allow/block deliberate, not accidental.

    What is included in the Alice Labs GEO Readiness Audit?

    The Alice Labs GEO Readiness Audit is a proprietary 30-point checklist that expands the 12 standard categories into concrete subchecks. Output includes a scored report, evidence per finding, benchmark comparison against the Alice Labs LLMO Citation Benchmark covering 100 SaaS brands, and a sequenced 90-day fix backlog with owners.

    How often should I re-run the audit?

    Run the audit at three points: a baseline before optimization starts, a checkpoint at 90 days to validate the backlog is moving the score, and an annual full re-audit to catch new categories as the AI search landscape evolves. Re-measurement is what turns the audit from a one-off artifact into a feedback loop.

    Can I run a GEO audit myself or do I need a consultant?

    The 12-category checklist is designed to be self-serviceable for in-house teams with strong SEO and technical chops. Consultants add value when speed matters, when benchmarks against peer brands are needed, or when the team wants an outside perspective on the prioritization. Either path works — the most important step is starting.

    Previous in AI Search & LLMO

    AI Crawler Management: GPTBot, ClaudeBot, PerplexityBot & More

    Next in AI Search & LLMO

    GEO Strategy: How to Optimize for Google AI Overviews (2026)

    Further reading

    Related reading

    Sources

    1. Aggarwal et al. — GEO: Generative Engine Optimization (arXiv:2311.09735, 2024)(accessed 2026-05-06)
    2. Jeremy Howard / Answer.AI — llms.txt proposal (September 2024)(accessed 2026-05-06)
    3. Schema.org — official vocabulary (founded 2011 by Google, Bing, Yahoo, Yandex)(accessed 2026-05-06)
    4. Schema.org Validator(accessed 2026-05-06)
    5. Google — Search Quality Evaluator Guidelines (E-E-A-T framework)(accessed 2026-05-06)
    6. Google — AI Overviews launch (May 2024)(accessed 2026-05-06)
    7. OpenAI — ChatGPT Search launch and OAI-SearchBot documentation (October 31, 2024)(accessed 2026-05-06)
    8. SparkToro / Datos — 2024 zero-click search analysis (~60%)(accessed 2026-05-06)

    Next scheduled review:

    Want a GEO audit on your site?

    The Alice Labs GEO Readiness Audit is a proprietary 30-point checklist benchmarked against 100 SaaS brands. We deliver a scored report and a 90-day prioritized fix backlog.

    Request a GEO audit
    Share

    Get in Touch!

    The lab usually responds within 24 hours.

    Need help with AI?Get in touch