{"checkSetVersion":"v1.9-2026-08-11","kind":"page","methodology":{"version":"v1.1 · 2026-07-09","intro":"This audit measures whether a page is ready to be read, understood and used by AI search systems. Every check is deterministic and measured on the page itself, nothing is simulated. How checks combine into a score is evidence-informed editorial judgment, published and versioned here so anyone can reproduce a result.","ladder":[{"label":"Confirmed","weight":1,"meaning":"First-party documented mechanism (vendor documentation) or replicated measurement. The check measures something the platform operators themselves describe."},{"label":"Supported (single study)","weight":0.85,"meaning":"Causal support from controlled experiments (GEO, KDD 2024) corroborated observationally at scale (arXiv 2604.25707), but resting on a thin study base we name openly."},{"label":"Directional","weight":0.7,"meaning":"Recommended practice or indirect evidence: the mechanism is plausible and vendor-recommended, but no controlled study isolates the effect. Down-weighted symmetrically, an uncertain check counts less whether it passes or fails."}],"weightsNote":"Labels encode certainty, not effect size. Pillar weights (AI Access 20 · Content 15 · Evidence 30 · Brand Identity 15 · Structured Data 10 · Freshness & Sourcing 10) reflect the strength of interventional evidence per pillar; they are versioned editorial judgment, reviewed quarterly, not empirically derived constants.","bandsNote":"Score bands (Unreachable, then eight 10-point steps: Weak · Limited · Developing · Moderate · Solid · Strong · Excellent · Best-practice) describe what was verified on the page, readiness, never a citation prediction. No tool can honestly promise citations. A critical access failure (blocked, gated, or JavaScript-only content) caps the score at 20: if crawlers can't read the page, nothing else matters yet.","disclosure":"Integrity notes: counter-evidence is cited within the checks themselves (including the strongest null results against structured data). The makers of this audit publish provenance markup in their own work; that signal is therefore held to a stricter evidence standard and scored as unproven. Sources are primary only, vendor documentation and peer-reviewed or large-N studies, never secondhand summaries."},"checks":{"crawler-matrix":{"name":"AI crawler access","pillar":"crawler","evidenceLabel":"Confirmed","why":"Every AI assistant runs crawlers with different jobs: some collect training data, some build the search index behind answers, some fetch a page live because a user just asked about it. robots.txt controls each of these separately, per bot, per purpose. Blocking a search-purpose crawler removes a page from that engine's answer surfacing; blocking a training crawler is a data-policy choice that doesn't affect search; and user-triggered fetchers may access pages regardless of robots.txt. Only the search surface is scored, because only the search surface changes whether this page can be cited. The rest is reported alongside it, unweighted, so the file's full contents are visible without pretending they all mean the same thing.","evidence":"The per-surface consequences are vendor-documented, not inferred: OpenAI states that sites opted out of OAI-SearchBot \"will not be shown in ChatGPT search answers\", while Perplexity documents that its user-triggered fetcher generally ignores robots.txt. Google-Extended is the most commonly misunderstood token, it governs Gemini training only and has no effect on AI Overviews. Content-Signal is read and reported but never scored: the policy is a reservation of rights under Article 4 of EU Directive 2019/790, enforced through law rather than through crawler behaviour, and no AI vendor documents honouring it, so on this ladder it does not yet reach even Directional. An absent signal is reported as absent, because the policy states that a missing signal neither grants nor restricts the corresponding use.","sources":[{"id":"openai-bots","title":"OpenAI crawler documentation","publisher":"OpenAI","url":"https://developers.openai.com/api/docs/bots","grade":"Official","established":"The four OpenAI bots and their purposes; opting out of OAI-SearchBot removes a site from ChatGPT search answers; ChatGPT-User may fetch regardless of robots.txt."},{"id":"anthropic-crawlers","title":"Anthropic crawler documentation","publisher":"Anthropic","url":"https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler","grade":"Official","established":"Current Anthropic user-agent tokens (ClaudeBot, Claude-User, Claude-SearchBot) and their robots.txt semantics; earlier tokens are deprecated."},{"id":"perplexity-bots","title":"Perplexity crawler documentation","publisher":"Perplexity","url":"https://docs.perplexity.ai/guides/bots","grade":"Official","established":"PerplexityBot (search surfacing) respects robots.txt; Perplexity-User \"generally ignores robots.txt\" because a user requested the fetch."},{"id":"google-crawlers","title":"Google common crawlers reference","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers","grade":"Official","established":"Google-Extended controls Gemini training and grounding only, it does not affect Google Search or AI Overviews inclusion."},{"id":"content-signals","title":"Content Signals Policy","publisher":"Cloudflare","url":"https://contentsignals.org/","grade":"Official","established":"Defines the Content-Signal robots.txt directive and its three named uses (search, ai-input, ai-train); states that a signal set to yes grants and no forbids the corresponding use, that an omitted signal \"neither grants nor restricts permission\", and that any restriction is an express reservation of rights under Article 4 of EU Directive 2019/790."},{"id":"cloudflare-managed-robots","title":"Managed robots.txt","publisher":"Cloudflare","url":"https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/","grade":"Official","established":"Documents the three `use` values, immediate (\"Interact, but store and reuse nothing\"), reference (\"Index, excerpt, and link back\") and full (\"Summarize and reproduce\"), and that Cloudflare's managed file applies use=reference with search=yes,ai-train=no by default, so the directive on a client site is usually a platform default rather than a decision."}]},"bot-response":{"name":"What AI crawlers receive from your server","pillar":"crawler","evidenceLabel":"Directional","why":"robots.txt is a request; the server and its firewall decide what actually gets served. Some sites block AI crawlers at the WAF/CDN level even with an open robots.txt. This check probes the server with AI-crawler user-agents, but honestly: modern firewalls verify real crawlers by IP address, so a blocked probe can mean the site blocks AI bots or that its anti-spoofing simply works. Only an 'allowed' result is strong evidence.","evidence":"Cloudflare documents that requests with a known bot's user-agent from an unverified IP are flagged as fake bots and blocked by design. That is why this check reports asymmetrically and includes a control probe, ground truth about verified-crawler treatment requires server logs.","sources":[{"id":"cloudflare-fakebots","title":"Fake-bot managed rules","publisher":"Cloudflare","url":"https://developers.cloudflare.com/waf/troubleshooting/fake-bot-managed-rules/","grade":"Official","established":"WAFs verify real crawlers by IP, not user-agent; a spoofed bot user-agent from an unverified IP is blocked by design, which is why a blocked probe is ambiguous."},{"id":"openai-bots","title":"OpenAI crawler documentation","publisher":"OpenAI","url":"https://developers.openai.com/api/docs/bots","grade":"Official","established":"The four OpenAI bots and their purposes; opting out of OAI-SearchBot removes a site from ChatGPT search answers; ChatGPT-User may fetch regardless of robots.txt."}]},"gate-parity":{"name":"Cookie / consent gating","pillar":"rendered","evidenceLabel":"Confirmed","why":"Indexing and on-demand crawlers browse cookieless and never click buttons. A server-side age gate, cookie wall or consent wall that swaps out the content therefore makes the entire page invisible to AI, the crawler sees only the gate, forever. Overlay banners with the full content underneath are fine; server-side gating is the failure mode. This matters most in gated industries: alcohol, supplements, finance, pharma.","evidence":"Google documents crawling stateless (cookies cleared between loads) and advises against content-blocking interstitials; no AI vendor documents any gate-dismissal capability. One nuance: consent walls are often geo-targeted, and crawler traffic egresses mostly from the US, so a gate seen from Europe may not be what the crawler sees.","sources":[{"id":"google-interstitials","title":"Interstitials and dialogs guidance","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/appearance/avoid-intrusive-interstitials","grade":"Official","established":"Crawlers do not dismiss interstitials; content behind server-side gates is not indexed."},{"id":"google-ai-guide","title":"AI features and your website","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/fundamentals/ai-optimization-guide","grade":"Official","established":"Google's own guidance for AI features: people-first content and clear structure; no special structured data is required; content need not be pre-chunked."}]},"raw-html":{"name":"Content in raw HTML","pillar":"rendered","evidenceLabel":"Confirmed","why":"Most AI crawlers download a page's HTML and read exactly that, they do not run JavaScript. Content that only appears after scripts execute (client-side apps, script-injected text) simply does not exist for those pipelines. The page may still surface as a bare link via third-party indexes, but nothing on it can be read, quoted or grounded.","evidence":"The largest crawler study to date (1B+ requests) found no major AI crawler executes JavaScript, GPTBot even downloads JS files without running them. The documented exceptions: Google's Gemini/AI Overviews, Applebot and Bingbot/Copilot do render, which is why the finding says \"unreadable to ~87% of assistant traffic\", not \"invisible to AI\".","sources":[{"id":"vercel-merj","title":"The Rise of the AI Crawler","publisher":"Vercel × MERJ","url":"https://vercel.com/blog/the-rise-of-the-ai-crawler","grade":"Tier A","established":"1B+ requests analyzed: none of the major AI crawlers execute JavaScript (GPTBot fetches JS files in 11.5% of requests and runs none); AI crawlers show notably high 404 rates."},{"id":"statcounter","title":"AI chatbot market share","publisher":"Statcounter Global Stats","url":"https://gs.statcounter.com/ai-chatbot-market-share","grade":"Tier B","established":"Live usage shares of AI assistants, the basis for estimating what fraction of assistant traffic runs on non-rendering fetch pipelines."},{"id":"google-ai-guide","title":"AI features and your website","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/fundamentals/ai-optimization-guide","grade":"Official","established":"Google's own guidance for AI features: people-first content and clear structure; no special structured data is required; content need not be pre-chunked."}]},"canonical":{"name":"Canonical web address","pillar":"rendered","evidenceLabel":"Directional","why":"The canonical tag tells search engines which URL is the 'real' one when several serve similar content. No AI engine reads this tag directly, but AI answers frequently cite pages via Google's and Bing's indexes, and those indexes follow canonical decisions. A page that defers its canonical elsewhere is telling the whole chain to cite the other URL.","evidence":"Google documents rel=canonical as a strong hint (not a directive) and picks a canonical automatically when it's missing, which is why a missing tag on a simple URL costs little, while conflicting canonicals genuinely confuse indexing.","sources":[{"id":"google-canonical","title":"Consolidate duplicate URLs (canonicalization)","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls","grade":"Official","established":"rel=canonical is a strong hint, not a directive; if absent, Google picks a canonical automatically; absolute URLs recommended."}]},"sitemap":{"name":"Sitemap listing","pillar":"rendered","evidenceLabel":"Directional","why":"A sitemap is a machine-readable list of the pages you want discovered. AI crawlers cover sites incompletely and hit dead links at notably high rates, so a current sitemap is cheap discovery insurance, but it is hygiene, not a requirement, and no AI vendor documents sitemaps as an inclusion factor.","evidence":"Google's own guidance states that small, well-linked sites (~500 pages or fewer) might not need a sitemap at all, which is why this check never fails a page outright. The crawler study behind the raw-HTML check also documents AI crawlers' high 404 rates, the practical argument for keeping sitemaps current.","sources":[{"id":"google-sitemaps","title":"Sitemaps overview","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/crawling-indexing/sitemaps/overview","grade":"Official","established":"Sites of roughly 500 pages or fewer with sound internal linking \"might not need a sitemap\", sitemaps are discovery hygiene, not a requirement."},{"id":"vercel-merj","title":"The Rise of the AI Crawler","publisher":"Vercel × MERJ","url":"https://vercel.com/blog/the-rise-of-the-ai-crawler","grade":"Tier A","established":"1B+ requests analyzed: none of the major AI crawlers execute JavaScript (GPTBot fetches JS files in 11.5% of requests and runs none); AI crawlers show notably high 404 rates."}]},"snippet-description":{"name":"Search snippet summary","pillar":"rendered","evidenceLabel":"Directional","why":"A result snippet is often the only text a reader or an assistant sees before deciding whether to open the page. Google generates it from the page, or from your description when you provide one, so an absent description hands that sentence over to a generator, and an over-long one gets cut mid-sentence. Assistants increasingly quote the description verbatim as a page summary.","evidence":"Google documents the mechanism (snippets come from content or the description; descriptions may be truncated) but states NO character limit, and truncation is by pixel width. The 120-165 band is practitioner convention, so this check is graded Directional: the mechanism is documented, the threshold is a proxy.","sources":[{"id":"google-snippet-doc","title":"Control your snippets in search results","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/appearance/snippet","grade":"Official","established":"Google generates result snippets from page content or from the meta description, and states that descriptions may be truncated; no character limit is specified."},{"id":"ms-seo-reference","title":"SEO reference for Microsoft Learn contributors","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/en-us/contribute/content/seo-reference","grade":"Tier B","established":"House style guide: title 30-65 characters with the keyword front-loaded, description 120-165 characters, slug short and keyword-bearing. Convention, not vendor specification."}]},"stats-density":{"name":"Facts & figures","pillar":"evidence","evidenceLabel":"Supported (single study)","why":"AI answers are assembled from extractable facts. A page that argues with adjectives ('extremely light, highly versatile') gives an engine nothing to quote; a page that states exact values (density, mass, cost, percentages) provides the raw material answers are built from. This is the single best-evidenced content lever.","evidence":"In controlled experiments, adding statistics lifted generative-engine visibility 30–40%, one of only three interventions that worked (keyword stuffing did nothing). The effect replicated (+15.6pp absorption) and holds at observational scale (+62% answer influence across 21k citations). Caveat we cite ourselves: the effect is domain-dependent, strongest for factual/regulatory content, and being topically relevant still outweighs any on-page tactic.","sources":[{"id":"geo-paper","title":"GEO: Generative Engine Optimization (KDD 2024)","publisher":"arXiv 2311.09735 · Aggarwal et al.","url":"https://arxiv.org/abs/2311.09735","grade":"Tier A","established":"Controlled experiments: adding statistics, quotations, or source citations lifted generative-engine visibility 30–41%; keyword stuffing showed little to no improvement; effects are domain-dependent."},{"id":"citation-absorption","title":"From Citation Selection to Citation Absorption","publisher":"arXiv 2604.25707 · Zhang, He & Yao","url":"https://arxiv.org/abs/2604.25707","grade":"Tier A","established":"21,143 AI-search citations measured (observational): pages with statistics (+62%), definitions (+57%), and comparisons (+55%) carry more answer influence; Q&A format alone showed −5.7%, formatting without evidence does not help."},{"id":"beyond-seo","title":"Beyond SEO: Transformer-based web content optimisation","publisher":"arXiv 2507.03169","url":"https://arxiv.org/abs/2507.03169","grade":"Tier A","established":"Replication: adding statistical evidence and credible citations raised absorption of page content into AI answers by +15.6 percentage points."},{"id":"sprinklr","title":"What Gets Cited (252,000 paired trials)","publisher":"arXiv 2605.25517","url":"https://arxiv.org/abs/2605.25517","grade":"Tier A","established":"Across 6 LLMs and 18 content factors, topical relevance and list position outweigh any on-page content factor, evidence density matters after a page is retrieved, not instead of being relevant."}]},"definitions":{"name":"Definitions","pillar":"evidence","evidenceLabel":"Supported (single study)","why":"A crisp one-sentence definition ('X is a Y that …') is the most quotable object a page can contain: self-contained, unambiguous, and exactly the shape an answer engine needs when a user asks 'what is X?'. Pages without a single definitional sentence force engines to synthesize one, usually from someone else's page.","evidence":"In the largest citation-absorption measurement, definition markers were among the strongest signals, particularly for Google's AI surfaces (+57% answer influence). The finding is correlational, and detection here is pattern-based, so phrasing definitions clearly matters more than gaming any pattern.","sources":[{"id":"citation-absorption","title":"From Citation Selection to Citation Absorption","publisher":"arXiv 2604.25707 · Zhang, He & Yao","url":"https://arxiv.org/abs/2604.25707","grade":"Tier A","established":"21,143 AI-search citations measured (observational): pages with statistics (+62%), definitions (+57%), and comparisons (+55%) carry more answer influence; Q&A format alone showed −5.7%, formatting without evidence does not help."},{"id":"google-ai-guide","title":"AI features and your website","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/fundamentals/ai-optimization-guide","grade":"Official","established":"Google's own guidance for AI features: people-first content and clear structure; no special structured data is required; content need not be pre-chunked."}]},"comparisons":{"name":"Comparisons","pillar":"evidence","evidenceLabel":"Supported (single study)","why":"A large share of commercial AI queries are comparisons ('X vs Y', 'is X lighter than Y?'). Engines answering them retrieve pages that already contain the comparison, with numbers, not adjectives. A concrete comparison table against the obvious alternative makes a page the natural source for that entire query class.","evidence":"Comparison content correlates with +55% answer influence across 21k measured citations. Honest caveat: in the wild, the biggest comparison-content effects accrue to third-party comparison pages; a first-party table earns retrieval for comparison queries rather than third-party neutrality.","sources":[{"id":"citation-absorption","title":"From Citation Selection to Citation Absorption","publisher":"arXiv 2604.25707 · Zhang, He & Yao","url":"https://arxiv.org/abs/2604.25707","grade":"Tier A","established":"21,143 AI-search citations measured (observational): pages with statistics (+62%), definitions (+57%), and comparisons (+55%) carry more answer influence; Q&A format alone showed −5.7%, formatting without evidence does not help."},{"id":"sprinklr","title":"What Gets Cited (252,000 paired trials)","publisher":"arXiv 2605.25517","url":"https://arxiv.org/abs/2605.25517","grade":"Tier A","established":"Across 6 LLMs and 18 content factors, topical relevance and list position outweigh any on-page content factor, evidence density matters after a page is retrieved, not instead of being relevant."}]},"entity-density":{"name":"References & standards","pillar":"evidence","evidenceLabel":"Directional","why":"Named standards, institutes and certifications do two jobs: they make claims verifiable (an engine can anchor a claim to a recognised standard rather than trust an adjective), and they ground the page in the entity graph engines use to understand what things are. A page rich in checkable references reads as evidence; a page without them reads as marketing.","evidence":"This check is a hypothesis-level heuristic, we say so openly: no study measures on-page entity density directly. The adjacent evidence is causal for citing verifiable sources (+30–40% in controlled experiments) and correlational for off-page entity presence. The detector counts standard codes, institutions and brand tokens, never bare capitalized words.","sources":[{"id":"geo-paper","title":"GEO: Generative Engine Optimization (KDD 2024)","publisher":"arXiv 2311.09735 · Aggarwal et al.","url":"https://arxiv.org/abs/2311.09735","grade":"Tier A","established":"Controlled experiments: adding statistics, quotations, or source citations lifted generative-engine visibility 30–41%; keyword stuffing showed little to no improvement; effects are domain-dependent."},{"id":"sprinklr","title":"What Gets Cited (252,000 paired trials)","publisher":"arXiv 2605.25517","url":"https://arxiv.org/abs/2605.25517","grade":"Tier A","established":"Across 6 LLMs and 18 content factors, topical relevance and list position outweigh any on-page content factor, evidence density matters after a page is retrieved, not instead of being relevant."}]},"headings":{"name":"Headings & structure","pillar":"evidence","evidenceLabel":"Supported (single study)","why":"Retrieval systems don't read pages top to bottom, they split them into chunks and fetch the chunk that answers the query. Headings are the natural cut lines. A 600-word wall under one heading becomes one bloated, unfocused chunk; heading-bounded sections of ~200 words or less survive as coherent, quotable retrieval units.","evidence":"Independent retrieval research converges on compact passages: 64–128 tokens optimal for factual answers, ~100-token chunks winning across strategies. Google's own guidance asks for well-organized, sectioned content while noting chunking isn't *required*, this check measures the structure Google recommends, with thresholds anchored in the retrieval literature.","sources":[{"id":"chunk-size","title":"Rethinking chunk size for long-document retrieval","publisher":"arXiv 2505.21700","url":"https://arxiv.org/abs/2505.21700","grade":"Tier A","established":"Retrieval works best on compact, self-contained passages: 64–128 tokens optimal for fact-based answers, the engineering basis for heading-bounded sections."},{"id":"chunk-twice","title":"Chunk Twice, Embed Once","publisher":"arXiv 2506.17277","url":"https://arxiv.org/abs/2506.17277","grade":"Tier A","established":"Across chunking strategies, ~100-token recursive chunks consistently win, long unstructured blocks fragment badly at retrieval time."},{"id":"google-ai-guide","title":"AI features and your website","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/fundamentals/ai-optimization-guide","grade":"Official","established":"Google's own guidance for AI features: people-first content and clear structure; no special structured data is required; content need not be pre-chunked."}]},"answer-chunks":{"name":"Answer sections","pillar":"evidence","evidenceLabel":"Directional","why":"AI search decomposes user questions into sub-queries and retrieves per sub-query. Sections framed as a question with a complete, self-contained answer align with that machinery, the section can be lifted whole. But the question format alone is worthless: a question heading over an evidence-free paragraph gives an engine nothing to absorb.","evidence":"AI Overviews trigger on 64.7% of question-form queries (vs 13.7% overall), the retrieval-side support. The absorption-side caution comes from the same measurement study this audit uses elsewhere: Q&A format alone showed −5.7% influence, and FAQ markup earns nothing (which is why it isn't scored here). Question sections work when they carry the evidence the other checks in this pillar measure.","sources":[{"id":"aio-questions","title":"AI Overviews triggering study","publisher":"arXiv 2605.14021","url":"https://arxiv.org/abs/2605.14021","grade":"Tier A","established":"AI Overviews appear on 64.7% of question-form queries vs 13.7% overall, question-shaped content aligns with how AI search decomposes queries."},{"id":"citation-absorption","title":"From Citation Selection to Citation Absorption","publisher":"arXiv 2604.25707 · Zhang, He & Yao","url":"https://arxiv.org/abs/2604.25707","grade":"Tier A","established":"21,143 AI-search citations measured (observational): pages with statistics (+62%), definitions (+57%), and comparisons (+55%) carry more answer influence; Q&A format alone showed −5.7%, formatting without evidence does not help."},{"id":"google-ai-guide","title":"AI features and your website","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/fundamentals/ai-optimization-guide","grade":"Official","established":"Google's own guidance for AI features: people-first content and clear structure; no special structured data is required; content need not be pre-chunked."}]},"name-consistency":{"name":"Name consistency","pillar":"entity","evidenceLabel":"Confirmed","why":"Machines resolve 'who is this page about?' by matching names across structured data, metadata and titles. When those fields disagree, one name in JSON-LD, another in og:site_name, a third in the title, resolvers get contradictory keys for the same entity. Multiple names are normal business reality; the sanctioned way to declare them is alternateName/legalName, which this check never penalizes.","evidence":"Google documents reading exactly these fields for site naming and asks for consistency across them, with alternateName as the mechanism for variants. Effects on non-Google AI engines are inferred from how entity resolution works generally, documented mechanism on Google, directional beyond it.","sources":[{"id":"google-site-names","title":"Site names documentation","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/appearance/site-names","grade":"Official","established":"Google reads WebSite structured data, og:site_name, titles and headings for site/entity naming, and asks for consistency across them."},{"id":"google-org","title":"Organization structured data","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/appearance/structured-data/organization","grade":"Official","established":"sameAs is a recommended Organization property; Google names leiCode/iso6523Code as the properties used to disambiguate organizations."}]},"topic-continuity":{"name":"Topic continuity","pillar":"entity","evidenceLabel":"Directional","why":"Retrieval lifts a heading and its passage out of the page. If the heading, the browser title and the summary each name the subject differently, a retrieved fragment cannot be resolved back to one topic, and the page competes with itself for the same query. Naming the same subject in all three is the cheapest disambiguation there is.","evidence":"Google documents that the title link may be taken from the title element OR from on-page headings, which is a first-party reason for the two to agree. That the agreement improves retrieval is plausible and widely recommended but not isolated by any controlled study, so this is Directional. Measured as term overlap, never against a declared keyword: a keyword field would reward stuffing, and a product page's subject is its product, not its brand.","sources":[{"id":"google-title-links","title":"Influencing your title links in search results","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/appearance/title-link","grade":"Official","established":"The title link may be generated from the <title> element, on-page headings such as the H1, or other prominent page text; consistent descriptive text across them is recommended."},{"id":"google-site-names","title":"Site names documentation","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/appearance/site-names","grade":"Official","established":"Google reads WebSite structured data, og:site_name, titles and headings for site/entity naming, and asks for consistency across them."},{"id":"ms-seo-reference","title":"SEO reference for Microsoft Learn contributors","publisher":"Microsoft Learn","url":"https://learn.microsoft.com/en-us/contribute/content/seo-reference","grade":"Tier B","established":"House style guide: title 30-65 characters with the keyword front-loaded, description 120-165 characters, slug short and keyword-bearing. Convention, not vendor specification."}]},"sameas":{"name":"Verified-profile links","pillar":"entity","evidenceLabel":"Directional","why":"sameAs links connect your Organization markup to its profiles elsewhere, LinkedIn, business registers, Wikidata. That corroboration helps knowledge systems confirm which real-world entity your site is, instead of guessing among similarly named companies. Independent, resolvable profiles are the point; self-declared links to nothing prove nothing.","evidence":"Google recommends sameAs on Organization markup and names leiCode/iso6523Code as its organization-disambiguation properties. A Wikidata entry is a genuine upgrade, Wikidata feeds knowledge graphs used in AI grounding, but Wikidata's own notability policy requires serious public references, which is why it's scored as an upgrade to earn, never a gate. No controlled study ties sameAs to AI-answer inclusion; this check is honest hygiene.","sources":[{"id":"google-org","title":"Organization structured data","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/appearance/structured-data/organization","grade":"Official","established":"sameAs is a recommended Organization property; Google names leiCode/iso6523Code as the properties used to disambiguate organizations."},{"id":"wikidata-notability","title":"Wikidata notability policy","publisher":"Wikidata","url":"https://www.wikidata.org/wiki/Wikidata:Notability","grade":"Official","established":"Wikidata items require serious, publicly available references, why a Wikidata entry is an upgrade to earn, not a box to tick."}]},"jsonld":{"name":"Structured data (JSON-LD)","pillar":"structured","evidenceLabel":"Confirmed","why":"JSON-LD is a machine-readable description of what a page is and offers. Google and Microsoft confirm consuming it for indexing, grounding and Copilot, so broken markup poisons a channel that is actually read, and empty boilerplate leaves it unused. What structured data does *not* do, on current evidence, is win AI citations by itself.","evidence":"The consumption side is first-party confirmed (Microsoft: schema helps Bing's LLMs; Google: structured data aids grounding). The impact side is where we cite the counter-evidence ourselves: the strongest study, 1,885 pages adding schema vs 4,000 controls, found no meaningful AI-citation uplift, and Google states no special markup is needed for AI features. Hence: hygiene, not a lever.","sources":[{"id":"bing-schema","title":"Microsoft: Bing/Copilot use schema for LLMs","publisher":"Search Engine Land (reporting Fabrice Canel, Microsoft)","url":"https://searchengineland.com/microsoft-bing-copilot-use-schema-for-its-llms-453455","grade":"Tier B","established":"Microsoft states schema markup helps Bing's LLMs and Copilot understand content, the first-party-confirmed consumption side of structured data."},{"id":"ahrefs-schema","title":"Schema markup and AI citations (difference-in-differences)","publisher":"Ahrefs","url":"https://ahrefs.com/blog/schema-ai-citations/","grade":"Tier B","established":"1,885 pages that added JSON-LD vs ~4,000 controls: no meaningful AI-citation uplift, the strongest counter-evidence to markup-as-lever, cited here deliberately."},{"id":"google-ai-guide","title":"AI features and your website","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/fundamentals/ai-optimization-guide","grade":"Official","established":"Google's own guidance for AI features: people-first content and clear structure; no special structured data is required; content need not be pre-chunked."},{"id":"searchviu","title":"Hidden structured data extraction experiment","publisher":"searchVIU","url":"https://www.searchviu.com/en/schema-markup-and-ai-in-2025-what-chatgpt-claude-perplexity-gemini-real","grade":"Tier B","established":"Values placed only in hidden JSON-LD were extracted by 0 of 5 AI systems during live retrieval, engines read the visible page."}]},"parity":{"name":"Markup / text parity","pillar":"structured","evidenceLabel":"Confirmed","why":"Everything claimed in structured data must also be visible on the page. Markup-only content fails twice: it violates Google's structured-data policy (a manual-action category), and, more practically, AI engines read the rendered page, so a claim that exists only in markup can never be quoted. It is policy exposure plus wasted effort.","evidence":"Google's policy language is explicit: \"Don't mark up content that is not visible to readers of the page.\" The invisibility half is experimental: values placed only in hidden markup were extracted by zero of five AI systems at retrieval time. Note: FAQ *rich results* ended in May 2026, visible answer quality, not FAQ markup, is where answer-engine value lives.","sources":[{"id":"google-sd-policies","title":"Structured data general policies","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/appearance/structured-data/sd-policies","grade":"Official","established":"\"Don't mark up content that is not visible to readers of the page\", markup must be a true representation of page content; violations are a manual-action category."},{"id":"searchviu","title":"Hidden structured data extraction experiment","publisher":"searchVIU","url":"https://www.searchviu.com/en/schema-markup-and-ai-in-2025-what-chatgpt-claude-perplexity-gemini-real","grade":"Tier B","established":"Values placed only in hidden JSON-LD were extracted by 0 of 5 AI systems during live retrieval, engines read the visible page."}]},"freshness":{"name":"Modification dates","pillar":"freshness","evidenceLabel":"Directional","why":"AI answers skew measurably toward recently updated content, and undated pages are discounted, engines can't rank what they can't date. The right response is honest dating: machine-readable modification dates that match the visible ones. The wrong response is bumping dates without changes, a documented counter-signal, since engines cross-check dates against crawl history.","evidence":"The largest citation dataset (17M) shows AI-cited content is 25.7% fresher than organic results; freshness ranks among the top citation-associated properties in framework studies. Two caveats we carry into the scoring: recency bias is category-dependent (reference content ages far slower than commercial content, an older dated page scores partial, never fail), and Google explicitly warns against unearned date changes.","sources":[{"id":"ahrefs-fresh","title":"Fresh content study (17M citations)","publisher":"Ahrefs","url":"https://ahrefs.com/blog/fresh-content/","grade":"Tier A","established":"AI-cited content is 25.7% fresher than organic results; ChatGPT cites URLs roughly 400+ days newer than classic search, the largest-N recency evidence."},{"id":"geo-16","title":"GEO-16: AI answer engine citation behavior framework","publisher":"arXiv 2509.10762","url":"https://arxiv.org/abs/2509.10762","grade":"Tier A","established":"Metadata & freshness rank among the on-page properties most strongly associated with being cited by AI answer engines (correlational)."},{"id":"google-dates","title":"Publication dates guidance","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/appearance/publication-dates","grade":"Official","established":"Google recommends datePublished/dateModified in structured data, cross-checks them against visible dates and crawl history, and never relies on a single date signal."},{"id":"seer-recency","title":"AI brand visibility and content recency","publisher":"Seer Interactive","url":"https://www.seerinteractive.com/insights/study-ai-brand-visibility-and-content-recency","grade":"Tier B","established":"Recency bias is category-dependent: financial content decays fast while instructional/reference content holds up 10–15 years, freshness thresholds must not punish evergreen material."}]},"provenance":{"name":"Source attribution","pillar":"freshness","evidenceLabel":"Supported (single study)","why":"Visibly citing sources, linked datasheets, named institutions, referenced test reports, is one of the few content changes with causal evidence behind it: engines prefer to absorb claims they can trace, and users trust answers built on traceable pages. Machine-readable citation markup is the forward-looking complement; no engine currently documents reading it, so it earns a note, not points.","evidence":"In controlled experiments, 'cite sources' was a top-3 tactic (+30–40% visibility); a preregistered RCT shows citations increase user trust in generative answers. The markup half is scored strictly: hidden structured data was read by zero of five engines at retrieval time. Disclosure: the makers of this audit publish provenance markup in their own work, precisely why that signal is held to the stricter standard here and scored as unproven.","sources":[{"id":"geo-paper","title":"GEO: Generative Engine Optimization (KDD 2024)","publisher":"arXiv 2311.09735 · Aggarwal et al.","url":"https://arxiv.org/abs/2311.09735","grade":"Tier A","established":"Controlled experiments: adding statistics, quotations, or source citations lifted generative-engine visibility 30–41%; keyword stuffing showed little to no improvement; effects are domain-dependent."},{"id":"trust-rct","title":"Human trust in AI search (preregistered RCT)","publisher":"arXiv 2504.06435","url":"https://arxiv.org/abs/2504.06435","grade":"Tier A","established":"Adding reference links and citations measurably increases user trust in generative answers (~12,000 queries, preregistered)."},{"id":"searchviu","title":"Hidden structured data extraction experiment","publisher":"searchVIU","url":"https://www.searchviu.com/en/schema-markup-and-ai-in-2025-what-chatgpt-claude-perplexity-gemini-real","grade":"Tier B","established":"Values placed only in hidden JSON-LD were extracted by 0 of 5 AI systems during live retrieval, engines read the visible page."}]}}}