Methodology

How Legible scores a store.

Twenty-five checks, four pillars, one number. Everything below is the same code that runs your scan, so the weights here are the weights you get.

Scoring

Each check returns pass (1), warn (0.5), fail (0) or info (not scored). A few checks return a partial score, for example Product schema completeness or the share of ACP prerequisites met.

A pillar score is the weighted average of its checks, from 0 to 100. The overall score is the weighted average of the four pillars:

Access

30%

Understanding

30%

Trust

20%

Transaction

20%

Caps

Some failures make everything else irrelevant. When one of these fails, the overall score cannot exceed the cap, however good the rest is. The report says when a cap applied and what the uncapped score was.

CheckConditionMax score
Homepage answers an agentFails because the homepage did not answer10
No bot wall in front of the storeFails because a bot wall turns agents away before they see any page29
AI search and shopping agents are allowedFails because robots.txt blocks the main AI search agents44
Product facts are in the raw HTMLFails because the product page is empty without JavaScript59

Grades

  1. A

    90+

    Agent-ready

  2. B

    75+

    Mostly legible

  3. C

    60+

    Partly legible

  4. D

    45+

    Hard to read

  5. E

    30+

    Mostly invisible

  6. F

    0+

    Invisible to agents

01 · 30% of the score

Access

Can an agent reach your pages at all?

Homepage answers an agent

Every agent starts with an HTTP request. Anything other than a fast 200 with HTML ends the visit before it begins.

access.reachable

3/ 19 pts

No bot wall in front of the store

Bot-protection that challenges every non-browser client blocks AI agents even when robots.txt welcomes them. Agents cannot solve JavaScript challenges or CAPTCHAs.

access.botwall

3/ 19 pts

AI search and shopping agents are allowed

Search agents (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot) build the indexes assistants search before recommending a product. User-triggered agents (ChatGPT-User, Claude-User, Perplexity-User) fetch your page while a shopper waits. Blocking either removes you from the answer.

access.robots-search

5/ 19 pts

Training crawler policy

Training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot and others) collect content for future models. Blocking them is a legitimate business choice and does not affect AI search visibility. Legible reports your stance but does not score it.

access.robots-training

Not scored

robots.txt is valid

RFC 9309 says a robots.txt that errors with 5xx means "crawl nothing", and an HTML page served as robots.txt is noise. A clean file also points agents at your sitemap.

access.robots-file

1/ 19 pts

XML sitemap lists your products

Crawlers that build AI search indexes discover product URLs from sitemaps far more reliably than from links, especially on large catalogues.

access.sitemap

2/ 19 pts

Fast first byte

User-triggered agents fetch pages while someone waits for an answer, with short timeouts. A slow server is a skipped store.

access.ttfb

2/ 19 pts

Lean HTML

Agents download and tokenise your HTML. Megabytes of inline scripts and styles bury product facts and can hit size caps before the content.

access.weight

1/ 19 pts

Pages are not marked noindex

A noindex meta tag or X-Robots-Tag header tells search agents to drop the page from their index, even when robots.txt allows crawling.

access.indexable

2/ 19 pts

02 · 30% of the score

Understanding

Can it read what you sell without running JavaScript?

Product facts are in the raw HTML

Most AI fetchers do not run JavaScript. If the name, price and description only appear after scripts run, the agent sees an empty page and moves on.

understanding.raw-html

5/ 16 pts

Complete Product structured data

Schema.org Product and Offer JSON-LD is the most reliable way for agents and shopping indexes to read price, currency, stock, identifiers, shipping and returns without guessing from layout.

understanding.product-schema

5/ 16 pts

Organization and WebSite schema

Organization and WebSite markup tells an agent the store's official name, logo, URL and social profiles, so it can tell your shop apart from resellers and lookalikes.

understanding.site-schema

2/ 16 pts

OpenGraph product preview

Chat interfaces render link previews and many agents fall back to OpenGraph when JSON-LD is missing. og:title, og:image and og:type are the minimum.

understanding.opengraph

1/ 16 pts

Canonical URL and language

A canonical URL collapses filter and tracking variants into one product, and a lang attribute tells the agent which language and market the price and policies apply to.

understanding.canonical-lang

1/ 16 pts

One clear H1 with the product name

Extractors use the heading outline to split a page into sections. One H1 naming the product, then H2s for details and reviews, makes the page easy to chunk.

understanding.headings

1/ 16 pts

llms.txt guide for agents

llms.txt is a proposed markdown index of a site's most useful pages. No major assistant has confirmed it uses the file, so it carries little weight here, but it is cheap and gives agents a clean map of products and policies.

understanding.llms-txt

1/ 16 pts

03 · 20% of the score

Trust

Can it verify who you are and what the terms are?

HTTPS everywhere, with HSTS

Agents that hand a shopper off to checkout will not send them to an insecure page. HSTS proves the store never falls back to plain HTTP.

trust.https

3/ 14 pts

Policy pages are linked

Before recommending a purchase, assistants look for shipping costs, the return window and who to contact. In the EU, consumers have a 14-day right of withdrawal for most online purchases (Directive 2011/83/EU, Art. 9), so returns terms are a deciding fact.

trust.policies

4/ 14 pts

Business identity is machine-readable

Agents weigh whether a merchant is real. A legal name, postal address and a contact point in Organization schema are the strongest signals short of a marketplace listing.

trust.identity

3/ 14 pts

Visible price matches structured data

When the price a shopper sees differs from the price in JSON-LD, agents quote the wrong number or distrust both. Google Merchant Center suspends listings for mismatches, and under EU consumer law the price shown must be the price charged.

trust.price-consistency

4/ 14 pts

04 · 20% of the score

Transaction

Can it start a purchase through a protocol?

UCP business profile

The Universal Commerce Protocol (UCP), announced by Google in January 2026, lets agents discover a store's capabilities from /.well-known/ucp and complete checkout through a standard API. It powers buying in Google AI Mode and Gemini. Current spec version: 2026-08-25.

transaction.ucp

5/ 11 pts

Machine-readable catalogue

Agent channels (ChatGPT shopping, Google AI Mode, Copilot) prefer a structured product feed they can refresh over crawling pages. A feed keeps price and stock current between crawls.

transaction.catalogue

3/ 11 pts

Agentic Commerce Protocol readiness (signal)

OpenAI and Stripe's Agentic Commerce Protocol (ACP) launched in September 2025 to let shoppers buy inside ChatGPT. There is no public discovery file: merchants are onboarded through OpenAI and submit a product feed. In March 2026 OpenAI moved checkout back to merchants' own sites, so discovery in ChatGPT now depends on that feed and on crawlable pages. Legible scores the prerequisites, not enrolment.

transaction.acp

1/ 11 pts

Add-to-cart works without scripts

Browser-driving agents (ChatGPT agent, Gemini in Chrome and similar) click through your real checkout. A plain form or link for add-to-cart in the server HTML is the most robust path when no protocol is available.

transaction.cart-path

2/ 11 pts

Platform and remediation path

What you can fix, and how fast, depends on the platform. Shopify merchants can switch most of this on; WooCommerce, Magento and custom stores need it built or configured.

transaction.platform

Not scored

AI agents evaluated in robots.txt

Legible applies RFC 9309 matching (named group first, then *, longest rule wins, allow wins ties) to the homepage and the product page for each agent. Search and user-triggered agents are scored; training crawlers are reported only. List checked against vendor documentation on 2026-10-07.

TokenVendorTypePurposerobots.txt
OAI-SearchBotOpenAISearch / indexIndexes pages so they can appear in ChatGPT search and shopping results.Honours robots.txt
ChatGPT-UserOpenAIUser-triggeredFetches a page when a ChatGPT user or custom GPT asks about it.May ignore robots.txt
GPTBotOpenAITraining crawlerCollects public content that may be used to train OpenAI models.Honours robots.txt
Claude-SearchBotAnthropicSearch / indexIndexes pages to improve the quality of Claude search results.Honours robots.txt
Claude-UserAnthropicUser-triggeredFetches a page when a Claude user asks a question that needs it.Honours robots.txt
ClaudeBotAnthropicTraining crawlerCollects public content that may be used to train Claude models.Honours robots.txt
PerplexityBotPerplexitySearch / indexIndexes and links pages in Perplexity answers. Perplexity says it is not used for model training.Honours robots.txt
Perplexity-UserPerplexityUser-triggeredVisits a page live to answer a user question. Perplexity says it generally ignores robots.txt.May ignore robots.txt
GooglebotGoogleSearch / indexGoogle Search crawler. Google AI Mode, AI Overviews and Shopping draw on its index.Honours robots.txt
Google-ExtendedGoogleTraining crawlerControl token, not a crawler: opts content out of Gemini training and grounding. No effect on Google Search.Control token
BingbotMicrosoftSearch / indexBing index crawler. Microsoft Copilot answers draw on the Bing index.Honours robots.txt
Applebot-ExtendedAppleTraining crawlerControl token: opts content out of Apple foundation-model training. Applebot itself still crawls for Siri and Spotlight.Control token
Amzn-SearchBotAmazonSearch / indexImproves search in Amazon products such as Alexa. Amazon says it is not used for model training.Honours robots.txt
AmazonbotAmazonTraining crawlerImproves Amazon products and services and may be used to train Amazon AI models.Honours robots.txt
meta-externalagentMetaTraining crawlerCrawls content for training Meta AI models and indexing for Meta products.Honours robots.txt
CCBotCommon CrawlTraining crawlerBuilds the open Common Crawl dataset that many models are trained on.Honours robots.txt
BytespiderByteDanceTraining crawlerCollects content for ByteDance AI models.Honours robots.txt

Limits and ethics

  • One product page. Legible reads the homepage and one product page. Category pages and variants are not crawled. Point the scan at your best seller for the most useful result.
  • No JavaScript. Deliberate: it mirrors how most AI fetchers read pages. Google's crawler renders JavaScript, so Gemini may see more than Legible reports.
  • From the outside. ACP enrolment, Merchant Center feeds and private APIs cannot be verified from public URLs. Those checks are signals, and they say so.
  • Polite by design. At most 15 requests per scan, never more than two at a time, 8-second timeouts, 2.5 MB per response, an identifying user agent, and a rate limit per visitor.
  • Safe by design. Only public http(s) addresses on standard ports. Every redirect hop is re-validated and its DNS checked; private, loopback, link-local and reserved addresses are refused.