Glossary

AI search and web technology, defined

20 terms. Each one traced to the specification or documentation that defines it, not paraphrased from another glossary.

Canonical URL A canonical URL is the address chosen to represent a piece of content when several addresses serve the same or nearly... ClaudeBot ClaudeBot is Anthropic's training crawler. Anthropic describes it as collecting web content that could contribute to ... Entity An entity is a particular thing that can be identified and referred to in its own right: one person, one business, on... Entity authority Entity authority is an industry term for how consistently and credibly a specific business, person or organization is... Google-Extended Google-Extended is a robots.txt control token rather than a crawler. It manages whether content Google has already cr... Googlebot Googlebot is Google's main search crawler. Google says crawling preferences addressed to it "affect Google Search (in... GPTBot GPTBot is OpenAI's training crawler. OpenAI documents it as the robots.txt token that governs whether a site's conten... Grounding Grounding is the practice of tying a model's output to specific sources supplied at the moment of the request, rather... JSON-LD JSON-LD is a JSON-based format for serializing linked data, so an ordinary JSON object can carry types and identifier... Knowledge graph A knowledge graph stores things and the relationships between them rather than documents and keywords. Google introdu... llms.txt llms.txt is a proposed markdown file offering AI agents a short, curated index of a site's most useful pages. It was ... OAI-SearchBot OAI-SearchBot is the OpenAI crawler that decides whether a website can be shown in ChatGPT's search answers. OpenAI s... PerplexityBot PerplexityBot is the crawler Perplexity uses to index pages for its search results. Perplexity states that it "is des... Retrieval-augmented generation Retrieval-augmented generation is a technique in which a language model fetches documents at question time and condit... robots.txt robots.txt is a plain-text file at the root of a site that tells automated clients which paths they may fetch. It is ... Schema markup Schema markup is structured data written with the schema.org vocabulary, a shared set of types and properties maintai... Structured data Structured data is machine-readable markup embedded in a page that states what the page is about in a fixed format, i... User agent A user agent is any client program that initiates an HTTP request, which covers browsers, crawlers, command-line tool... Web crawler A web crawler is an automated client that requests pages over HTTP without a person driving each request, typically f... XML sitemap An XML sitemap is a file that lists the URLs on a site so search engines can learn about pages available for crawling...

These are definitions. The explanations are next door.

A glossary entry tells you what a word means. If you want the mechanism, the Learn section covers it properly.