Glossary · 49 terms · English and Spanish

GEO glossary: 49 AI search terms, defined and sourced

Generative engine optimization (GEO) is the practice of increasing how often and how prominently a source’s content is used and cited in answers written by generative AI search systems. This glossary defines 49 terms from GEO and AI search in one quotable sentence each, labels the strength of the evidence behind every term, and links each fact to a dated primary source.

Not sure which of these terms matter for your brand? See if AI cites me: the free diagnosis runs an initial set of buying prompts on five AI engines.

How to read an entry

Each entry opens with a one-sentence definition you can quote on its own. It then explains the mechanism, gives the origin of the term, corrects common misconceptions with their sources, and lists its sources with dates. Boxes marked “In practice” come from our client work, with names withheld. Every entry carries one evidence label:

Official documentation
The company or body that runs the system documents it: Google, OpenAI, Anthropic, Perplexity, Microsoft or the IETF.
Peer-reviewed
Published in, or accepted by, a peer-reviewed journal or conference.
Industry study
Research by a company, usually a vendor. The entry says who ran it and what they sell.
Practitioner consensus
Widely used and consistent across vendors, but not formally documented or tested.
Contested or unproven
Sources disagree, there is no standard definition, or the claimed effect has not been shown.

Foundations

4 terms

What GEO is, what else it is called, and the engines it deals with.

Generative engine optimization (GEO) #

Foundations Evidence: Peer-reviewed

Generative engine optimization (GEO) is the practice of increasing how often and how prominently a source’s content is used and cited in answers written by generative AI search systems.

Read the full entry

Naming variants: AI SEO, LLMO (LLM optimization), GSO (generative search optimization) and AIO; no primary source we found defines them as separate disciplines. Spanish: optimización de buscadores generativos (Google’s Spanish guide).

The paper that introduced GEO built GEO-bench, 10,000 queries from 25 domains, and rewrote web sources in different ways to test which changes raised their visibility in generated answers. Answers came from gpt-3.5-turbo working on the top five Google results. In the main results (arXiv v3, Table 1), adding quotations raised Position-Adjusted Word Count (a source’s share of the answer’s words, weighted by position) by about 41% over unmodified sources, adding statistics by about 31% and citing sources by about 27%; keyword stuffing scored below the unmodified baseline. On Perplexity.ai the gains reached up to 37%, and results varied by domain. C-SEO Bench (NeurIPS 2025) retested eight of the paper’s methods with newer models and found most such methods “largely ineffective”, with traditional SEO working better.

Origin

Introduced by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande in “GEO: Generative Engine Optimization”, first posted to arXiv on November 16, 2023, and accepted to KDD 2024. In our engine test (September 29, 2026), ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews all defined GEO correctly in English and Spanish, but only Claude credited this paper.

Common misconceptions

  • “GEO is a separate discipline from SEO.” Google’s guide: “From Google Search’s perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.”
  • “Close to 50% of searches already happen inside AI-generated answers.” In the same engine test, Perplexity, answering in Spanish, repeated this claim from an SEO blog without citing any data. The closest independent measure we found: in Pew Research Center’s browsing data for 900 US adults (March 2025), 18% of Google searches produced an AI summary.
Evidence
Peer-reviewed origin; a later peer-reviewed benchmark found most such optimization methods largely ineffective.

Sources

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2023, November 16). GEO: Generative engine optimization. arXiv:2311.09735 (v3, June 28, 2024; accepted to KDD 2024). Figures from Table 1 of v3. arxiv.org
  2. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  3. Puerto, H., Gubri, M., Green, T., Oh, S. J., & Yun, S. (2025, June 6). C-SEO Bench: Does conversational SEO work? NeurIPS 2025 Datasets and Benchmarks Track. arXiv:2506.11097. arxiv.org
  4. Chapekis, A., & Lieb, A. (2025, July 22). Google users are less likely to click on links when an AI summary appears in the results. Pew Research Center. pewresearch.org

Cite this entry

Sourcelift. (2026, September 29). Generative engine optimization (GEO). In GEO glossary. https://sourcelift.ai/geo-glossary#generative-engine-optimization

Last reviewed September 29, 2026

Answer engine optimization (AEO) #

Foundations Evidence: Practitioner consensus

Answer engine optimization (AEO) is the practice of making content more likely to be used or cited when voice assistants and AI search systems answer a question directly.

Read the full entry

Spanish: optimización de buscadores de respuestas (Google’s Spanish guide). Google lists it next to GEO as a label for the same kind of work.

The label grew out of voice search. A Search Engine Watch article of February 7, 2018, argued that voice search was “transforming search engines into ‘answer engines’”, since most voice searches get a single answer read out by a digital assistant, and that optimizing for those answers needed a different strategy from classic SEO. The term now covers AI answers as well. Google’s guide says of AEO and GEO: “These are both terms you may see used to describe work specifically focused on improving visibility in AI search experiences.” It adds that, from Google Search’s perspective, this work is still SEO. Google’s page on third-party SEO advice names both labels and recommends checking such advice against its official guidance. In our engine test (September 29, 2026), every definition of AEO we collected was correct.

Origin

Disputed. Jason Barnard, CEO and founder of Kalicube, which sells brand intelligence services, says he coined the term in 2017. The earliest independent record we found, Rebecca Sentance’s Search Engine Watch article of February 7, 2018, quotes Barnard on the topic but says only that the strategy “has come to be known as AEO”, crediting no one.

Common misconceptions

  • “AEO requires FAQPage or HowTo markup and answer blocks of 40 to 60 words.” In the same test, Gemini gave this advice in English and Spanish. Google’s guide says: “Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add”, and “You don’t need to write in a specific way just for generative AI search.” Google also stopped showing FAQ rich results on May 7, 2026.
Evidence
A practitioner label with a disputed origin; Google names it but treats the work as SEO.

Sources

  1. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  2. Google Search Central. (2026, June 5). Google Search’s guidance on using third-party SEO tools, services, and advice. developers.google.com
  3. Google Search Central. (2026, May 8). Latest documentation updates: Deprecating the FAQ rich result feature. developers.google.com
  4. Sentance, R. (2018, February 7). The rise of answer engine optimization: Why voice search matters. Search Engine Watch. searchenginewatch.com
  5. Kalicube. (2026, May 15). Answer Engine Optimization. Kalicube sells brand intelligence services. kalicube.pro

Cite this entry

Sourcelift. (2026, September 29). Answer engine optimization (AEO). In GEO glossary. https://sourcelift.ai/geo-glossary#answer-engine-optimization

Last reviewed September 29, 2026

Generative engine (answer engine, AI search) #

Foundations Evidence: Peer-reviewed

A generative engine is a search system that uses a large language model to gather information from several sources and summarize it into one answer to the user’s query.

Read the full entry

Spanish: buscador generativo, buscador de respuestas, búsqueda con IA (Google’s Spanish guide says búsqueda con IA generativa).

The GEO paper introduced the term for search engines “that use generative models to gather and summarize information to answer user queries”, typically by “synthesizing information from multiple sources”. Answer engine is older: in 2018 Search Engine Watch used it for search engines that give one spoken answer to a voice query, and Perplexity, which presents its product as one, defines an answer engine as “a tool designed to give you direct, detailed answers to your questions”. AI search has no formal definition; Google applies it to “generative AI features on Google Search (such as AI Overviews and AI Mode)”. The three labels overlap: each describes a system that replies with a written answer built from sources, where a classic search engine returns a list of links.

Origin

Generative engine: Aggarwal et al., first posted to arXiv on November 16, 2023 (“we formalize under the unified framework of generative engines (GEs)”). Answer engine: unknown; already used for voice search in 2018. AI search: unknown; a generic label.

Common misconceptions

  • “Generative engines are replacing traditional search.” The GEO paper’s abstract said in 2023 that they were “rapidly replacing traditional search engines like Google and Bing”. Panel data does not show that for Google: SparkToro, using Similarweb data, found that only 0.34% of US Google searches made their way to AI Mode from January to April 2026. The study does not cover ChatGPT or other assistants.
Evidence
Only generative engine has a peer-reviewed definition; answer engine and AI search are industry labels.

Sources

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2023, November 16). GEO: Generative engine optimization. arXiv:2311.09735 (accepted to KDD 2024). arxiv.org
  2. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  3. Perplexity. (n.d.). What is an answer engine and how does Perplexity work? Perplexity Help Center. perplexity.ai
  4. Sentance, R. (2018, February 7). The rise of answer engine optimization: Why voice search matters. Search Engine Watch. searchenginewatch.com
  5. Fishkin, R. (2026, June 8). In 2026, less than one third of Google searches still send a click. SparkToro, with Similarweb data. SparkToro sells audience research software. sparktoro.com

Cite this entry

Sourcelift. (2026, September 29). Generative engine. In GEO glossary. https://sourcelift.ai/geo-glossary#generative-engine

Last reviewed September 29, 2026

How AI search works

11 terms

How AI systems find, choose and cite the pages behind an answer.

Large language model (LLM) #

How AI search works Evidence: Official documentation

A large language model (LLM) is a language model with a very large number of parameters, trained on large amounts of text to predict the next token.

Read the full entry

Spanish: modelo de lenguaje grande (Google’s Spanish glossary); Google’s Spanish documentation also uses modelo de lenguaje extenso.

Google’s Machine Learning Glossary defines an LLM as “At a minimum, a language model having a very high number of parameters” and, informally, “any Transformer-based language model, such as Gemini or GPT”. An LLM writes by predicting one token after another; Google’s Machine Learning Crash Course calls LLMs “essentially autocomplete mechanisms”. What the model learns in training is stored in its parameters, which can number in the hundreds of billions or trillions, so fluent output can still be wrong: the course says “LLMs hallucinate, meaning their predictions often contain mistakes.” In AI search, the model is paired with a retrieval step: the system first finds pages, then the model writes the answer from them. An answer can therefore reflect training, retrieved pages or both.

Origin

Who first used the phrase in today’s sense is unknown. It already appears in the title of a June 2007 paper, “Large Language Models in Machine Translation” by Thorsten Brants, Ashok C. Popat, Peng Xu, Franz J. Och and Jeffrey Dean (EMNLP-CoNLL 2007).

Common misconceptions

  • “LLMs look facts up.” On its own, a model generates text from what it learned in training: Google’s crash course says LLMs “predict a token or sequence of tokens”. Facts from the web reach an answer only when the system adds a retrieval or grounding step.

Sources

  1. Google for Developers. (2026, April 13). Machine Learning Glossary: Generative AI. developers.google.com
  2. Google for Developers. (2026, January 2). LLMs: What’s a large language model? Machine Learning Crash Course. developers.google.com
  3. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  4. Brants, T., Popat, A. C., Xu, P., Och, F. J., & Dean, J. (2007, June). Large language models in machine translation. Proceedings of EMNLP-CoNLL 2007, 858-867. ACL Anthology. aclanthology.org

Cite this entry

Sourcelift. (2026, September 29). Large language model (LLM). In GEO glossary. https://sourcelift.ai/geo-glossary#large-language-model

Last reviewed September 29, 2026

Retrieval-augmented generation (RAG) #

How AI search works Evidence: Peer-reviewed

Retrieval-augmented generation (RAG) is a technique in which a system retrieves relevant documents from an external source and gives them to a language model as context for its answer.

Read the full entry

Spanish: generación aumentada por recuperación (RAG). Google Search Central treats RAG and grounding as synonyms.

Lewis and colleagues described RAG models as ones “which combine pre-trained parametric and non-parametric memory for language generation”: a sequence-to-sequence model holds the parametric memory, and a dense vector index of Wikipedia, searched by a neural retriever, holds the non-parametric memory. The retrieved passages give the model information it did not store in training. Google’s guide defines RAG as “A technique (also known as grounding) used to improve the quality, accuracy, and freshness of AI responses”, and says it relies on Google’s core Search ranking systems to retrieve relevant, up-to-date pages from the Search index. So for Google’s AI features, a page has to be retrievable first: to appear as a supporting link in AI Overviews or AI Mode, it “must be indexed and eligible to be shown in Google Search with a snippet”.

Origin

Named by Patrick Lewis and 11 co-authors in “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, first posted to arXiv on May 22, 2020, and accepted at NeurIPS 2020. In our engine test (September 29, 2026), none of the ten answers about RAG that we read credited this paper.

Common misconceptions

  • “RAG eliminates hallucinations.” Magesh et al. (Journal of Empirical Legal Studies, 2025) tested RAG-based legal research tools whose providers had claimed to avoid or eliminate hallucinations; the tools from LexisNexis and Thomson Reuters “each hallucinate between 17% and 33% of the time”.

Sources

  1. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020, May 22). Retrieval-augmented generation for knowledge-intensive NLP tasks. NeurIPS 2020. arXiv:2005.11401. arxiv.org
  2. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  3. Google Search Central. (2025, December 10). AI features and your website. developers.google.com
  4. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. (2025). Hallucination-free? Assessing the reliability of leading AI legal research tools. Journal of Empirical Legal Studies, 22(2), 216-242. arxiv.org

Cite this entry

Sourcelift. (2026, September 29). Retrieval-augmented generation (RAG). In GEO glossary. https://sourcelift.ai/geo-glossary#retrieval-augmented-generation

Last reviewed September 29, 2026

Grounding #

How AI search works Evidence: Official documentation

Grounding is the practice of basing a generative AI model’s answer on verifiable external information, such as search results or documents, rather than only on its training.

Read the full entry

Spanish: fundamentación (Google’s Spanish documentation). Google Search Central uses it as a synonym of retrieval-augmented generation (RAG).

Google uses the word in two ways. In the Gemini API, “Grounding with Google Search connects the Gemini model to real-time web content”, and the response carries annotations linking parts of the answer to their sources; Google says this lets Gemini “cite verifiable sources beyond its knowledge cutoff”. Google Search Central, by contrast, defines retrieval-augmented generation as “also known as grounding”. For teaching, it helps to keep them apart: grounding is the goal, an answer tied to sources that can be checked, and RAG is one way to reach it. Google’s AI features page also separates grounding from training: to limit either in some of Google’s other systems, it points site owners to Google-Extended.

Origin

The word is older than AI search: in “The Symbol Grounding Problem” (Physica D, June 1990), Stevan Harnad asked how the symbols of a formal system can carry meaning of their own. Who first used it for tying model answers to sources is unknown.

Common misconceptions

  • “A grounded answer is always correct.” Google presents Grounding with Google Search as a way to “Reduce model hallucinations”, not to remove them, and its AI Overviews help page warns: “AI responses may include mistakes.”

Sources

  1. Google AI for Developers. (2026, September 23). Grounding with Google Search. Gemini API documentation. ai.google.dev
  2. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  3. Google Search Central. (2025, December 10). AI features and your website. developers.google.com
  4. Google Search Help. (n.d.). Find information in faster & easier ways with AI Overviews in Google Search. support.google.com
  5. Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1-3), 335-346. arxiv.org

Cite this entry

Sourcelift. (2026, September 29). Grounding. In GEO glossary. https://sourcelift.ai/geo-glossary#grounding

Last reviewed September 29, 2026

Query fan-out #

How AI search works Evidence: Official documentation

Query fan-out is a technique in which an AI search system turns one question into several related searches, runs them concurrently and combines the results into one answer.

Read the full entry

Also called fan-out queries. Spanish: ramificación de búsquedas (Google also uses distribución múltiple de consultas).

Google named the technique when it launched AI Mode, and its documentation says both AI Overviews and AI Mode may use it. Google’s own example: for “how to fix a lawn that’s full of weeds”, fan-out queries might include “best herbicides for lawns”, “remove weeds without chemicals” and “how to prevent weeds in lawn”. Each sub-query brings back its own results, so a page can be cited for one part of a question without ranking for the question as a whole. Google does not publish how many sub-queries a question produces or which ones it runs, so tools that list “your fan-out queries” are estimating them.

Origin

Google’s AI Mode launch post (Robby Stein, VP of Product, Google Search, March 5, 2025) said AI Mode “uses a ‘query fan-out’ technique, issuing multiple related searches concurrently across subtopics”. It is the earliest use of the name we found.

Common misconceptions

  • “Publish a page for every fan-out query.” Google’s guide says that creating separate content for every variation of a query, fan-out queries included, “primarily to manipulate rankings or generative AI responses in Google Search violates Google’s scaled content abuse spam policy.”
  • “Fan-out runs a known number of searches.” Google publishes no figure. In our engine test (September 29, 2026), Google’s AI Overview said fan-out typically runs 5 to 12 searches and Gemini, asked in Spanish, said “between 5 and more than 20”, neither with a source. Claude (in both languages) and ChatGPT (in English) said there is no official number.

Sources

  1. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  2. Google Search Central. (2025, December 10). AI features and your website. developers.google.com
  3. Stein, R. (2025, March 5). Expanding AI Overviews and introducing AI Mode. Google. blog.google

Cite this entry

Sourcelift. (2026, September 29). Query fan-out. In GEO glossary. https://sourcelift.ai/geo-glossary#query-fan-out

Last reviewed September 29, 2026

Training data vs live retrieval (knowledge cutoff) #

How AI search works Evidence: Peer-reviewed

Training data vs live retrieval is the distinction between what a model learned in training, up to its knowledge cutoff, and what it fetches from the web when answering.

Read the full entry

Also framed as parametric vs non-parametric memory. Spanish: datos de entrenamiento frente a búsqueda en tiempo real; knowledge cutoff is fecha límite de conocimiento in OpenAI’s Spanish help center.

Research calls the two sources parametric and non-parametric memory. Petroni et al. (2019) analyzed the relational knowledge that pre-trained language models absorb from their training data, and Lewis et al. (2020), starting from the finding that such models “store factual knowledge in their parameters”, added a searchable index as a second memory. Parametric knowledge ends at the model’s knowledge cutoff, the date of its most recent training data, and stays fixed until the model is retrained. Retrieved knowledge can be as recent as the index being searched. For GEO, the consequence is direct: a page published after a model’s cutoff can reach that model’s answers only through retrieval, while material the model absorbed in training can shape answers even when nothing is searched.

Origin

The parametric vs non-parametric framing comes from Lewis et al., first posted to arXiv on May 22, 2020. Petroni et al. had earlier analyzed the relational knowledge held in pre-trained models in “Language Models as Knowledge Bases?” (first posted September 3, 2019; EMNLP 2019). Who coined “knowledge cutoff” is unknown.

Common misconceptions

  • “AI assistants always check the live web.” OpenAI’s help center says its models “are trained on data up to a certain point” and responses “do not incorporate information about events beyond that, unless tools are used”; ChatGPT “may search the web automatically when your question would benefit from current information”.
  • “A model’s knowledge stops cleanly at its stated cutoff.” Cheng et al. (“Dated Data”, 2024) found that “effective cutoffs often differ from reported cutoffs”, partly because newer web crawls contain older material and deduplication misses near-duplicates.

Sources

  1. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020, May 22). Retrieval-augmented generation for knowledge-intensive NLP tasks. NeurIPS 2020. arXiv:2005.11401. arxiv.org
  2. Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., & Riedel, S. (2019, September 3). Language models as knowledge bases? EMNLP 2019. arXiv:1909.01066. arxiv.org
  3. Cheng, J., Marone, M., Weller, O., Lawrie, D., Khashabi, D., & Van Durme, B. (2024, March 19). Dated data: Tracing knowledge cutoffs in large language models. arXiv:2403.12958. arxiv.org
  4. OpenAI. (n.d.). Does ChatGPT tell the truth? OpenAI Help Center. help.openai.com
  5. OpenAI. (n.d.). Searching the web with ChatGPT. OpenAI Help Center. help.openai.com

Cite this entry

Sourcelift. (2026, September 29). Training data vs live retrieval. In GEO glossary. https://sourcelift.ai/geo-glossary#training-data-vs-live-retrieval

Last reviewed September 29, 2026

Citation (source attribution) #

How AI search works Evidence: Peer-reviewed

A citation in an AI answer is a link or numbered reference to a source that the answer presents as support for a specific statement.

Read the full entry

Also called source attribution or a source link. Google’s documentation calls the links in AI Overviews and AI Mode “supporting links”. Spanish: cita.

AI search products attach citations so readers can check where a statement came from: ChatGPT responses that use web search may include citations, Perplexity adds numbered citations to each answer, and Google’s AI Overviews and AI Mode show supporting links to web pages. Research treats this as a question of attribution: the Attributable to Identified Sources (AIS) framework of Rashkin et al. holds that generated statements about the world should be verifiable against an identified, independent source. A link alone does not guarantee this. In a human evaluation of four generative search engines (Bing Chat, NeevaAI, perplexity.ai and YouChat), Liu, Zhang and Liang found that, on average, only 51.5% of generated sentences were fully supported by their citations. For a website, being cited means one of its pages is linked inside the answer, which differs from the brand being named without a link.

Origin

Unknown in the product sense. The research framing comes from Rashkin et al.: the Attributable to Identified Sources (AIS) framework, posted on arXiv on December 23, 2021, and published in Computational Linguistics in December 2023. Liu, Zhang and Liang (Findings of EMNLP 2023) measured citation quality in generative search engines as citation recall and citation precision.

Common misconceptions

  • “A citation proves the sentence next to it.” In the four engines Liu, Zhang and Liang audited, “only 74.5% of citations support their associated sentence.” OpenAI’s own help page tells users to “Open a cited source to check that it supports the answer.”

Sources

  1. Rashkin, H., Nikolaev, V., Lamm, M., Aroyo, L., Collins, M., Das, D., Petrov, S., Tomar, G. S., Turc, I., & Reitter, D. (2023). Measuring attribution in natural language generation models. Computational Linguistics, 49(4), 777-840. First posted on arXiv on December 23, 2021. aclanthology.org
  2. Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating verifiability in generative search engines. Findings of the Association for Computational Linguistics: EMNLP 2023, 7001-7025. aclanthology.org
  3. Google Search Central. (2025, December 10). AI features and your website. developers.google.com
  4. OpenAI. (n.d.). Searching the web with ChatGPT. OpenAI Help Center. help.openai.com
  5. Perplexity. (n.d.). How does Perplexity work? Perplexity Help Center. perplexity.ai

Cite this entry

Sourcelift. (2026, September 29). Citation. In GEO glossary. https://sourcelift.ai/geo-glossary#citation

Last reviewed September 29, 2026

Google AI Overviews #

How AI search works Evidence: Official documentation

Google AI Overviews are AI-generated summaries shown in some Google Search results, giving a short answer to the query with links to web pages that support it.

Read the full entry

Tested from 2023 as SGE (Search Generative Experience) in Search Labs. Spanish: vistas creadas con IA, Google’s official name. In our engine test (September 29, 2026), ChatGPT and Claude, answering in Spanish, used other names for it, and Google’s own AI Overview used the English name.

Google first tested the feature in SGE (Search Generative Experience), a Search Labs experiment announced on May 10, 2023, and on May 14, 2024, began rolling out AI Overviews to everyone in the US, built on a Gemini model customized for Google Search. They arrived in Spain, in Spanish and English, on March 25, 2025, for signed-in users aged 18 or over. On May 20, 2025, Google said they were available in more than 200 countries and territories and in more than 40 languages. According to Google’s documentation, AI Overviews and AI Mode may use query fan-out, issuing several related searches across subtopics and data sources, and they show a wider and more varied set of links than a classic web search.

Origin

Google. Elizabeth Reid (Vice President and GM, Search) announced SGE in Search Labs on May 10, 2023. On May 14, 2024, Google began rolling out AI Overviews to everyone in the US, noting that people had already used them billions of times in the Search Labs experiment.

Common misconceptions

  • “AI Overviews are available in about 120 countries.” That figure is out of date: on May 20, 2025, Google announced availability in more than 200 countries and territories and in more than 40 languages. In our engine test (September 29, 2026), Google’s own AI Overview, asked in English, still said the feature was available in more than 120 countries.
  • “You need special markup or an AI file to appear in AI Overviews.” Google: “There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.”

Sources

  1. Google Search Central. (2025, December 10). AI features and your website. developers.google.com
  2. Reid, E. (2023, May 10). Supercharging Search with generative AI. Google. blog.google
  3. Reid, E. (2024, May 14). Generative AI in Search: Let Google do the searching for you. Google. blog.google
  4. Budaraju, H. (2025, March 25). We’re bringing the helpfulness of AI Overviews to more countries in Europe. Google. blog.google
  5. Budaraju, H. (2025, May 20). AI Overviews are now available in over 200 countries and territories, and more than 40 languages. Google. blog.google

Cite this entry

Sourcelift. (2026, September 29). Google AI Overviews. In GEO glossary. https://sourcelift.ai/geo-glossary#ai-overviews

Last reviewed September 29, 2026

Google AI Mode #

How AI search works Evidence: Official documentation

Google AI Mode is a conversational search mode in Google Search that answers complex questions with AI-generated responses, running several related searches and linking to supporting web pages.

Read the full entry

Spanish: modo IA, the name Google uses in its Spanish documentation.

Google introduced AI Mode on March 5, 2025, as an early experiment in Labs, inviting Google One AI Premium subscribers first, and began rolling it out to everyone in the US on May 20, 2025. It started rolling out globally in Spanish on September 23, 2025. Google describes it as “particularly helpful for queries where further exploration, reasoning, or complex comparisons are needed”. It relies on query fan-out: it breaks a question into subtopics, issues many searches at once and combines the results into one response with links. SparkToro, using Similarweb’s US desktop and mobile web panel (which excludes the Google app), found that 0.34% of Google searches reached AI Mode from January to April 2026.

Origin

Google. Robby Stein (VP of Product, Google Search) introduced AI Mode as a Labs experiment on March 5, 2025, in the same post that described its “query fan-out” technique.

Common misconceptions

  • “AI Mode launched in 2026.” Google introduced it in Labs on March 5, 2025, and began rolling it out to everyone in the US on May 20, 2025.
  • “Pages need separate optimization to appear in AI Mode.” Google’s documentation sets the same bar as for Search: to be shown as a supporting link in AI Overviews or AI Mode, a page “must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements.”
Evidence
The usage figure comes from an industry study by SparkToro, which sells audience research software, using data from Similarweb, which sells web traffic data.

Sources

  1. Stein, R. (2025, March 5). Expanding AI Overviews and introducing AI Mode. Google. blog.google
  2. Reid, E. (2025, May 20). AI in Search: Going beyond information to intelligence. Google. blog.google
  3. Google. (2025, September 23). Try AI Mode in Spanish today. blog.google
  4. Google Search Central. (2025, December 10). AI features and your website. developers.google.com
  5. Fishkin, R. (2026, June 8). In 2026, less than one third of Google searches still send a click. SparkToro, with data from Similarweb. SparkToro sells audience research software; Similarweb sells web and app traffic data. sparktoro.com

Cite this entry

Sourcelift. (2026, September 29). Google AI Mode. In GEO glossary. https://sourcelift.ai/geo-glossary#ai-mode

Last reviewed September 29, 2026

Hallucination #

How AI search works Evidence: Peer-reviewed

Hallucination is output from a generative AI model that sounds plausible but is false, or that the source material it was given does not support.

Read the full entry

Also: hallucinated content. Research distinguishes intrinsic from extrinsic hallucination. Spanish: alucinación.

Two definitions are in use. Google’s machine learning glossary treats it as a factual error: plausible-seeming output that is factually incorrect, from a model presenting it as a claim about the real world. Ji et al.’s survey (ACM Computing Surveys, 2023) measures it against the source content instead and, following earlier work, splits it in two: intrinsic hallucination contradicts the source, and extrinsic hallucination can be neither supported nor contradicted by it. In AI search, where the source is the set of retrieved pages, an intrinsic hallucination misstates what a cited page says and an extrinsic one adds a claim that no retrieved page makes. Research published by OpenAI (vendor research, September 5, 2025) argues that “language models hallucinate because standard training and evaluation procedures reward guessing over acknowledging uncertainty.”

Origin

Ji et al. note that the term “first appeared in Computer Vision (CV) in Baker and Kanade” with a positive meaning (for example, super-resolution), before work on image captioning and object detection used it for errors, the sense closest to its use in text generation. Maynez et al. (ACL 2020) had already split hallucinations in generated summaries into intrinsic and extrinsic.

Common misconceptions

  • “Hallucination means the model invented a fact.” Ji et al. also count intrinsic hallucination, output that “contradicts the source content”: an answer can misstate a real page that it cites.
  • “Answers that link to live search results don’t hallucinate.” Google’s help page on AI Overviews says: “AI Overviews can and will make mistakes.”

Sources

  1. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1-38. Open version on arXiv (2202.03629). arxiv.org
  2. Maynez, J., Narayan, S., Bohnet, B., & McDonald, R. (2020). On faithfulness and factuality in abstractive summarization. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 1906-1919. aclanthology.org
  3. Google for Developers. (2026, April 13). Machine Learning Glossary: Generative AI. developers.google.com
  4. Google. (n.d.). Find information in faster & easier ways with AI Overviews in Google Search. Google Search Help. support.google.com
  5. OpenAI. (2025, September 5). Why language models hallucinate. Vendor research: OpenAI sells ChatGPT and access to its models. openai.com

Cite this entry

Sourcelift. (2026, September 29). Hallucination. In GEO glossary. https://sourcelift.ai/geo-glossary#hallucination

Last reviewed September 29, 2026

AI agents (agentic browsing) #

How AI search works Evidence: Official documentation

AI agents are autonomous AI systems that carry out tasks for people, such as making a reservation or comparing products, and may visit and operate websites to do so.

Read the full entry

Also: browser agents. Google’s guide speaks of “agentic experiences”. “Agentic search” is also used, with no formal definition. Spanish: agentes de IA.

Google’s guide to its generative AI features defines AI agents as autonomous systems that can perform tasks on behalf of people, and says browser agents may gather what they need from a site by analyzing screenshots, inspecting the DOM and interpreting the accessibility tree. The web.dev guide (April 1, 2026) adds that modern agents combine these inputs: they take a structured list of interactive elements from the DOM and the accessibility tree and cross-check it against a visual rendering. Its advice follows from that: semantic HTML (real buttons and links rather than styled div elements), labels linked to their inputs, stable layouts and no transparent overlays over controls. OpenAI’s ChatGPT agent (July 17, 2025) works on its own virtual computer with a visual browser, a text-based browser and a terminal; ChatGPT Atlas (October 21, 2025) is a browser built around ChatGPT, with an agent mode.

Origin

Unknown. “Agent” also has an established sense in reinforcement learning, which Google’s machine learning glossary records alongside the current one. We found no coinage or formal definition for “agentic search” or “agentic browsing”.

Common misconceptions

  • “Agents read a site the way search crawlers do.” Google says browser agents may analyze “visual renderings (like screenshots)”, inspect “the DOM structure” and interpret “the accessibility tree”. The web.dev guide warns that an agent’s visual analysis “might discard nodes that are covered, even if the node appears transparent.”
Evidence
“Agentic search” has no formal definition; Google documents AI agents and browser agents.

Sources

  1. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  2. Kulikowski, K., & More, O. (2026, April 1). Build agent-friendly websites. web.dev. web.dev
  3. OpenAI. (2025, July 17). Introducing ChatGPT agent: bridging research and action. openai.com
  4. OpenAI. (2025, October 21). Introducing ChatGPT Atlas. openai.com
  5. Google for Developers. (2026, April 13). Machine Learning Glossary: Generative AI. developers.google.com

Cite this entry

Sourcelift. (2026, September 29). AI agents. In GEO glossary. https://sourcelift.ai/geo-glossary#ai-agents

Last reviewed September 29, 2026

Measurement

9 terms

How AI visibility is counted. Every number here depends on the prompts, engines and dates behind it.

AI visibility #

Measurement Evidence: Practitioner consensus

AI visibility is an umbrella term for how often and how prominently a brand or website appears, as a mention or cited source, in AI-generated answers.

Read the full entry

Tools use the same label for different metrics: Semrush’s AI Visibility score and Peec AI’s Visibility. Spanish: visibilidad en IA.

It has no single formula. Semrush’s AI Visibility is a benchmark score from 0 to 100 combining topic coverage (how many topics include the brand in AI answers, compared with all other domains) and mention consistency (how often the brand is mentioned across responses within those topics). Peec AI’s Visibility is a mention rate: responses mentioning the brand ÷ total responses × 100. The GEO paper defined a website’s visibility in a generative engine’s answer through impression metrics, such as the share of the answer’s words in sentences citing it. Microsoft’s AI Performance report in Bing Webmaster Tools (public preview, February 10, 2026) tracks citations instead, with Total Citations, Average Cited Pages and Grounding Queries (the phrases the AI used when retrieving cited content). Two visibility figures are comparable only if calculated the same way, on the same prompts, engines and dates.

Origin

Unknown. No one is documented as having coined the phrase. Aggarwal et al. gave a formal definition of a website’s visibility in generative engine answers in the GEO paper (arXiv, November 16, 2023; accepted to KDD 2024).

Common misconceptions

  • “AI visibility tools measure it exactly.” Semrush’s own data page says: “AI search and LLM responses are fast-changing and highly personalized, which means no platform can provide exact numbers on visibility.”
Evidence
Practitioners agree on the broad meaning, but each tool calculates it with its own formula.

Sources

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2023, November 16). GEO: Generative engine optimization. arXiv:2311.09735 (v3, June 28, 2024; accepted to KDD 2024). arxiv.org
  2. Madhavan, K., Merchant, M., Canel, F., & Nigam, S. (2026, February 10). Introducing AI Performance in Bing Webmaster Tools public preview. Microsoft Bing Blogs. blogs.bing.com
  3. Semrush. (n.d.). AI visibility metrics. Semrush Knowledge Base. Semrush sells AI visibility software. semrush.com
  4. Semrush. (n.d.). Where does the data in Semrush’s AI Visibility Toolkit come from? Semrush Knowledge Base. Semrush sells AI visibility software. semrush.com
  5. Peec AI. (n.d.). Visibility. Peec AI Docs. Peec AI sells AI search tracking software. docs.peec.ai

Cite this entry

Sourcelift. (2026, September 29). AI visibility. In GEO glossary. https://sourcelift.ai/geo-glossary#ai-visibility

Last reviewed September 29, 2026

AI share of voice #

Measurement Evidence: Contested or unproven

AI share of voice is a brand’s share of all mentions or citations that it and named competitors receive in AI answers to a fixed set of prompts.

Read the full entry

Also called AI SOV. Bing Webmaster Tools reports a citations-only version called Citation Share. Spanish: cuota de voz en IA.

The term comes from advertising, where share of voice meant a brand’s share of the media exposure in its category. AI tools count different things under the same name: Semrush counts the brand’s mentions against competitors’, Ahrefs weights mentions by the search volume of the prompts (a share of “impressions”), and Microsoft’s Bing Webmaster Tools reports Citation Share, a site’s share of all citations shown for one grounding query. A figure means little without its formula, prompt set, engines, dates and competitor list. Answers also vary from run to run: in a SparkToro study of 2,961 runs, there was less than a 1 in 100 chance that ChatGPT or Google’s AI returned the same list of brands twice, so the share should be averaged over many prompts and repeated runs.

Origin

Share of voice is an advertising metric. John Philip Jones tied it to market share in the Harvard Business Review in 1990, describing it as “the brand’s share of the total value of the main media exposure in that product category”. Who first applied it to AI answers is unknown.

Common misconceptions

  • “AI share of voice is one standard number.” The tools above count mentions, search-volume-weighted impressions or citations per query, so two tools can report different shares for the same brand on the same day. Check the formula before comparing.
Evidence
No standard formula: vendors define it differently.

Sources

  1. Jones, J. P. (1990). Ad spending: Maintaining market share. Harvard Business Review, 68(1), 38-42. pubmed.ncbi.nlm.nih.gov
  2. Semrush. (n.d.). AI visibility metrics. Semrush Knowledge Base. Semrush sells AI visibility software. semrush.com
  3. Ahrefs. (2026). AI visibility metrics. Ahrefs Help Center. Ahrefs sells Brand Radar. help.ahrefs.com
  4. Madhavan, K., Merchant, M., Nigam, S., & Shah, T. (2026, June 16). New AI visibility insights in Bing Webmaster Tools: Intents, topics, Citation Share, compare. Microsoft Bing Blogs. blogs.bing.com
  5. Fishkin, R. (2026, January 28). New research: AIs are highly inconsistent when recommending brands or products. SparkToro, with Gumshoe.ai. SparkToro sells audience research software. sparktoro.com

Cite this entry

Sourcelift. (2026, September 29). AI share of voice. In GEO glossary. https://sourcelift.ai/geo-glossary#ai-share-of-voice

Last reviewed September 29, 2026

Mention rate #

Measurement Evidence: Practitioner consensus

Mention rate is the percentage of AI answers to a fixed set of prompts that name a brand at least once, whether or not they cite it.

Read the full entry

Peec AI calls it Visibility (Visibility Score). Spanish: tasa de menciones.

Formula: mention rate = answers that mention the brand ÷ total answers × 100. Vendors calculate it the same way under different names: Peec AI’s Visibility Score uses exactly this formula, and ZipTie calls it mention rate (“AI responses naming your brand ÷ Total relevant AI responses monitored”). Semrush instead reports Mentions as a count: “The total number of prompts in which a brand is included in AI responses.” A mention only requires the brand to appear in the answer text; no link or citation is needed. The rate depends on the prompt set and on the run: in a SparkToro study, the same prompt rarely returned the same list of brands twice. Average it over many prompts and repeated runs, and report it with its prompts, engines and dates.

Origin

Unknown. It is a vendor metric with no documented coinage.

Common misconceptions

  • “A mention is the same as a citation.” Ahrefs counts them separately: a mention when “a brand appears at least once in an AI generated response”, a citation when “a page appears at least once as a cited source in an AI generated response”. The two counts can differ for the same brand.
Evidence
Vendors agree on the formula but not on the name.

Sources

  1. Peec AI. (n.d.). Visibility. Peec AI Docs. Peec AI sells AI search tracking software. docs.peec.ai
  2. Ahmed, I. (2026, March 23; updated 2026, September 15). Citation rate vs mention rate: Two AI visibility metrics every marketer must track. ZipTie. ZipTie sells AI search monitoring software. ziptie.ai
  3. Ahrefs. (2026). AI visibility metrics. Ahrefs Help Center. Ahrefs sells Brand Radar. help.ahrefs.com
  4. Semrush. (n.d.). AI visibility metrics. Semrush Knowledge Base. Semrush sells AI visibility software. semrush.com
  5. Fishkin, R. (2026, January 28). New research: AIs are highly inconsistent when recommending brands or products. SparkToro, with Gumshoe.ai. SparkToro sells audience research software. sparktoro.com

Cite this entry

Sourcelift. (2026, September 29). Mention rate. In GEO glossary. https://sourcelift.ai/geo-glossary#mention-rate

Last reviewed September 29, 2026

Citation rate #

Measurement Evidence: Contested or unproven

Citation rate is how often AI answers cite a page or website as a source, divided by a base (answers monitored, or times retrieved) that differs between tools.

Read the full entry

Spanish: tasa de citas. Bing Webmaster Tools reports Total Citations and Citation Share instead.

No standard formula exists. Peec AI defines it as “how often a URL is cited per AI answer relative to how often it is retrieved”: citations ÷ retrievals, where a retrieval means the page made it into the list of candidates the model might use for its answer. Because one answer can cite a URL several times, this rate can exceed 1; in Peec’s benchmarks, 2.0 means a URL “was cited twice per every AI answer” on average. ZipTie divides “AI responses citing your URL” by “Total relevant AI responses monitored”, a share that cannot exceed 100%. Microsoft’s first-party metrics in Bing Webmaster Tools include Total Citations, the number of citations shown as sources in AI answers over a period, and Citation Share, a site’s percentage of all citations shown for one grounding query. Check which base a figure uses before comparing it.

Origin

Unknown. No one is documented as having coined it.

Common misconceptions

  • “Citation rate is one standard number.” Peec AI divides citations by retrievals; ZipTie divides answers citing the URL by all answers monitored. The same page can get different rates from two tools on the same day.
Evidence
No standard formula: vendors divide by different bases.

Sources

  1. Madhavan, K., Merchant, M., Canel, F., & Nigam, S. (2026, February 10). Introducing AI Performance in Bing Webmaster Tools public preview. Microsoft Bing Blogs. blogs.bing.com
  2. Madhavan, K., Merchant, M., Nigam, S., & Shah, T. (2026, June 16). New AI visibility insights in Bing Webmaster Tools: Intents, topics, Citation Share, compare. Microsoft Bing Blogs. blogs.bing.com
  3. Peec AI. (2026, February). Citation rate benchmarks differ sharply by engine. Peec AI Docs. Peec AI sells AI search tracking software. docs.peec.ai
  4. Wells, T. (2026, February 27). What does a good citation rate look like? Benchmarks from over 1 million AI citations. Peec AI Blog. Peec AI sells AI search tracking software. peec.ai
  5. Ahmed, I. (2026, March 23; updated 2026, September 15). Citation rate vs mention rate: Two AI visibility metrics every marketer must track. ZipTie. ZipTie sells AI search monitoring software. ziptie.ai

Cite this entry

Sourcelift. (2026, September 29). Citation rate. In GEO glossary. https://sourcelift.ai/geo-glossary#citation-rate

Last reviewed September 29, 2026

Citation gap #

Measurement Evidence: Contested or unproven

A citation gap is a prompt, or a source cited in AI answers, where named competitors appear but a given brand does not, measured against a stated prompt set.

Read the full entry

Related vendor terms: AI visibility gap and Missing Sources (Semrush), brand gap analysis (Ahrefs). Spanish: brecha de citas.

No standard definition exists, and every tool builds the gap from its own prompts and competitor list. Semrush’s Missing Sources report lists “External websites cited in AI answers that mention competitors but not your brand.” Semrush’s guide to AI visibility gaps separates five kinds (mention, prompt, source, citation and narrative) and describes the citation gap as: “The AI doesn’t cite content from your website in its responses.” Ahrefs calls its version brand gap analysis, which “measures the difference between your brand’s potential visibility and its actual presence across Google, AI results, and the wider web”. A gap is only meaningful with its prompt set, engines, dates and competitors stated; change any of them and the gap can change. Answers also vary between runs: in a SparkToro study, the same prompt rarely returned the same list of brands twice, so a gap seen once may not hold.

Origin

Unknown. It is a practitioner term with no documented coinage and no shared definition.

Evidence
No standard definition: each tool computes the gap from its own prompts and competitor list.

Sources

  1. Semrush. (n.d.). AI visibility metrics. Semrush Knowledge Base. Semrush sells AI visibility software. semrush.com
  2. Harsel, L. (2026, July 27). How to find AI visibility gaps with Semrush. Semrush Blog. Semrush sells AI visibility software. semrush.com
  3. Gavoyannis, D. (2025, October 30; modified 2026, June 5). Brand gap analysis: Find out why you’re invisible in AI search. Ahrefs Blog. Ahrefs sells Brand Radar. ahrefs.com
  4. Fishkin, R. (2026, January 28). New research: AIs are highly inconsistent when recommending brands or products. SparkToro, with Gumshoe.ai. SparkToro sells audience research software. sparktoro.com

Cite this entry

Sourcelift. (2026, September 29). Citation gap. In GEO glossary. https://sourcelift.ai/geo-glossary#citation-gap

Last reviewed September 29, 2026

Prompt set and prompt tracking (branded vs non-branded prompts) #

Measurement Evidence: Industry study

A prompt set is a fixed list of questions run through AI assistants and AI search features to measure visibility; prompt tracking repeats those runs on a schedule.

Read the full entry

Semrush’s feature is called Prompt Tracking. Search Console classifies Google Search queries, not AI prompts, as branded or non-branded. Spanish: set de prompts, seguimiento de prompts.

Because AI answers change from run to run, visibility is measured over many prompts and runs. In a SparkToro study with Gumshoe.ai (January 28, 2026), 600 volunteers ran 12 prompts through ChatGPT, Claude and Google’s AI Overviews 2,961 times; there was less than a 1 in 100 chance that ChatGPT or Google’s AI gave the same list of brands in any two responses. Rand Fishkin concluded that visibility “across dozens to hundreds of prompts run multiple times is a reasonable metric”. Semrush tracks “daily visibility for a custom set of prompts”; Peec AI reruns prompts every 24 hours. Branded prompts name the brand; non-branded prompts don’t. Peec AI’s default mix is 20% branded. Search Console’s branded queries filter (announced November 2025) makes the same split for Google Search queries only, and Google warns that “Some queries might be incorrectly identified as branded or non-branded.”

Origin

Unknown. Both are vendor feature names with no documented coinage.

Common misconceptions

  • “A tool can give you your ranking position in AI answers.” After the study, Rand Fishkin wrote that “any tool that gives a ‘ranking position in AI’ is full of baloney”, and that visibility percentages across many prompts and runs are a reasonable measure.
Evidence
The data on answer variability come from SparkToro (sells audience research software) with Gumshoe.ai (sells AI visibility monitoring).

Sources

  1. Semrush. (n.d.). Prompt Tracking on Semrush. Semrush Knowledge Base. Semrush sells AI visibility software. semrush.com
  2. Peec AI. (n.d.). Setting up your prompts. Peec AI Docs. Peec AI sells AI search tracking software. docs.peec.ai
  3. Google. (n.d.). Performance report (Search results): Dimensions and data groupings. Search Console Help. support.google.com
  4. Google Search Central. (2025, November). Introducing the branded queries filter in Search Console. Google Search Central Blog. developers.google.com
  5. Fishkin, R. (2026, January 28). New research: AIs are highly inconsistent when recommending brands or products. SparkToro, with Gumshoe.ai. SparkToro sells audience research software. sparktoro.com

Cite this entry

Sourcelift. (2026, September 29). Prompt set and prompt tracking. In GEO glossary. https://sourcelift.ai/geo-glossary#prompt-set

Last reviewed September 29, 2026

Position-Adjusted Word Count (and Subjective Impression) #

Measurement Evidence: Peer-reviewed

Position-Adjusted Word Count is a research metric that scores a source by the share of an AI answer’s words in sentences citing it, giving earlier citations more weight.

Read the full entry

Both are impression metrics from the GEO paper. Spanish: recuento de palabras ajustado por posición; impresión subjetiva.

Aggarwal et al. proposed it in the GEO paper to measure a source’s visibility inside a generative answer. The simpler Word Count metric is the words in sentences citing the source ÷ all words in the answer, with a sentence’s words split equally when several sources cite it. The position-adjusted version reduces each sentence’s weight “by an exponentially decaying function of the citation position”, so material cited early counts more. Scores are normalized so the impressions of all citations in an answer sum to 1. Subjective Impression covers what word counts miss, including relevance to the query, influence, uniqueness and the probability of clicking the citation; an LLM, GPT-3.5, scored each facet with a method similar to G-Eval. The paper applied both to GEO-bench, 10,000 queries across 25 domains. They are research metrics; the vendor metrics in this glossary are simpler counts of answers or citations.

Origin

Introduced by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande in “GEO: Generative Engine Optimization” (arXiv, first posted November 16, 2023; accepted to KDD 2024).

Common misconceptions

  • “The GEO paper shows a 40% visibility gain in ChatGPT or Google.” The “up to 40%” was measured with these metrics on answers “generated by the gpt3.5-turbo model” from the top five Google results. On Perplexity.ai the paper reports improvements of up to 37%, and it says efficacy “varies across domains”.
Evidence
Accepted to KDD 2024, per the paper’s arXiv record.

Sources

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2023, November 16). GEO: Generative engine optimization. arXiv:2311.09735 (v3, June 28, 2024; accepted to KDD 2024). arxiv.org
  2. Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., & Zhu, C. (2023, March 29). G-Eval: NLG evaluation using GPT-4 with better human alignment. arXiv:2303.16634. arxiv.org

Cite this entry

Sourcelift. (2026, September 29). Position-Adjusted Word Count. In GEO glossary. https://sourcelift.ai/geo-glossary#position-adjusted-word-count

Last reviewed September 29, 2026

AI referral traffic #

Measurement Evidence: Official documentation

AI referral traffic is website visits that come from links in AI-generated answers, identified by analytics tools from the referring site.

Read the full entry

Google Analytics channel: AI Assistant (Spanish interface: Asistente de IA). Spanish: tráfico de referencia desde IA.

Google Analytics added an AI Assistant channel to its default channel group on May 13, 2026; the Spanish interface calls it Asistente de IA. Google defines it as “the channel by which users arrive at your site from sources like ChatGPT, Gemini, Deepseek, Copilot, or Grok”. Classification depends on the referrer: GA sets the medium to ai-assistant and the campaign to (ai-assistant) when the referrer matches its list of AI assistants. Visits from Google’s AI Overviews and AI Mode are not in this channel; they count as Organic Search. Search Console’s generative AI performance report covers AI Overviews and AI Mode but defines only impressions, so neither of Google’s tools documents a separate count of visits from its own AI features.

Origin

Unknown for the general term. Google Analytics added its AI Assistant channel on May 13, 2026, according to its release notes.

Common misconceptions

  • “Clicks from AI Overviews and AI Mode show up in the AI Assistant channel.” Google says the channel “excludes Google’s AI Overviews and AI Mode”, and that Organic Search covers non-ad links in organic results, “including Google’s AI Overviews and AI Mode.”

Sources

  1. Google. (n.d.). Default channel group. Analytics Help. support.google.com
  2. Google. (2026, May 13). What’s new in Google Analytics [release note on the AI Assistant channel]. Analytics Help. support.google.com
  3. Google. (n.d.). Generative AI performance report (Search). Search Console Help. support.google.com

Cite this entry

Sourcelift. (2026, September 29). AI referral traffic. In GEO glossary. https://sourcelift.ai/geo-glossary#ai-referral-traffic

Last reviewed September 29, 2026

Generative AI performance report (Search Console) #

Measurement Evidence: Official documentation

The generative AI performance report is a Search Console report that shows how many times links to a site appeared in generative AI features on Google Search.

Read the full entry

Spanish interface: informe de rendimiento de IA generativa. Discover has its own generative AI performance report.

Google announced the report on June 3, 2026, and first rolled it out to a subset of site owners in the UK for testing; the Help page says it reached all websites worldwide as of August 31, 2026. It includes impressions for AI Overviews and AI Mode, leaves out Search Labs experiments, and draws on the Web search type of the standard Performance report. Data can be grouped by page, country, date and device. Impressions count “how many times links to your site were shown to a user in a generative AI feature on Google Search”. Google Discover has a separate report. Google announced the report together with the Search generative AI control, an opt-out: sites that opt out “will not receive traffic or impressions from our generative AI features”.

Origin

Google. Announced on June 3, 2026, by Mrinalini Loew, General Manager, Google Search Ecosystem, with a first rollout to a subset of UK site owners; rolled out to all websites worldwide as of August 31, 2026.

Common misconceptions

  • “The report shows how many clicks AI Overviews and AI Mode send you.” The Help page defines only impressions and does not mention clicks, click-through rate or position; Google’s launch post describes “impressions metrics and information about which pages appear in AI responses and in what countries.”
  • “Google Analytics’ AI Assistant channel shows the visits these features send.” That channel “excludes Google’s AI Overviews and AI Mode”; those visits count as Organic Search.

Sources

  1. Google. (n.d.). Generative AI performance report (Search). Search Console Help. support.google.com
  2. Loew, M. (2026, June 3; updated 2026, August 31). New opportunities, control and insights for website owners. Google. blog.google
  3. Google. (n.d.). Search generative AI control. Search Console Help. support.google.com
  4. Google. (n.d.). Default channel group. Analytics Help. support.google.com

Cite this entry

Sourcelift. (2026, September 29). Generative AI performance report. In GEO glossary. https://sourcelift.ai/geo-glossary#generative-ai-performance-report

Last reviewed September 29, 2026

Technical access

8 terms

What decides whether AI crawlers and Google’s AI features can read, use or leave out your pages.

AI crawlers and user agents #

Technical access Evidence: Official documentation

AI crawlers are automated clients that AI companies use to fetch web pages for model training, search indexing or answering a user’s question, each with its own user agent name.

Read the full entry

Also called AI bots. Spanish: rastreadores de IA.

AI companies now run separate bots for separate jobs, and robots.txt rules apply to each name on its own. Training: OpenAI’s GPTBot and Anthropic’s ClaudeBot collect content that may be used to train models. Search: OAI-SearchBot, Claude-SearchBot and PerplexityBot surface pages in the assistants’ search answers, and Applebot feeds search in Spotlight, Siri and Safari. User requests: ChatGPT-User, Claude-User and Perplexity-User may visit a page when someone asks a question. OpenAI says robots.txt rules “may not apply” to these visits and Perplexity says its fetcher “generally ignores robots.txt rules”; Anthropic says disabling Claude-User prevents them. Applebot-Extended “does not crawl webpages”: it is how publishers opt out of training Apple’s models. For AI Overviews and AI Mode, Google points to Googlebot, because AI “is built into Search”.

Origin

Unknown as a general term. Each company names and documents its own bots, and the vendor pages we checked do not say when each bot was introduced.

Common misconceptions

  • “Blocking the training bot removes you from AI answers.” OpenAI says “Each setting is independent of the others”: a site can disallow GPTBot and still allow OAI-SearchBot to appear in ChatGPT search. Apple says pages that disallow Applebot-Extended “can still be included in search results.”

Sources

  1. OpenAI. (n.d.). Overview of OpenAI crawlers. developers.openai.com
  2. Anthropic. (2026, April 7). Does Anthropic crawl data from the web, and how can site owners block the crawler? Claude Help Center. support.claude.com
  3. Perplexity. (n.d.). Perplexity crawlers. docs.perplexity.ai
  4. Apple. (2026, September 4). About Applebot. Apple Support. support.apple.com
  5. Google Search Central. (2025, December 10). AI features and your website. developers.google.com

Cite this entry

Sourcelift. (2026, September 29). AI crawlers and user agents. In GEO glossary. https://sourcelift.ai/geo-glossary#ai-crawlers

Last reviewed September 29, 2026

robots.txt #

Technical access Evidence: Official documentation

robots.txt is a text file at the root of a website whose rules, defined by the Robots Exclusion Protocol (RFC 9309), tell crawlers which URLs they may access.

Read the full entry

Formally, the Robots Exclusion Protocol (RFC 9309). Spanish: archivo robots.txt.

The file must sit at the top level of a site as /robots.txt, in lowercase, and is read in groups: one or more user-agent lines naming crawlers, followed by allow and disallow rules for URL paths. Google says the file is used mainly to avoid overloading a site with requests. It governs crawling, not indexing: a URL blocked in robots.txt can still appear in Google’s results without a description, and can be indexed if other sites link to it. For AI, it is where you address each company’s bots by name. OpenAI, for example, lets a site disallow its training crawler, GPTBot, while allowing OAI-SearchBot, the bot that surfaces sites in ChatGPT search.

Origin

Martijn Koster defined the method in 1994, according to RFC 9309. It became an IETF Proposed Standard, RFC 9309, in September 2022, written by Koster with Gary Illyes, Henner Zeller and Lizzi Sassman of Google.

Common misconceptions

  • “robots.txt keeps a page out of Google.” Google says it “is not a mechanism for keeping a web page out of Google”; to do that, block indexing with noindex or password-protect the page.
  • “Crawlers have to obey robots.txt.” RFC 9309: “These rules are not a form of access authorization.” Compliance depends on each operator: OpenAI says robots.txt rules “may not apply” to user-initiated ChatGPT-User visits, and Perplexity says Perplexity-User “generally ignores robots.txt rules”.

Sources

  1. Koster, M., Illyes, G., Zeller, H., & Sassman, L. (2022, September). Robots Exclusion Protocol (RFC 9309). Internet Engineering Task Force. rfc-editor.org
  2. Google Search Central. (2025, December 10). Introduction to robots.txt. developers.google.com
  3. OpenAI. (n.d.). Overview of OpenAI crawlers. developers.openai.com
  4. Perplexity. (n.d.). Perplexity crawlers. docs.perplexity.ai

Cite this entry

Sourcelift. (2026, September 29). robots.txt. In GEO glossary. https://sourcelift.ai/geo-glossary#robots-txt

Last reviewed September 29, 2026

llms.txt #

Technical access Evidence: Contested or unproven

llms.txt is a proposed Markdown file, placed at a website’s root, that gives large language models a curated summary of the site and links to its most useful pages.

Read the full entry

Spanish: archivo llms.txt.

Jeremy Howard proposed the format in September 2024 so that language models and agents could read a site’s key content without wading through navigation, scripts and ads. The only required element is an H1 with the site’s name; a short summary in a blockquote and H2 sections listing links are optional. Unlike robots.txt, it allows or blocks nothing, and the proposal has no fields for training permissions, usage rights or attribution. Google’s guide says Google Search ignores these files, including for its generative AI features. Google’s Chrome team takes a different angle: Lighthouse checks for the file because, without it, agents “may spend more time crawling the site to understand its high-level structure”. OpenAI’s and Perplexity’s crawler documentation does not say their bots read it.

Origin

Proposed by Jeremy Howard on September 3, 2024, at llmstxt.org. It is a proposal, not a standard adopted by a standards body.

Common misconceptions

  • “llms.txt helps you appear in Google’s AI Overviews.” Google says adding one “will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them.”
  • “llms.txt tells AI companies what they may use for training.” The proposal has no such field. Training opt-outs go through robots.txt rules for each company’s crawler token, such as GPTBot or Google-Extended.
Evidence
A proposal with no documented effect on Google Search.

Sources

  1. Howard, J. (2024, September 3). The /llms.txt file. llmstxt.org. llmstxt.org
  2. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. The llms.txt note was added on June 15, 2026, per the documentation changelog. developers.google.com
  3. Chrome for Developers. (2026, May 5). llms.txt (Lighthouse audit). developer.chrome.com
  4. OpenAI. (n.d.). Overview of OpenAI crawlers. developers.openai.com
  5. Perplexity. (n.d.). Perplexity crawlers. docs.perplexity.ai

Cite this entry

Sourcelift. (2026, September 29). llms.txt. In GEO glossary. https://sourcelift.ai/geo-glossary#llms-txt

Last reviewed September 29, 2026

Server-side rendering (vs client-side JavaScript) #

Technical access Evidence: Industry study

Server-side rendering is the generation of a page’s HTML on the server, so the content arrives in the first response instead of being assembled by JavaScript in the browser.

Read the full entry

Also called SSR. Related: pre-rendering. The alternative is client-side rendering (CSR). Spanish: renderizado en el servidor.

With client-side rendering, the browser builds the page by running JavaScript that modifies the DOM, so a bot that does not execute scripts may get little text. Googlebot processes such pages in three phases (crawling, rendering and indexing) and queues every page that returns HTTP 200 for rendering, unless told not to index it. Google nonetheless says “server-side or pre-rendering is still a great idea”, because it makes sites faster “for users and crawlers, and not all bots can run JavaScript.” For AI assistants the evidence is one industry study from 2024: monitoring Vercel’s network, Vercel and MERJ found that the major AI crawlers they measured, from OpenAI, Anthropic, Meta, ByteDance and Perplexity, did not render JavaScript, while Gemini (through Googlebot’s infrastructure) and Applebot did.

Origin

Unknown. Server-side and client-side rendering are general web-engineering terms. web.dev’s “Rendering on the Web” (Addy Osmani and Jason Miller, February 6, 2019) defines server-side rendering as “Rendering an app on the server to send HTML, rather than JavaScript, to the client.”

Common misconceptions

  • “AI search can’t see JavaScript content.” Not in Google’s case: its AI features draw on Googlebot, and Google’s generative AI guide says “Google is able to process content within JavaScript as long as it isn’t blocked.”
  • “Every AI crawler renders JavaScript like Googlebot.” Vercel and MERJ reported that “none of the major AI crawlers currently render JavaScript”, with Gemini (via Googlebot) and Applebot as the exceptions. The data is from December 2024.
Evidence
Google’s advice is official documentation; the AI-crawler finding rests on one 2024 study by Vercel, which sells frontend cloud hosting, with the consultancy MERJ.

Sources

  1. Google Search Central. (2026, March 4). Understand the JavaScript SEO basics. developers.google.com
  2. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  3. Google Search Central. (2025, December 10). AI features and your website. developers.google.com
  4. Zecchini, G., Moore, A. A., Ubl, M., & Siddle, R. (2024, December 17). The rise of the AI crawler. Vercel, with MERJ. Vercel sells frontend cloud hosting. vercel.com
  5. Osmani, A., & Miller, J. (2019, February 6). Rendering on the Web. web.dev. web.dev

Cite this entry

Sourcelift. (2026, September 29). Server-side rendering. In GEO glossary. https://sourcelift.ai/geo-glossary#server-side-rendering

Last reviewed September 29, 2026

Structured data (schema.org) #

Technical access Evidence: Official documentation

Structured data is code added to a web page in a standardized format, usually with schema.org vocabulary, that describes what the page is about and classifies its content.

Read the full entry

Also called schema markup. Formats: JSON-LD, Microdata and RDFa. Spanish: datos estructurados.

Google recommends JSON-LD and also reads Microdata and RDFa. It says the markup gives it “explicit clues about the meaning of a page” and can make the page eligible for rich results; most markup uses schema.org vocabulary, but Google treats its own documentation as definitive for Search. For its generative AI features, Google’s guide says structured data is not required and no special schema.org markup is needed. Rich result types also change: the FAQ rich result stopped appearing in Google Search on May 7, 2026. In our engine test (September 29, 2026), Gemini, in both languages, still recommended FAQPage and HowTo schema for answer engines.

Origin

Schema.org was announced on June 2, 2011, as “a new initiative from Google, Bing and Yahoo!” to create a common vocabulary for structured data markup (Ramanathan Guha, Google Fellow). Schema.org lists Google, Microsoft, Yahoo and Yandex as its founding companies, and since April 2015 the W3C Schema.org Community Group has been its main forum.

Common misconceptions

  • “AI Overviews need special schema.” Google: “Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add.”
  • “FAQ markup still earns an FAQ rich result.” Google’s changelog: “This feature will no longer appear in Google Search starting May 7, 2026.”

Sources

  1. Google Search Central. (2025, December 10). Introduction to structured data markup in Google Search. developers.google.com
  2. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  3. Google Search Central. (2026, May 8). Latest documentation updates (entry on the FAQ rich result). developers.google.com
  4. Guha, R. (2011, June 2). Introducing schema.org: Search engines come together for a richer web. Inside Search (Google). search.googleblog.com
  5. Schema.org. (n.d.). About Schema.org. schema.org

Cite this entry

Sourcelift. (2026, September 29). Structured data. In GEO glossary. https://sourcelift.ai/geo-glossary#structured-data

Last reviewed September 29, 2026

Snippet controls (nosnippet, data-nosnippet, max-snippet) #

Technical access Evidence: Official documentation

Snippet controls are robots meta rules and an HTML attribute that limit how much page text may be shown as a search snippet or used to generate AI answers.

Read the full entry

Google also calls them preview controls. Spanish: controles de fragmentos (Google’s Spanish pages say controles de vista previa).

Google defines them in its robots meta tag specification. nosnippet removes the text snippet and video preview and also stops the content from being used as a direct input for AI Overviews and AI Mode. max-snippet:[number] caps the snippet at that many characters (0 is equivalent to nosnippet) and also limits how much content can serve as direct input for those features. The data-nosnippet attribute, on span, div or section elements, excludes only the marked text. Google’s AI features page lists these rules, plus noindex, as the way to limit what Search shows from your pages. The trade-off: to be shown as a supporting link in AI Overviews or AI Mode, a page “must be indexed and eligible to be shown in Google Search with a snippet”.

Origin

Google. nosnippet was already an “existing option” when John Mueller announced max-snippet and data-nosnippet as new settings in September 2019. Google’s own pages call the family both “snippet controls” and “preview controls”.

Common misconceptions

  • “Snippet controls only affect the classic blue-link snippet.” Google’s specification says nosnippet “applies to all forms of search results (at Google: web search, Google Images, Discover, AI Overviews, AI Mode)”.
  • “Leaving AI Overviews means losing snippets everywhere.” Since August 31, 2026, the Search generative AI control in Search Console lets a site opt out of AI Overviews, AI Mode and generative AI features in Discover, and Google says it “isn’t used as a ranking or inclusion signal affecting other parts of Search.”

Sources

  1. Google Search Central. (2026, March 24). Robots meta tag, data-nosnippet, and X-Robots-Tag specifications. developers.google.com
  2. Google Search Central. (2025, December 10). AI features and your website. developers.google.com
  3. Mueller, J. (2019, September). More options to help websites preview their content on Google Search. Google Search Central Blog. developers.google.com
  4. Search Console Help. (n.d.). Search generative AI control. Google. support.google.com
  5. Loew, M. (2026, June 3; updated August 31). New opportunities, control and insights for website owners. Google. blog.google

Cite this entry

Sourcelift. (2026, September 29). Snippet controls. In GEO glossary. https://sourcelift.ai/geo-glossary#snippet-controls

Last reviewed September 29, 2026

Google-Extended #

Technical access Evidence: Official documentation

Google-Extended is a robots.txt product token that controls whether content Google crawls from a site may train Gemini models and ground answers in Gemini Apps and Vertex AI.

Read the full entry

A robots.txt product token, not a crawler. Spanish: Google-Extended (not translated).

Disallowing Google-Extended in robots.txt covers two uses of content Google has crawled: training “future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini”, and grounding (supplying content from the Search index to the model at prompt time) in Gemini Apps and in Grounding with Google Search on Vertex AI. Google’s AI features page sends site owners to Google-Extended “to limit AI training and grounding in some of Google’s other systems”, while robots.txt rules for Googlebot govern Search itself. The Search Console Help page for the Search generative AI control adds a use inside Search: “To limit training of the models used to generate responses in Search generative AI features, use Google-Extended.” Read together, Google-Extended governs training and Gemini grounding; appearing in AI Overviews and AI Mode depends on Googlebot, snippet controls and that Search Console setting.

Origin

Google announced it on September 28, 2023 (Danielle Romain, VP, Trust) as a control over whether sites help improve “Bard and Vertex AI generative APIs”. Google’s current documentation describes it in terms of Gemini models.

Common misconceptions

  • “Blocking Google-Extended removes your site from AI Overviews.” Google says it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search”, and AI Overviews and AI Mode are part of Search. To leave them, use the Search generative AI control or snippet controls.
  • “Google-Extended is Google’s AI crawler.” Google says it “doesn’t have a separate HTTP request user agent string”: crawling is done with Google’s existing user agents, and the token is read only as a robots.txt control.
Evidence
Two Google pages describe its scope differently; the entry quotes both.

Sources

  1. Google. (2026, July 14). List of Google’s common crawlers. Google Crawling Infrastructure. developers.google.com
  2. Romain, D. (2023, September 28). An update on web publisher controls. Google. blog.google
  3. Google Search Central. (2025, December 10). AI features and your website. developers.google.com
  4. Search Console Help. (n.d.). Search generative AI control. Google. support.google.com

Cite this entry

Sourcelift. (2026, September 29). Google-Extended. In GEO glossary. https://sourcelift.ai/geo-glossary#google-extended

Last reviewed September 29, 2026

Search generative AI control (Search Console) #

Technical access Evidence: Official documentation

The Search generative AI control is a Search Console setting that lets a site opt out of Google’s AI Overviews, AI Mode and generative AI features in Discover.

Read the full entry

In Search Console: Settings > Search generative AI. Spanish: control de la IA generativa de la Búsqueda.

It offers three options: include the site’s links and content (the default for every property), exclude them, or inherit the value of a parent property. Included content can appear as links and help ground AI responses in those features. Excluded content is neither shown nor linked there, so the site gets no traffic or impressions from them; Google says exclusion generally takes a few days. The control does not cover model training: for that, the Help page points to Google-Extended. Until it arrived, Google’s documented ways to limit these features were robots.txt rules for Googlebot and page-level controls (nosnippet, data-nosnippet, max-snippet, noindex), all of which also change how a page appears in regular results.

Origin

Google announced it on June 3, 2026 (Mrinalini Loew, General Manager, Google Search Ecosystem) as a test with a subset of website owners in the UK, while engaging with regulators such as the UK’s Competition and Markets Authority. The Help page says it was rolled out to all websites worldwide as of August 31, 2026.

Common misconceptions

  • “It does the same job as Google-Extended.” The control governs whether content can appear in Search AI features. For training, Google says: “To limit training of the models used to generate responses in Search generative AI features, use Google-Extended.”
  • “Opting out lowers your regular rankings.” Google: “this control isn’t used as a ranking or inclusion signal affecting other parts of Search.”

Sources

  1. Search Console Help. (n.d.). Search generative AI control. Google. support.google.com
  2. Loew, M. (2026, June 3; updated August 31). New opportunities, control and insights for website owners. Google. blog.google
  3. Google Search Central. (2025, December 10). AI features and your website. developers.google.com
  4. Google Search Central. (2026, March 24). Robots meta tag, data-nosnippet, and X-Robots-Tag specifications. developers.google.com

Cite this entry

Sourcelift. (2026, September 29). Search generative AI control. In GEO glossary. https://sourcelift.ai/geo-glossary#search-generative-ai-control

Last reviewed September 29, 2026

Content and authority

9 terms

What makes a page or a brand worth citing, and how strong the evidence is.

E-E-A-T #

Content and authority Evidence: Official documentation

E-E-A-T is the set of quality criteria in Google’s Search Quality Rater Guidelines (experience, expertise, authoritativeness and trust), of which trust is the most important.

Read the full entry

Stands for experience, expertise, authoritativeness and trustworthiness (the rater guidelines say Trust). Formerly E-A-T. Spanish: experiencia, conocimiento, autoridad y fiabilidad.

E-E-A-T comes from the guidelines for Google’s search quality raters, people who tell Google whether its algorithms “seem to be providing good results”. Google says its ranking systems use several factors to identify content with good E-E-A-T and give it even more weight on “Your Money or Your Life” (YMYL) topics, which could significantly affect people’s health, financial stability or safety, or the welfare of society. The guidelines define Experience as the creator’s “first-hand or life experience for the topic” and call Trust “the most important member of the E-E-A-T family”: an untrustworthy page has low E-E-A-T however expert it seems. Content does not have to show all four. Google’s AI features retrieve pages through the same core ranking systems. In our engine test (September 29, 2026), all ten answers we read said E-E-A-T is not a direct ranking factor, in line with Google.

Origin

Google’s Search Quality Rater Guidelines. Google added “Experience” in December 2022, turning E-A-T into E-E-A-T, in the Search Central blog post “Our latest update to the quality rater guidelines: E-A-T gets an extra E for Experience”. We could not confirm when E-A-T first appeared in the guidelines.

Common misconceptions

  • “E-E-A-T is a ranking factor with its own score.” Google: “While E-E-A-T itself isn’t a specific ranking factor, using a mix of factors that can identify content with good E-E-A-T is useful.”
  • “Quality raters’ ratings push pages up or down.” Google: “Search raters have no control over how pages rank. Rater data is not used directly in our ranking algorithms.”

Sources

  1. Google Search Central. (2025, December 10). Creating helpful, reliable, people-first content. developers.google.com
  2. Google. (2025, September 11). General Guidelines (Search Quality Rater Guidelines). guidelines.raterhub.com
  3. Google Search Central. (2022, December). Our latest update to the quality rater guidelines: E-A-T gets an extra E for Experience. Google Search Central Blog. developers.google.com
  4. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com

Cite this entry

Sourcelift. (2026, September 29). E-E-A-T. In GEO glossary. https://sourcelift.ai/geo-glossary#e-e-a-t

Last reviewed September 29, 2026

Non-commodity content (original research) #

Content and authority Evidence: Official documentation

Non-commodity content is content that offers unique expert or first-hand insight beyond common knowledge, such as original research, reporting or analysis, rather than restating what is widely available.

Read the full entry

Opposite: commodity content. Spanish: contenido no genérico (Google’s Spanish name), as opposed to contenido genérico.

Google uses the term in its guide to optimizing for generative AI features, added on May 15, 2026. Commodity content, the guide says, “is often based on common knowledge, which could originate from anyone, and typically adds little unique insight for readers”, while non-commodity content “provides unique expert or experienced takes that go beyond common knowledge and the ordinary”. Its examples: a generic “7 Tips for First-Time Homebuyers” versus a first-hand account of why a buyer waived an inspection and saved money. The idea predates AI search: Google’s self-assessment asks “Does the content provide original information, reporting, research, or analysis?”, and its original content systems aim to show original reporting “ahead of those who merely cite it”. Google recommends non-commodity content for its own AI features; its effect on citations in other AI engines is not documented.

Origin

Google’s generative AI optimization guide, added to Search Central on May 15, 2026, is the earliest Google use we found. Who first used “commodity content” in this sense is unknown.

Common misconceptions

  • “First-party data is another name for original research.” In Google’s usage, “First-party data is information customers have consented to provide, like an email address or phone number that your business directly collects and owns.” Original research is what Google’s self-assessment calls “original information, reporting, research, or analysis”.

Sources

  1. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. Added on May 15, 2026, per the documentation changelog. developers.google.com
  2. Google Search Central. (2025, December 10). Creating helpful, reliable, people-first content. developers.google.com
  3. Google Search Central. (2025, December 10). A guide to Google Search ranking systems. developers.google.com
  4. Google Ads. (n.d.). Power your ad privacy strategy with first-party data. business.google.com

Cite this entry

Sourcelift. (2026, September 29). Non-commodity content. In GEO glossary. https://sourcelift.ai/geo-glossary#non-commodity-content

Last reviewed September 29, 2026

Entity and Knowledge Graph #

Content and authority Evidence: Official documentation

An entity is a distinct real-world thing, such as a person, place or organization; Google’s Knowledge Graph is its collection of facts about entities and their relationships.

Read the full entry

Related: knowledge panel. Spanish: entidad; Google’s Spanish documentation calls the Knowledge Graph “Gráfico de conocimiento” and knowledge panels “paneles de información”.

Google introduced the Knowledge Graph on May 16, 2012, to move search from matching keywords to understanding “things, not strings”. Its launch example was the query Taj Mahal, which could mean the monument or the musician: the graph lets Google tell them apart. At launch it held “more than 500 million objects, as well as more than 3.5 billion facts about and relationships between these different objects”. Knowledge panels are the information boxes Google shows when someone searches for an entity in the graph. Google’s public Knowledge Graph Search API returns entries from the graph, but Google says it “is not suitable for use as a production-critical service”. Google’s guide to its generative AI features does not mention entities or the Knowledge Graph, so their role in AI Overviews is not documented there.

Origin

Google, in a blog post by Amit Singhal (SVP, Engineering) on May 16, 2012. The idea of entities is older than Google’s use of it; who first applied it to search is unknown.

Common misconceptions

  • “You can create or order a knowledge panel.” Google: “Knowledge panels are automatically generated, and information that appears in a knowledge panel comes from various sources across the web.” What the subject or an official representative can do is claim an existing panel and suggest changes.

Sources

  1. Singhal, A. (2012, May 16). Introducing the Knowledge Graph: things, not strings. Google. blog.google
  2. Google for Developers. (2024, April 26). Google Knowledge Graph Search API. developers.google.com
  3. Google. (n.d.). About knowledge panels. Knowledge Panel Help. support.google.com
  4. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com

Cite this entry

Sourcelift. (2026, September 29). Entity and Knowledge Graph. In GEO glossary. https://sourcelift.ai/geo-glossary#entity

Last reviewed September 29, 2026

Entity consistency (NAP) #

Content and authority Evidence: Practitioner consensus

Entity consistency is the practice of presenting a business’s name, address, phone number and other core facts identically across its website, business profiles, directories and other sources.

Read the full entry

NAP stands for name, address and phone number. Also called NAP consistency. Spanish: coherencia de entidad, coherencia NAP.

The idea comes from local SEO. Google’s Business Profile guidelines say a profile’s name should “reflect your business’s real-world name, as used consistently on your storefront, website, stationery, and as known to customers”, and Google says businesses “with complete and accurate info are more likely to show up in local search results”. Its generative AI guide adds that Business Profiles “can help your products and services to be visible in both AI responses and other Google Search results”. Beyond Google’s own profiles, the case rests on practitioner reasoning: AI engines take business facts from third-party sources (in a small BrightLocal test, “All LLMs are using directories and citations for business information across every industry”), so mismatched names, addresses or phone numbers hand them conflicting facts. We found no engine documentation or controlled study that measures this effect on AI answers.

Origin

Unknown. NAP is a local SEO term with no documented coinage that we found; “entity consistency” is a newer practitioner label, also of unknown origin.

Common misconceptions

  • “Adding a service keyword or a location to the business name helps.” Google’s guidelines list service or product information and location information among the things a business name must not include, and warn: “Including unnecessary information in your business name isn’t permitted, and could result in the suspension of your Business Profile.”
Evidence
Google documents name rules for Business Profiles; the effect of consistency across other sites on AI answers is undocumented.

Sources

  1. Google. (n.d.). Guidelines for representing your business on Google. Google Business Profile Help. support.google.com
  2. Google. (n.d.). Tips to improve your local ranking on Google. Google Business Profile Help. support.google.com
  3. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  4. Herbert-Smith, K. (2025, July 22). AI search makes local listings more important than ever. BrightLocal. BrightLocal sells local SEO software, including listings management and citation building. brightlocal.com

Cite this entry

Sourcelift. (2026, September 29). Entity consistency (NAP). In GEO glossary. https://sourcelift.ai/geo-glossary#entity-consistency

Last reviewed September 29, 2026

Brand mentions #

Content and authority Evidence: Industry study

Brand mentions are references to a brand’s name on web pages, with or without a link, as distinct from mentions of the brand inside AI-generated answers.

Read the full entry

Includes linked and unlinked mentions; Ahrefs calls them branded web mentions. Not the same as mentions inside AI answers (see mention rate). Spanish: menciones de marca.

GEO looks at web mentions because AI features draw on what other sites say: Google’s guide says its AI features “can show what’s being said about products and services across the web, including in blogs, videos, and forum discussions”. An industry study by Ahrefs, which sells SEO and AI visibility software, looked at 75,000 brands. “Branded web mentions” (the brand’s name anywhere on the web) had the strongest correlation with how often Google’s AI Overviews mentioned a brand (Spearman 0.664), ahead of branded anchor text (0.527) and the number of backlinks (0.218). The study covers AI Overviews only and shows correlation, not cause (“correlation ≠ causation”, Ahrefs notes). It counts mentions with or without a link, so it cannot tell linked from unlinked mentions apart.

Origin

Unknown. The term predates GEO; we found no primary source that coined it.

Common misconceptions

  • “More mentions will get you named in AI answers, so plant as many as you can.” The Ahrefs data is correlational, and Google’s guide warns: “Seeking inauthentic ‘mentions’ across the web isn’t as helpful as it might seem. Our core ranking systems focus on high-quality content while other systems block spam; our generative AI features depend on both.”
Evidence
Industry study by Ahrefs, which sells SEO and AI visibility software; correlation only, and Google AI Overviews only.

Sources

  1. Linehan, L., & Guan, X. (2025, May 26; updated 2026, April 27). An analysis of AI Overview brand visibility factors (75K brands studied). Ahrefs. Ahrefs sells SEO and AI visibility software, including Brand Radar. ahrefs.com
  2. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com

Cite this entry

Sourcelift. (2026, September 29). Brand mentions. In GEO glossary. https://sourcelift.ai/geo-glossary#brand-mentions

Last reviewed September 29, 2026

Third-party sources in AI answers (reviews, directories, forums) #

Content and authority Evidence: Industry study

Third-party sources in AI answers are cited pages that the brand being discussed does not control, such as review platforms, directories and forums.

Read the full entry

Includes review platforms, business directories and forums such as Reddit. Spanish: fuentes de terceros en las respuestas de IA.

Vendor studies show AI answers citing these sites, but each measures a different slice at a different time. Local directories: BrightLocal, which sells listings and citation-building software, ran 20 local searches across 10 niches on ChatGPT search, Gemini, Perplexity and Google’s AI Mode, and found Yelp used as a source in 33% of its searches (published July 22, 2025; search dates not given). Software review sites: SE Ranking, which sells SEO software including an AI Overviews tracker, checked 30,000 US commercial keywords on December 1, 2025; 34.5% of the AI Overview responses cited at least one of 23 review platforms, such as G2 and Capterra. Forums: Semrush, which sells SEO and AI visibility software, found ChatGPT cited Reddit in close to 60% of responses in early August 2025 and around 10% by mid-September, in a study covering July 14 to October 12, 2025.

Origin

Unknown. The grouping is descriptive; the studies above date from 2025 and 2026.

Common misconceptions

  • “AI answers draw mainly on your own website.” BrightLocal found that “All LLMs are using directories and citations for business information across every industry”, and Google says its AI features “can show what’s being said about products and services across the web, including in blogs, videos, and forum discussions”.
  • “Reddit is the top source in AI answers.” In Semrush’s data, Reddit’s share of ChatGPT responses fell from close to 60% to around 10% in about six weeks of 2025. A share means little without its engine and measurement window.
Evidence
Studies by BrightLocal, SE Ranking and Semrush, which sell local SEO, SEO and AI visibility software; samples and dates differ.

Sources

  1. Herbert-Smith, K. (2025, July 22). AI search makes local listings more important than ever. BrightLocal. BrightLocal sells local SEO software, including listings management and citation building. brightlocal.com
  2. Deda, Y. (2026, January 29). Despite 90% traffic loss, review platforms top AI Overview citations. SE Ranking. SE Ranking sells SEO software, including an AI Overviews tracker. seranking.com
  3. Harsel, L. (2025, November 10). The most-cited domains in AI: A 3-month study. Semrush. Semrush sells SEO and AI visibility software. semrush.com
  4. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com

Cite this entry

Sourcelift. (2026, September 29). Third-party sources in AI answers. In GEO glossary. https://sourcelift.ai/geo-glossary#third-party-sources

Last reviewed September 29, 2026

Citable content (statistics, quotations, sources) #

Content and authority Evidence: Contested or unproven

Citable content is content that includes elements an AI answer can reuse and attribute, such as specific statistics, quotations from credible sources and references to sources.

Read the full entry

Covers three methods from the GEO paper: Cite Sources, Quotation Addition and Statistics Addition. Spanish: contenido citable.

The idea comes from the paper that introduced GEO (Aggarwal et al., first posted November 16, 2023, accepted to KDD 2024), which tested nine ways of rewriting source pages. Cite Sources adds citations from credible sources, Quotation Addition adds quotations from them, and Statistics Addition replaces qualitative statements with quantitative data where possible. In the paper’s main results table (arXiv version 3), these raised the Position-Adjusted Word Count score from 19.3 for unmodified pages to 27.2 (quotations), 25.2 (statistics) and 24.6 (sources); Fluency Optimization scored 24.7 and Keyword Stuffing 17.7. The answers were generated by gpt-3.5-turbo from the top five Google results, and a test on Perplexity.ai showed gains of up to 37%. C-SEO Bench (NeurIPS 2025) retested these three methods with two 2024 models across six domains and found no significant gain from any of them.

Origin

Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande defined the three methods in the GEO paper (arXiv, November 16, 2023). Who first used the label “citable content” is unknown.

Common misconceptions

  • “Adding statistics raises AI visibility by 40%.” The paper’s “up to 40%” is a best case from a 2023 benchmark built on gpt-3.5-turbo; its largest gain came from quotations, not statistics, and the authors found that “the efficacy of these strategies varies across domains”. It did not test today’s ChatGPT or Google AI Overviews.
  • “You need to rewrite your content in a special style for AI.” Google: “You don’t need to write in a specific way just for generative AI search.”
Evidence
Peer-reviewed studies disagree: the 2023 GEO paper found gains, and a 2025 benchmark found no significant gain from these methods.

Sources

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2023, November 16). GEO: Generative engine optimization. arXiv:2311.09735 (v3, June 28, 2024; accepted to KDD 2024). arxiv.org
  2. Puerto, H., Gubri, M., Green, T., Oh, S. J., & Yun, S. (2025, June 6). C-SEO Bench: Does conversational SEO work? arXiv:2506.11097 (v3, October 20, 2025; NeurIPS 2025 Datasets and Benchmarks Track). arxiv.org
  3. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com

Cite this entry

Sourcelift. (2026, September 29). Citable content. In GEO glossary. https://sourcelift.ai/geo-glossary#citable-content

Last reviewed September 29, 2026

Freshness #

Content and authority Evidence: Official documentation

Freshness is how recent and up to date a page’s content is, which search systems weigh more heavily for queries where users expect recent information.

Read the full entry

Google’s name for its freshness systems: “query deserves freshness” systems. Spanish: actualidad del contenido; Google’s Spanish documentation says “sistemas de actualización”.

Google’s ranking systems guide says: “We have various ‘query deserves freshness’ systems designed to show fresher content for queries where it would be expected.” Its example: someone searching for a movie that has just been released probably wants recent reviews rather than articles from when production began. Freshness also shapes AI answers. Google describes retrieval-augmented generation, which it also calls grounding, as a technique to improve “the quality, accuracy, and freshness of AI responses” by retrieving “up-to-date web pages” from its Search index.

Origin

Google announced a ranking change for fresher results on November 3, 2011 (Amit Singhal, Google Fellow), building on its Caffeine indexing system. It said the change affected “roughly 35 percent of searches”, then clarified that this meant at least one result on the page, with 6 to 10% of searches noticeably affected. Who coined “query deserves freshness” is unverified.

Common misconceptions

  • “Changing the date makes a page fresh.” Google lists this among the warning signs of search engine-first content: “Are you changing the date of pages to make them seem fresh when the content has not substantially changed?”
  • “Newer content always ranks higher.” The systems target “queries where it would be expected”. Google’s example: a search for “earthquake” normally brings back preparation resources, and news articles may appear after a recent earthquake.

Sources

  1. Google Search Central. (2025, December 10). A guide to Google Search ranking systems. developers.google.com
  2. Google Search Central. (2025, December 10). Creating helpful, reliable, people-first content. developers.google.com
  3. Singhal, A. (2011, November 3). Giving you fresher, more recent search results. Official Google Blog. googleblog.blogspot.com
  4. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  5. Google Search Central. (2025, December 10). Google Search’s core updates and your website. developers.google.com

Cite this entry

Sourcelift. (2026, September 29). Freshness. In GEO glossary. https://sourcelift.ai/geo-glossary#freshness

Last reviewed September 29, 2026

Risks and policy

8 terms

Google’s spam policies and updates, which now also cover its AI answers, and the ways AI answers go wrong.

Core update vs spam update #

Risks and policy Evidence: Official documentation

A core update is a broad change Google makes to its search systems several times a year; a spam update is a notable improvement to its automated spam-detection systems.

Read the full entry

Google calls a spam update that deals specifically with link spam a link spam update. Spanish: actualización principal frente a actualización de spam.

Core updates are “significant, broad changes” to Google’s search algorithms and systems that “don’t target specific sites or individual web pages”, so a drop during one is not a sanction aimed at that site; Google points owners to its content self-assessment. Spam updates improve automated systems that run all the time. A site that fixes violations can improve as those systems learn “over a period of months” that it complies, but after a link spam update any ranking benefit from spammy links is lost and “cannot be regained”. Google announces both on its Search Status Dashboard. As of September 29, 2026, it listed 2026 core updates starting March 27 and May 21, and spam updates starting March 24, June 24, August 18 and September 24; the September update, which Google said applies globally and to all languages, was still rolling out.

Origin

Both are Google’s own labels. The oldest entry on its Search Status Dashboard is the November 2021 spam update, which started on November 3, 2021. When Google first used either term is unverified.

Common misconceptions

  • “A drop after a core update means Google penalized the site.” Google says these changes “are broad in nature, and don’t target specific sites or individual web pages”.
  • “Recovery has to wait for the next core update.” Google: “you don’t necessarily have to wait for a major core update to see the effect of your improvements”, although it “could take several months” for its systems to confirm that a site as a whole has improved.

Sources

  1. Google Search Central. (2025, December 10). Google Search’s core updates and your website. developers.google.com
  2. Google Search Central. (2025, December 10). Google Search spam updates and your site. developers.google.com
  3. Google. (n.d.). History for Ranking. Google Search Status Dashboard. Checked on September 29, 2026. status.search.google.com
  4. Google. (2026, September 24). September 2026 spam update. Google Search Status Dashboard. status.search.google.com

Cite this entry

Sourcelift. (2026, September 29). Core update vs spam update. In GEO glossary. https://sourcelift.ai/geo-glossary#core-update

Last reviewed September 29, 2026

Manual action #

Risks and policy Evidence: Official documentation

A manual action is a sanction a human reviewer at Google applies when a site’s pages break Google’s spam policies; some or all of the site stops appearing in results.

Read the full entry

Informally called a Google penalty. Spanish: acción manual.

Manual actions are the exception. Google says its algorithms are “extremely good at detecting spam” and in most cases find and remove it automatically; people step in to protect the quality of its index. When a reviewer issues one, Google notifies the owner in Search Console’s Manual actions report and message center, and the report names the issue, such as “Unnatural links to your site” or “Cloaking and/or sneaky redirects”. Once every listed issue is fixed on every page, the owner can select Request Review; Google says most reconsideration reviews “can take several days or weeks”, and link-related ones may take longer. If the report is empty, a ranking drop has another cause, such as automated spam systems or a core update. Google’s documentation now says the impact of site reputation manual actions doesn’t apply to results shown to users in the European Economic Area.

Origin

Google’s term, used in Search Console’s Manual actions report. When Google first used it is unverified.

Common misconceptions

  • “Every traffic drop is a Google penalty.” Google says its algorithms catch most spam on their own (“in most cases we automatically discover it and remove it from our search results”); a manual action is a human decision that Google reports “in the Manual actions report and in the Search Console message center”.

Sources

  1. Google. (n.d.). Manual actions report. Search Console Help. support.google.com
  2. Google Search Central. (2026, August 28). Spam policies for Google web search. developers.google.com
  3. Google Search Central. (2025, December 10). Google Search spam updates and your site. developers.google.com
  4. Google Search Central. (2025, December 10). Google Search’s core updates and your website. developers.google.com

Cite this entry

Sourcelift. (2026, September 29). Manual action. In GEO glossary. https://sourcelift.ai/geo-glossary#manual-action

Last reviewed September 29, 2026

Scaled content abuse #

Risks and policy Evidence: Official documentation

Scaled content abuse is Google’s term for generating many pages primarily to manipulate search rankings rather than help users, whether AI tools, people or both produce them.

Read the full entry

Spanish: abuso de contenido a gran escala.

Google’s spam policies list examples: using generative AI or similar tools to generate many pages without adding value for users; scraping feeds, search results or other content, including through automated transformations such as synonymizing or translating, where little value is provided; stitching together content from different pages without adding value; and creating multiple sites to hide the scaled nature of the content. Google’s guide to AI search extends the policy to AI answers: separate content for every variation of a query, fan-out queries included, made primarily to manipulate rankings or generative AI responses violates it. A related policy, doorway abuse, covers sites or pages created to rank for specific, similar queries. In our engine test (September 29, 2026), all ten answers we read, in English and Spanish, described the policy correctly.

Origin

Google, March 5, 2024, in a post by Elizabeth Tucker, Director of Product Management. Google had long had a policy against using automation to generate low-quality or unoriginal content at scale; the new policy targets producing content at scale to boost rankings “whether automation, humans or a combination are involved”.

Common misconceptions

  • “Any AI-generated content is spam.” Google’s guidance says using generative AI tools “to generate many pages without adding value for users may violate” the policy: the issue is volume without value, not the tool.
  • “Content written by people can’t break this policy.” The policy targets large amounts of unoriginal, low-value content “no matter how it’s created”.

Sources

  1. Google Search Central. (2026, August 28). Spam policies for Google web search. developers.google.com
  2. Tucker, E. (2024, March 5). New ways we’re tackling spammy, low-quality content on Search. Google. blog.google
  3. Google Search Central. (2026, July 10). Optimizing your website for generative AI features on Google Search. developers.google.com
  4. Google Search Central. (2025, December 10). Google Search’s guidance on using generative AI content on your website. developers.google.com

Cite this entry

Sourcelift. (2026, September 29). Scaled content abuse. In GEO glossary. https://sourcelift.ai/geo-glossary#scaled-content-abuse

Last reviewed September 29, 2026

Site reputation policy (formerly site reputation abuse) #

Risks and policy Evidence: Official documentation

The site reputation policy is Google’s spam policy against third-party content published on a host site mainly to exploit the ranking signals the host earned with its own content.

Read the full entry

Site reputation abuse was Google’s name for the policy when it was announced in 2024, and the name the European Commission used in 2025. Spanish: política de reputación de sitios (antes, abuso de reputación del sitio).

Third-party content comes from anyone separate from the host site, such as its users, freelancers or white-label services; the policy applies only when it is published mainly to benefit from the host’s ranking signals. In an update published on August 28, 2026, Google changed enforcement in the European Economic Area (EEA). In results shown to EEA users, affected pages may be categorized as separate from the main domain, so each part of the site ranks on its own merits, but manual actions no longer apply; Google said it would lift previous manual actions under the policy for those results, and it offers EEA sites a new reconsideration process and alternative dispute resolution. Outside the EEA, manual actions still apply. The European Commission opened a Digital Markets Act investigation into the policy on November 13, 2025.

Origin

Google announced it as site reputation abuse on March 5, 2024 (Elizabeth Tucker), with enforcement set to begin on May 5, 2024. Google’s spam policies now head the section “Site reputation policy”; when the name changed is unverified.

Common misconceptions

  • “Any third-party or sponsored content on a site breaks the policy.” Google: “Having third-party content alone isn’t inconsistent with the site reputation policy”; it becomes a problem only when the content is published mainly because of the host’s established ranking signals.

Sources

  1. Google Search Central. (2026, August 28). Spam policies for Google web search. developers.google.com
  2. Google Search Central. (2026, August 28). Update to the Site Reputation Policy. Google Search Central Blog. developers.google.com
  3. Tucker, E. (2024, March 5). New ways we’re tackling spammy, low-quality content on Search. Google. blog.google
  4. European Commission. (2025, November 13). Commission opens investigation into potential Digital Markets Act breach by Google in demoting media publishers’ content in search results. digital-markets-act.ec.europa.eu
  5. Google. (n.d.). Manual actions report. Search Console Help. support.google.com

Cite this entry

Sourcelift. (2026, September 29). Site reputation policy. In GEO glossary. https://sourcelift.ai/geo-glossary#site-reputation-policy

Last reviewed September 29, 2026

Cloaking #

Risks and policy Evidence: Official documentation

Cloaking is serving one version of a page to search engine crawlers and a different one to people, with the intent to manipulate rankings and mislead users.

Read the full entry

Google’s matching manual action is “Cloaking and/or sneaky redirects”. Spanish: encubrimiento.

Google’s spam policies give the example of inserting text or keywords into a page only when the user agent requesting it is a search engine, not a human visitor. Since May 15, 2026, Google’s documentation says the spam policies also apply to generative AI responses in Google Search. The same technique can target other AI systems, whose crawlers and fetchers identify themselves by user agent, such as OpenAI’s OAI-SearchBot and ChatGPT-User. In controlled experiments published in October 2025 by SPLX, which sells AI security software, a test site served AI crawlers an altered page, and AI tools including OpenAI’s Atlas browser repeated the altered version.

Origin

Unknown. Gyöngyi and Garcia-Molina’s “Web Spam Taxonomy” (AIRWeb workshop, 2005) already described spam servers that “return one specific HTML document to a regular web browser, while they return a different document to a web crawler”.

Common misconceptions

  • “Paywalls are cloaking.” Google: “we don’t consider this to be cloaking if Google can see the full content of what’s behind the paywall”, provided the publisher also follows its flexible sampling guidance.
Evidence
Google’s definition is official; the AI-crawler test is a vendor experiment.

Sources

  1. Google Search Central. (2026, August 28). Spam policies for Google web search. The clarification that the policies apply to generative AI responses was added on May 15, 2026, per the documentation changelog. developers.google.com
  2. Gyöngyi, Z., & Garcia-Molina, H. (2005). Web spam taxonomy. First International Workshop on Adversarial Information Retrieval on the Web (AIRWeb ‘05). airweb.cse.lehigh.edu
  3. OpenAI. (n.d.). Overview of OpenAI crawlers. developers.openai.com
  4. Vlahov, I., & Eymery, B. (2025, October 25). OpenAI’s new browser Atlas falls for AI-targeted cloaking attack. SPLX. SPLX, now part of Zscaler, sells AI security software. splx.ai
  5. Google. (n.d.). Manual actions report. Search Console Help. support.google.com

Cite this entry

Sourcelift. (2026, September 29). Cloaking. In GEO glossary. https://sourcelift.ai/geo-glossary#cloaking

Last reviewed September 29, 2026

AI answer manipulation (prompt injection, hidden text) #

Risks and policy Evidence: Official documentation

AI answer manipulation is the use of deceptive techniques, such as hidden text or instructions planted in web pages, to make AI answers feature, recommend or cite certain content.

Read the full entry

Includes indirect prompt injection. Spanish: manipulación de respuestas de IA (inyección de prompts, texto oculto).

Google’s spam policies define spam as techniques used to deceive users or manipulate its systems into featuring content prominently, including “attempting to manipulate generative AI responses in Google Search”. Its hidden text and link abuse policy covers content placed “solely to manipulate search engines and not to be easily viewable by human visitors”, such as white text on a white background. Google’s policies do not use the term prompt injection; researchers do. Greshake et al. (2023) showed how to compromise applications built on language models, including Bing’s GPT-4 powered chat, “by strategically injecting prompts into data likely to be retrieved”. Kumar and Lakkaraju (2024; an arXiv preprint, not peer-reviewed as far as we know) showed, with a catalog of fictitious coffee machines, that adding a “strategic text sequence” to a product’s information page “can significantly increase its likelihood of being listed as the LLM’s top recommendation”.

Origin

Simon Willison proposed the name “prompt injection” on September 12, 2022, a day after Riley Goodside showed malicious inputs ordering GPT-3 to ignore its previous directions. Greshake et al. introduced indirect prompt injection in February 2023; the paper won the AISec 2023 Best Paper Award. Google clarified on May 15, 2026, that its spam policies apply to generative AI responses in Google Search.

Common misconceptions

  • “Google’s spam policies cover rankings, not AI Overviews or AI Mode.” Google’s documentation changelog, May 15, 2026: “Clarified that our spam policies also apply to generative AI responses in Google Search.”
Evidence
The policy is Google’s; the attack research includes a peer-reviewed paper and an arXiv preprint.

Sources

  1. Google Search Central. (2026, August 28). Spam policies for Google web search. The clarification that the policies apply to generative AI responses was added on May 15, 2026, per the documentation changelog. developers.google.com
  2. Willison, S. (2022, September 12). Prompt injection attacks against GPT-3. simonwillison.net. simonwillison.net
  3. Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. 16th ACM Workshop on Artificial Intelligence and Security (AISec 2023), Best Paper Award. arXiv:2302.12173. arxiv.org
  4. Kumar, A., & Lakkaraju, H. (2024, April 11). Manipulating large language models to increase product visibility. arXiv:2404.07981 (preprint). arxiv.org

Cite this entry

Sourcelift. (2026, September 29). AI answer manipulation. In GEO glossary. https://sourcelift.ai/geo-glossary#ai-answer-manipulation

Last reviewed September 29, 2026

Misattribution #

Risks and policy Evidence: Industry study

Misattribution is an error in which an AI answer credits information to the wrong source, such as another publisher, a copy of the original or a fabricated or broken link.

Read the full entry

Also discussed as citation accuracy. Spanish: atribución errónea.

The Tow Center for Digital Journalism measured it in a study published in Columbia Journalism Review on March 6, 2025. Researchers took excerpts from 10 articles by each of 20 news publishers and asked eight AI search tools, including ChatGPT search, Gemini and Perplexity, to identify each article’s headline, original publisher, publication date and URL: 1,600 queries in all. Collectively, the tools gave incorrect answers to more than 60% of queries. More than half of the responses from Gemini and Grok 3 cited fabricated or broken URLs that led to error pages. On some occasions, the chatbots linked to syndicated copies on other platforms instead of the original, often even when the publisher was known to have a licensing deal with the AI company. The tool versions tested date from early 2025, so current results may differ.

Origin

Unknown as an AI-search term. Liu, Zhang and Liang (2023) measured citation accuracy in generative search engines, and the Tow Center (2025) measured it for news.

Common misconceptions

  • “Paid AI search tiers cite more reliably.” In the Tow Center study, the premium versions tested (Perplexity Pro and Grok 3) answered more prompts correctly than their free equivalents but “paradoxically also demonstrated higher error rates”, mainly by giving “definitive, but wrong, answers”.
  • “A citation proves the sentence next to it.” Auditing four generative search engines, Liu, Zhang and Liang (2023) found that “only 74.5% of citations support their associated sentence”.
Evidence
Study by the Tow Center for Digital Journalism at Columbia’s Graduate School of Journalism, an academic research center, not a vendor; it was not peer-reviewed.

Sources

  1. Jaźwińska, K., & Chandrasekar, A. (2025, March 6). AI search has a citation problem. Columbia Journalism Review, Tow Center for Digital Journalism. cjr.org
  2. Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating verifiability in generative search engines. Findings of EMNLP 2023. arXiv:2304.09848. arxiv.org

Cite this entry

Sourcelift. (2026, September 29). Misattribution. In GEO glossary. https://sourcelift.ai/geo-glossary#misattribution

Last reviewed September 29, 2026

What AI engines say about these terms: our test

On September 29, 2026, we asked ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews to define eight terms from this glossary, in English and in Spanish. None of the 80 answers got the definition wrong. The problems were in the details: invented numbers, outdated facts, wrong product names, advice Google contradicts and few links to primary sources.

Results in English: 40 of 40 answers

Question ChatGPTClaudeGeminiPerplexityGoogle AI Overviews
T1. What is GEO in digital marketing? Correct Correct a Correct Correct b Correct
T2. What is AEO (answer engine optimization)? Correct Correct Correct c Correct b Correct d
T3. What is llms.txt, and does Google Search use it? Correct Correct Correct Correct Correct
T4. What is query fan-out in Google’s AI search? Correct Correct Correct Correct b Correct e
T5. What are Google AI Overviews? Correct Correct Correct Correct b Correct f
T6. What is retrieval-augmented generation (RAG)? Correct Correct Correct Correct Correct
T7. What is scaled content abuse according to Google? Correct Correct Correct Correct Correct
T8. What is E-E-A-T, and is it a Google ranking factor? Correct Correct Correct Correct Correct

Engine test in English, September 29, 2026. “Correct” means the definition matched the primary source; letters point to the notes below the table.

Notes

  1. Credited the 2023 paper that introduced GEO, the only engine to do so.
  2. Answered in Spanish, the account’s interface language.
  3. Recommended FAQPage and HowTo markup and answer blocks of 40 to 60 words. Google says no special markup is needed, and it stopped showing FAQ rich results on May 7, 2026.
  4. Cited a Reddit thread.
  5. Said fan-out typically runs 5 to 12 searches, without a source. Google publishes no number.
  6. Said AI Overviews are available in more than 120 countries. Google announced more than 200 countries and territories on May 20, 2025.

Results in Spanish: 40 of 40 answers

Question ChatGPTClaudeGeminiPerplexityGoogle AI Overviews
T1. ¿Qué es el GEO en marketing digital? Correct Correct a Correct b Correct c Correct
T2. ¿Qué es el AEO (answer engine optimization)? Correct d Correct e Correct f Correct Correct g
T3. ¿Qué es llms.txt y lo usa Google Search? Correct h Correct i Correct Correct j Correct k
T4. ¿Qué es el query fan-out en la búsqueda con IA de Google? Correct Correct Correct l Correct m Correct
T5. ¿Qué son las AI Overviews de Google? Correct n Correct o Correct Correct Correct p
T6. ¿Qué es la generación aumentada por recuperación (RAG)? Correct Correct Correct q Correct Correct
T7. ¿Qué es el abuso de contenido a gran escala según Google? Correct Correct r Correct Correct Correct
T8. ¿Qué es E-E-A-T y es un factor de posicionamiento de Google? Correct Correct Correct Correct Correct

Engine test in Spanish, September 29, 2026. “Correct” means the definition matched the primary source; letters point to the notes below the table.

Notes

  1. Credited the 2023 paper that introduced GEO, the only engine to do so.
  2. Said, without a source, that adding statistics and expert quotes raises the chances of being cited. The evidence is contested (see Citable content).
  3. Repeated, from an SEO blog, an unsourced claim that close to 50% of searches already happen inside AI answers.
  4. Read in part: the temporary chat was lost on reload after about 3,000 of its 3,600 characters.
  5. Recommended FAQ and HowTo markup and noted Google’s 2023 cut to FAQ results, but not their removal on May 7, 2026.
  6. Recommended FAQPage and HowTo markup and answer blocks of 40 to 60 words. Google says no special markup is needed, and it stopped showing FAQ rich results on May 7, 2026.
  7. Cited a Facebook post.
  8. Said there is no official evidence that Google Search uses the file. Google’s guide states that it does not use it.
  9. Quoted an unverified vendor statistic about AI bot visits to llms.txt files.
  10. One of its sources says llms.txt tells models what content they may use; the proposal has no such field.
  11. Cited an Instagram post, did not name the proposal’s author and did not link Google’s own guide.
  12. Said fan-out generates between 5 and more than 20 searches, without a source.
  13. Used its own Spanish name, “abanico de consultas”; Google’s Spanish documentation says “ramificación de búsquedas”.
  14. Called the feature “Visión general creada por IA”; Google’s Spanish name is “vistas creadas con IA”.
  15. Called the feature “Resúmenes creados con IA”; Google’s Spanish name is “vistas creadas con IA”.
  16. Used the English name, AI Overviews, and never gave Google’s Spanish name, “vistas creadas con IA”, although the answer box itself was labeled “Vista creada con IA”.
  17. Overstated RAG: said it forces the model to answer from facts.
  18. Linked only vendor blogs; asked the same question in English, Claude linked Google’s spam policies.

What we found

  1. No wrong definitions. All 80 answers defined the term correctly.
  2. Invented precision. Google’s AI Overview, in English, said query fan-out typically runs 5 to 12 searches; Gemini, in Spanish, said between 5 and more than 20. Google publishes no number.
  3. Outdated facts. Google’s own AI Overview said AI Overviews are available in more than 120 countries. Google announced more than 200 countries and territories on May 20, 2025.
  4. Wrong names. In Spanish, ChatGPT translated AI Overviews as “Visión general creada por IA” and Claude as “Resúmenes creados con IA”, and Google’s own AI Overview used the English name. Google’s Spanish name is “vistas creadas con IA”.
  5. Advice Google contradicts. Gemini, in both languages, recommended FAQPage and HowTo markup and 40 to 60 word answer blocks for AEO. Google says its AI features need no special markup, and it stopped showing FAQ rich results on May 7, 2026.
  6. Origins go missing. Only Claude credited the 2023 paper that introduced GEO, and none of the 10 answers about RAG credited Lewis et al. (2020), the paper that named the technique.
  7. Few primary sources. Leaving out Google’s own answers, 14 of 32 English answers and 7 of 32 Spanish answers linked a primary source. Google’s AI answers cited a Reddit thread, a Facebook post and an Instagram post, and its Spanish answer about llms.txt did not link Google’s own guide.
  8. Logged-in answers are personalized. 12 of 16 ChatGPT answers were tailored to the account (its business, location or earlier chats), although we used temporary chats; Perplexity answered 4 of 8 English questions in Spanish, the account’s interface language; and Claude brought the account’s own business into three answers. Anyone tracking AI visibility should use clean sessions and repeat each prompt.

Limits

One run per question, in logged-in accounts, on one day. AI answers vary between runs: in a study of 2,961 runs by SparkToro and Gumshoe.ai, there was less than a 1 in 100 chance that ChatGPT or Google’s AI would give the same list of brands twice. One Spanish answer, from ChatGPT, was read only in part (see the notes). This test checks how engines explain this vocabulary; it is not a ranking of engines.

How we ran the test

Method

We used each engine’s consumer web app, logged in, with its default model and settings and web search left on automatic; ChatGPT ran in temporary chats. For Google, we searched google.com with the interface in each language and read the AI answer on the results page. Each question ran once. We read every answer in full, checked the definition against the primary source, and flagged any claim that was unsourced, outdated or contradicted by that source. A primary source is the originator’s own documentation or paper, such as Google Search Central for AI Overviews or the original arXiv paper for GEO. We logged the sources each answer showed and kept a screenshot of every answer; we do not publish them because they show account details.

The 16 prompts

In English

  1. What is GEO in digital marketing?
  2. What is AEO (answer engine optimization)?
  3. What is llms.txt, and does Google Search use it?
  4. What is query fan-out in Google’s AI search?
  5. What are Google AI Overviews?
  6. What is retrieval-augmented generation (RAG)?
  7. What is scaled content abuse according to Google?
  8. What is E-E-A-T, and is it a Google ranking factor?

In Spanish

  1. ¿Qué es el GEO en marketing digital?
  2. ¿Qué es el AEO (answer engine optimization)?
  3. ¿Qué es llms.txt y lo usa Google Search?
  4. ¿Qué es el query fan-out en la búsqueda con IA de Google?
  5. ¿Qué son las AI Overviews de Google?
  6. ¿Qué es la generación aumentada por recuperación (RAG)?
  7. ¿Qué es el abuso de contenido a gran escala según Google?
  8. ¿Qué es E-E-A-T y es un factor de posicionamiento de Google?

Opinion: Sourcelift’s view, with the evidence for and against it.

Our view: GEO is here to stay, and quality will count for more

We think GEO is a lasting discipline, not a passing label. The platforms now measure it themselves: Bing Webmaster Tools has reported AI citations since February 10, 2026, Google Analytics added an AI Assistant channel on May 13, 2026, and Search Console opened its generative AI performance report to every site on August 31, 2026. AI companies run separate crawlers for training, search and user requests, and Google now gives sites a setting for its AI features. And AI answers do their work without a click: Pew Research Center found that users clicked a link inside Google’s AI summary in 1% of visits to pages that showed one.

We also expect quality to be rewarded more over time. Google’s guide to its AI features, added on May 15, 2026, asks for non-commodity content, and that day Google made clear that its spam policies cover attempts to manipulate its AI answers. The rewriting tricks that looked promising in the 2023 GEO paper were largely ineffective in a 2025 benchmark, and answers change from run to run, so tactics that game one answer rarely last. What compounds is being worth citing: original research, consistent facts about your entity, and independent sources that confirm them.

The counterpoints are real. AI Mode drew 0.34% of US Google searches from January to April 2026 in SparkToro’s analysis of Similarweb data, which leaves out the Google app. Much of the evidence for GEO tactics is correlational or comes from vendors that sell GEO tools. And Google says that optimizing for its AI features is still SEO, so part of what is sold as GEO is SEO under a new name. That is why every entry here carries an evidence label.

What people call it: search demand for each label

The field has no settled name. In the US, “generative engine optimization” is the most searched label, but “answer engine optimization” and “AI search engine optimization” follow close behind, and at least six other labels have real demand. In Spain, the most searched labels mix languages, and literal translations of the English labels, such as “optimización para IA” or “SEO generativo”, barely register.

United States

LabelMonthly searches (US)
generative engine optimization 8,100
answer engine optimization 4,400
AI search engine optimization 4,400
AI search optimization 3,600
LLM SEO 1,600
SEO for AI 1,600
generative search optimization 1,300
LLM optimization 880
search everywhere optimization 480

Spain

LabelMonthly searches (Spain)
seo ia 590
seo para ia 390
generative engine optimization 320
que es geo 320
que es el geo 210
answer engine optimization 170
que es aeo 140
posicionamiento en ia 110
posicionamiento en inteligencia artificial 70
seo generativo / optimización para ia 10 or fewer

In Spanish, “GEO” also collides with geography and geomarketing: “geo españa” gets about 1,000 searches a month, which are about geography, and one of the sources Perplexity cited when we asked what GEO is was an article about geomarketing.

Source: Semrush keyword data, US and Spain databases, average monthly searches, retrieved September 13, 2026. Semrush sells SEO and AI visibility software, and its volumes are estimates.

How we built this glossary, and how we keep it right

We chose the 49 terms from three inputs: the vocabulary in the documentation of Google, OpenAI, Anthropic, Perplexity and Microsoft; search demand in Semrush; and a review of 11 existing GEO glossaries in English and Spanish, which showed where definitions disagree or go wrong. Every number, date, name and quotation traces to a source listed under its entry, opened and checked on September 29, 2026. Primary sources come first, and when we cite a vendor study we say what the vendor sells. Definitions are one sentence of 30 words or fewer, so they can be quoted without context. The Spanish version is written separately, not translated, and uses the names in Google’s Spanish documentation where they exist.

Every entry went through three checks before publication: sources reopened to confirm each quotation and figure; the English and Spanish versions compared fact by fact; and a final read to remove any claim without a source.

Found an error? Write to info@sourcelift.ai. We fix it, date the change in the changelog and credit you if you want. We plan to review the whole glossary, including crawler names and the engine test, every quarter and after each Google core or spam update. You are welcome to quote any definition with a link to its entry.

Written and maintained by Sourcelift. Sourcelift is a productized GEO and AEO agency that gets ecommerce, legal tech, health, fintech and SaaS brands cited inside AI answers. No definition recommends our services; links to them sit outside the definitions, in “Go deeper” lines and in clearly marked calls to action.

Changelog

  • September 29, 2026. First version: 49 terms in English and Spanish, an engine test of 80 answers, and search demand for the main labels.

From definitions to results

If you want to know where your brand stands in AI answers, start with the free diagnosis. It runs an initial set of buying prompts for your category on five engines and shows who gets named instead of you. The GEO audit measures your starting point on a locked prompt set and turns it into a 90-day plan; the sprint and the monthly programs do the work and report the change every month.

If AI doesn’t name you, it names a competitor.

The free diagnosis runs real buying prompts on five engines, shows who gets named instead of you and where you can win first.