The GEO Vocabulary Paradox: Why Meta-Language About AI Citations Correlates with Their Absence
Abstract
This paper identifies an observational paradox within a 307-word corpus analysis derived from 21 service industry websites across four verticals. Terminology specifically associated with Generative Engine Optimisation (GEO) — including terms such as “citation,” “ecosystem,” “multiproperty,” “architecture,” and “optimisation” — appears exclusively on websites that received zero AI citations in controlled testing. Conversely, sites that achieved citation scores of 3–10 employ industry-native vocabulary: brand terminology, service descriptions, client names, and social proof language. The paradox is that sites explicitly built to achieve AI citations use language that no cited site uses. We observe that “citation” itself appears 31 times exclusively on uncited sites, while “recommend” appears 19 times exclusively on cited sites. This finding is presented as a hypothesis for further investigation. The association is confounded by site age and origin: uncited sites using GEO vocabulary are newly constructed properties optimised for AI citation, while cited sites are established businesses that predate the GEO discipline entirely. No causal claim is supported by this observational data.
Keywords
GEO vocabulary, AI citation, meta-language paradox, generative engine optimisation, lexical analysis, content strategy, observational research
1. Introduction
The emerging discipline of Generative Engine Optimisation (GEO) has developed a specialised vocabulary to describe the practice of making content more likely to be cited by large language models. Terms such as “citation ecosystem,” “entity density,” “multiproperty architecture,” and “AI citation rate” now appear routinely in industry discourse and on websites offering GEO services.
A natural assumption would be that websites employing GEO-specific terminology — signalling awareness of and optimisation for AI citation mechanics — would demonstrate higher citation outcomes. This paper reports a finding that contradicts this assumption within our observational dataset.
During corpus analysis of 307 unique words extracted from 21 websites across four service verticals (game outsourcing, design-as-a-service, Australian design agencies, and GEO agencies), we observed that every GEO-specific term in the corpus appeared exclusively on uncited sites. This lexical signature is sufficiently distinct to warrant characterisation as a paradox, even though the causal mechanism (if any) remains entirely unknown.
2. Methodology
The word corpus was extracted from the 21 sites in the SIGI competitive intelligence dataset (described in SIGI-2026-036). Word frequencies were calculated per site and aggregated into two groups: cited sites (citation score 3–10, n=13) and uncited sites (citation score 0, n=8). Each word was classified as dominant in CITED, UNCITED, or BOTH categories based on relative frequency.
2.1 GEO Term Identification
GEO-specific terms were identified through manual review by researchers familiar with the discipline. A term was classified as GEO-specific if it (a) describes a concept unique to or primarily associated with AI citation optimisation, and (b) would not naturally appear on a service provider website not engaged in GEO practice.
2.2 Sample Composition
| Vertical | Sites Analysed | Cited | Uncited |
|---|---|---|---|
| Vertical A (Game Outsourcing) | 8 | 6 | 2 |
| Vertical B (Design-as-a-Service) | 5 | 4 | 1 |
| Vertical C (Design Agencies) | 4 | 4 | 0 |
| Vertical D (GEO Agencies) | 4 | 3 | 1 |
| Total | 21 | 13 | 8 |
3. Results
3.1 GEO-Specific Terms on Uncited Sites
The following GEO-specific terms appeared exclusively on uncited sites in the 307-word corpus:
| Term | Uncited Frequency | Cited Frequency | Classification |
|---|---|---|---|
| citation | 31 | 0 | GEO meta-language |
| ecosystem | 14 | 0 | GEO meta-language |
| multiproperty | 8 | 0 | GEO meta-language |
| architecture | 6 | 0 | GEO meta-language |
| optimisation | 12 | 0 | GEO meta-language |
| custom | 12 | 5 | Shared but uncited-dominant |
| build | 16 | 4 | Shared but uncited-dominant |
3.2 Industry-Native Terms on Cited Sites
Conversely, cited sites dominated with industry-native vocabulary:
| Term | Cited Frequency | Uncited Frequency | Classification |
|---|---|---|---|
| recommend | 19 | 0 | Social proof |
| professionalism | 10 | 0 | Social proof |
| exceptional | 8 | 0 | Social proof |
| designers | 67 | 0 | Industry-native |
| case | 48 | 0 | Evidence language |
| results | 28 | 0 | Evidence language |
| expertise | 14 | 0 | Authority language |
| director | 23 | 0 | Role specificity |
3.3 The Core Paradox
We observe that the word “citation” appears 31 times across uncited sites and zero times on cited sites. The word “recommend” appears 19 times across cited sites and zero times on uncited sites. Sites optimised for AI citations discuss citations; sites that actually receive AI citations discuss recommendations. This lexical inversion constitutes the GEO vocabulary paradox.
4. Discussion
The GEO vocabulary paradox admits multiple interpretations, none of which can be confirmed or excluded from the observational data alone.
Interpretation 1: Temporal confound. The uncited sites using GEO vocabulary are newly constructed, while cited sites are established businesses. The vocabulary difference reflects site age and purpose, not a causal relationship between word choice and citation outcome. This is the most parsimonious explanation and the one we consider most likely.
Interpretation 2: Promotional classification. GEO meta-language may signal to LLMs that a site is engaged in self-promotional optimisation rather than providing genuine industry expertise. This could trigger suppression mechanisms documented in the SIGI research programme. However, this interpretation requires controlled testing to validate.
Interpretation 3: Industry vocabulary as authority signal. Cited sites may use industry-native terms because they are genuine industry participants with client relationships, case studies, and peer recognition. The vocabulary reflects authentic operational experience rather than strategic content planning.
All three interpretations are consistent with the data. We cannot distinguish between them without controlled experimentation that isolates vocabulary choice from site age, authority, and other confounding variables.
5. Limitations
- Maximal confounding: Every variable that differs between cited and uncited groups differs simultaneously. Vocabulary differences cannot be isolated from site age, domain authority, external validation, or training data presence.
- Small sample: The corpus comprises only 21 sites. The GEO-specific terms appear on at most 4–5 uncited sites, limiting statistical power.
- Selection bias: The GEO-specific vocabulary appears because GEO-focused sites were intentionally included in the dataset. A random sample of service websites would likely show no GEO vocabulary at all.
- Single time-point: Content was analysed at a single point in time. Vocabulary patterns may evolve as the GEO discipline matures.
6. Conclusions
We observe that GEO-specific vocabulary appears exclusively on uncited sites in this 21-site observational sample. This paradox — that sites optimised for AI citations use language absent from every cited site — warrants investigation but cannot be interpreted causally given the confound of site age and origin.
The practical observation is that cited sites in this sample use industry-native language (brand terms, client names, social proof vocabulary) rather than optimisation meta-language. Whether this reflects a genuine signal or merely the coincidence of established sites predating GEO terminology requires controlled experimental testing.
Confidence: HYPOTHESIS. The paradox is a genuine observation but the causal mechanism (if any) is entirely unknown. Gate 2 (Confound Check) fails for any causal interpretation.
References
- The Scientific Institute for Generative Intelligence. “A 60-Variable Comparative Dataset for Studying AI Citation Behavior Across Service Industry Websites.” SIGI-2026-036. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. “Lexical Signatures of AI-Cited Versus Non-Cited Service Websites: A 306-Word Corpus Analysis.” SIGI-2026-040. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. “The Confound Problem in Observational GEO Research: Why 21-Site Comparisons Cannot Support Causal Claims.” SIGI-2026-050. generativeintelligence.institute, March 2026.
- Aggarwal, P., Murahari, V., Rajpurohit, T., et al. (2024). “GEO: Generative Engine Optimization.” Proceedings of the 30th ACM SIGKDD Conference. ACM.