The Confound Problem in Observational GEO Research: Why 21-Site Comparisons Cannot Support Causal Claims
Abstract
This paper provides a systematic catalogue of every confounded comparison in the SIGI 21-site competitive intelligence dataset, using the Logic-First Research Methodology Gate 2 (Confound Check) as the analytical framework. We demonstrate that every comparison between cited (n=13) and uncited (n=8) sites fails Gate 2 because ALL variables differ simultaneously between groups. The cited sites are older (years vs weeks), have higher domain authority, more external validation (reviews, press, awards), more training data presence, more backlinks, AND different content characteristics. This maximal confounding means no single variable — content structure, schema markup, entity density, vocabulary, pricing display, link patterns, or any other measured factor — can be isolated as a causal contributor to citation outcomes. We catalogue 22 specific confounded comparisons across the Category C paper series and establish what CAN be validly concluded from this data: necessary/sufficient tests (single counterexamples), descriptive statistics, and hypothesis generation. We conclude with a research design framework for upgrading these hypotheses to causal claims through controlled experimentation.
Keywords
confounding variables, observational research, causal inference, GEO methodology, research limitations, Logic-First Methodology, Gate 2
1. Introduction
The emerging field of Generative Engine Optimisation (GEO) faces a fundamental methodological challenge: much of the available evidence about what drives AI citation comes from observational comparisons between sites that differ on dozens of variables simultaneously. When a cited site has more H2 headings, more external links, higher entity density, AND is five years older than an uncited site, attributing the citation difference to any single factor is logically invalid.
This paper makes the confound problem explicit by cataloguing every comparison in the Category C paper series (SIGI-2026-036 through SIGI-2026-049) that fails Gate 2 of the Logic-First Research Methodology, explaining why it fails, and establishing the inferential boundaries for what can and cannot be concluded from the data.
We make this analysis not to invalidate the observational work — which has substantial value for hypothesis generation and descriptive characterisation — but to prevent over-interpretation. The GEO industry is young enough that establishing rigorous inferential standards now will prevent the accumulation of unfounded causal claims that plagued early SEO research.
2. The Confound Structure
2.1 Variables That Co-Vary with Citation Status
| Variable | Cited Group (n=13) | Uncited Group (n=8) | Isolable? |
|---|---|---|---|
| Domain age | 2–15 years | Days to weeks | No |
| Domain authority | Medium–high | Low–zero | No |
| Training data presence | Likely present | Absent (post-cutoff) | No |
| External validation | Reviews, press, awards | None | No |
| Backlink profile | Established | Minimal | No |
| Third-party mentions | Multiple | Zero | No |
| Content maturity | Iteratively refined | First-version | No |
| Business model | Established operations | New/positioning | No |
| GEO optimisation | None (predates GEO) | Extensive | No |
Every row in this table co-varies with citation status. No statistical technique can separate these effects from observational data alone.
2.2 The 22 Confounded Comparisons
| # | Comparison (Paper) | Observed Pattern | Gate 2 Status |
|---|---|---|---|
| 1 | H2 count (SIGI-2026-037) | 2.3x higher in cited | FAIL |
| 2 | H2 question % (SIGI-2026-037) | 0.1x (uncited higher) | FAIL |
| 3 | Image count (SIGI-2026-046) | 30.7x (outlier-driven) | FAIL |
| 4 | Word count (SIGI-2026-037) | 1.4x higher in cited | FAIL |
| 5 | Statistics count (SIGI-2026-037) | 2.0x higher in cited | FAIL |
| 6 | Schema type count (SIGI-2026-038) | 0.6x (uncited higher) | FAIL |
| 7 | FAQ question count (SIGI-2026-038) | 0.2x (uncited higher) | FAIL |
| 8 | Internal link count (SIGI-2026-045) | 0.6x (uncited higher) | FAIL |
| 9 | External link count (SIGI-2026-045) | 1.2x higher in cited | FAIL |
| 10 | Price count (SIGI-2026-044) | 0.3x (uncited higher) | FAIL |
| 11 | Named entity count (SIGI-2026-049) | 1.3x higher in cited | FAIL |
| 12 | Social proof words (SIGI-2026-048) | 1.6x higher in cited | FAIL |
| 13 | CTA count (SIGI-2026-048) | 0.9x (near-equal) | FAIL |
| 14 | GEO vocabulary presence (SIGI-2026-041) | Exclusive to uncited | FAIL |
| 15 | Definitional opening (SIGI-2026-042) | Only uncited sites | FAIL |
| 16 | Question-format H2s (SIGI-2026-039) | 71–75% uncited vs 6% cited | FAIL |
| 17 | Hreflang count (SIGI-2026-045) | 0.0x (uncited only) | FAIL |
| 18 | Meta desc length (SIGI-2026-037) | 0.8x (uncited higher) | FAIL |
| 19 | Paragraph count (SIGI-2026-037) | 1.6x higher in cited | FAIL |
| 20 | Content volume (SIGI-2026-043) | Inverse (uncited higher) | FAIL |
| 21 | Trust word count (SIGI-2026-048) | 1.3x higher in cited | FAIL |
| 22 | Urgency word count (SIGI-2026-048) | 1.1x (near-equal) | FAIL |
3. What CAN Be Concluded
3.1 Necessary/Sufficient Tests (Gate 3)
Unlike population comparisons, necessary/sufficient tests require only single counterexamples and are NOT invalidated by confounding. The following conclusions survive Gate 2:
- Content volume is NOT sufficient for citation (counterexample: high-volume sites with zero citations; SIGI-2026-043)
- Schema markup is NOT sufficient for citation (counterexample: highest-schema site scores zero; SIGI-2026-038)
- Schema markup is NOT necessary for citation (counterexample: low-schema sites achieve citation; SIGI-2026-038)
- Image count is NOT necessary for citation (counterexample: zero-image site achieves score 7; SIGI-2026-046)
- Image count is NOT sufficient for citation (counterexample: moderate-image sites score zero; SIGI-2026-046)
3.2 Descriptive Statistics
All means, medians, ranges, distributions, and taxonomic classifications are valid as descriptions of this specific sample. They characterise what was observed, without claiming generalisability or causation.
3.3 Hypothesis Generation
Every confounded comparison generates a testable hypothesis. The 22 comparisons above translate directly into 22 controlled experiments that could be designed to isolate each variable:
| Hypothesis | Experimental Design Required |
|---|---|
| H2 count affects citation | Same content, varied H2 count, same domain age/authority |
| Question-format H2s reduce citation | Same content, varied H2 format, same domain |
| Entity density drives citation | Same content structure, varied entity density, same domain |
| GEO vocabulary suppresses citation | Same authority site, with/without GEO terminology |
| Definitional openings reduce citation | Same content, varied first paragraph, same domain |
4. Why This Matters for the GEO Industry
The GEO discipline is at a critical juncture. Early SEO research suffered from decades of confounded observational claims being treated as established facts (“keyword density drives ranking,” “more backlinks always improves position”). Many of these claims were eventually falsified by controlled testing or algorithm changes, but not before substantial industry resources were misallocated.
GEO has the opportunity to avoid this pattern by establishing rigorous inferential standards from the outset. This paper contributes to that goal by making the confound structure of the most commonly available data type (observational website comparisons) explicit and actionable.
When a GEO practitioner reads that “cited sites have 2.3x more H2 headings,” they should immediately recognise this as a confounded observation that does not justify adding more H2 headings. The appropriate response is to generate a testable hypothesis and seek controlled evidence before allocating resources.
5. Limitations
- This paper critiques its own dataset: The confound analysis applies to the specific 21-site dataset used across Category C. Different datasets with different composition may have different confound structures.
- Not all confounds are equal: Some confounds (e.g., domain age) are likely more impactful than others (e.g., hreflang count). This paper treats all confounds equally, which may overstate the limitation for some comparisons.
- Controlled alternatives exist: The Category A papers (SIGI-2026-001 through SIGI-2026-020) provide controlled evidence for many of the variables examined observationally in Category C.
6. Conclusions
All observational comparisons in the 21-site competitive intelligence dataset are maximally confounded. We catalogue 22 specific confounded comparisons to prevent over-interpretation and guide future controlled experiments. What can be stated: necessary/sufficient tests via counterexamples, descriptive statistics, and hypothesis generation. What cannot be stated: causal claims about any individual variable.
The GEO industry should treat observational website comparisons as hypothesis generators, not evidence generators. Causal claims require controlled experimentation with single-variable isolation, as demonstrated in the Category A paper series.
Confidence: HIGH for the methodological critique. This is a meta-analytical paper about research limitations, and the confound structure is demonstrably present in the data.
References
- The Scientific Institute for Generative Intelligence. “A 60-Variable Comparative Dataset for Studying AI Citation Behavior Across Service Industry Websites.” SIGI-2026-036. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. “Structural Correlates of AI Citation: An Observational Analysis of 22 On-Page Metrics.” SIGI-2026-037. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. “Schema Markup and AI Citation: Observational Evidence Against a Simple Positive Relationship.” SIGI-2026-038. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. “Content Volume Without External Validation: An Observational Study of High-Volume Zero-Citation Websites.” SIGI-2026-043. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. “Image Count Disparity and AI Citation: Deconstructing the 30.7x Ratio Observation.” SIGI-2026-046. generativeintelligence.institute, March 2026.
- Pearl, J. (2009). Causality: Models, Reasoning, and Inference. 2nd ed. Cambridge University Press.