The Confound Epidemic in Published GEO Research: How Observational Comparisons Are Misrepresented as Causal Findings
Abstract
The emerging field of generative engine optimisation (GEO) is building its knowledge base primarily on observational comparisons between websites that differ on multiple dimensions simultaneously. This paper analyses the systematic pattern by which confounded observations are presented as causal findings, using examples from both the SIGI research programme's own observational data and common patterns in published GEO commentary. We demonstrate that comparisons between newly launched, uncited websites and established, cited websites fail Gate 2 (Confound Check) of the Logic-First Methodology because every variable -- site age, domain authority, training data presence, external validation, backlink profile, content maturity, and brand recognition -- differs simultaneously. From our own 21-site dataset, we catalogue 10 specific correlational observations that cannot support causal claims despite appearing striking (including a 30.7x image count ratio and an inverse schema-citation correlation). We identify the valid logical operations that remain available from confounded data -- specifically, necessary/sufficient tests through single counterexamples -- and propose a mandatory confound matrix disclosure standard for all GEO research publications.
Keywords
confounds, causal inference, observational research, GEO methodology, Gate 2, confound matrix, necessary conditions, sufficient conditions
1. Introduction
Every week, new GEO research is published claiming to identify factors that "drive," "cause," or "determine" AI citation behaviour. The typical methodology is to compare websites that are cited by LLMs with websites that are not, identify differences, and attribute citation outcomes to those differences. This methodology has a fundamental flaw: the groups being compared differ on every measurable dimension simultaneously.
This is not a subtle statistical concern. It is a basic violation of causal inference that would be rejected in any established scientific field. In medical research, no journal would publish a study claiming that a treatment works based on comparing patients who chose the treatment (who may differ in age, health, socioeconomic status, and motivation) with patients who did not. The GEO field routinely publishes the equivalent.
We write this paper as a critique of our own programme as much as the broader field. The SIGI research programme's Category C data (21-site observational analysis) contains correlations that appear striking but are maximally confounded. We use our own data as the primary case study to demonstrate the problem.
2. The Anatomy of a Confounded GEO Comparison
In the SIGI 21-site dataset, the uncited websites share a cluster of characteristics: they are the newest sites (days to weeks old), have the lowest domain authority, have zero external validation (no third-party reviews, no press mentions, no industry citations), have zero training data presence, and have the highest levels of GEO-specific optimisation (schema markup, FAQ structures, question-format headings).
The cited websites are simultaneously older, higher-authority, richer in external validation, more present in training data, and less explicitly optimised for AI citation. When we observe that question-format H2 headings correlate with lower citation, we cannot determine whether the correlation reflects the headings themselves or the fact that only new, low-authority sites use that heading format in our sample.
| Variable | Uncited sites (n=8) | Cited sites (n=13) | Isolated? |
|---|---|---|---|
| Site age | Days to weeks | Months to years | No |
| Domain authority | Near zero | Moderate to high | No |
| Training data presence | Absent | Present | No |
| External reviews | Zero | Multiple | No |
| Press mentions | Zero | One or more | No |
| Backlink profile | Minimal | Established | No |
| Content maturity | First version | Iterated | No |
| Brand recognition | None | Established | No |
| Schema type count | Higher (avg 11) | Lower (avg 7) | No |
| Question-format H2s | 54% | 6% | No |
With zero isolated variables, no causal conclusion about any individual factor is possible from this comparison. The confound matrix makes this visually obvious.
3. Catalogue of Confounded Claims
We catalogue 10 specific correlational observations from our own data that cannot support causal claims, despite their apparent significance:
- Image count ratio of 30.7x -- driven by portfolio-heavy outliers, not a predictive signal
- Question-format H2 inverse correlation (0.1x) -- confounded by site age and authority
- Schema type count inversely correlated -- uncited sites have more schema types, but they also differ on every other dimension
- Internal link density negatively correlated -- uncited sites have more internal links but are also newer and more architecturally complex
- GEO vocabulary appearing only on uncited sites -- because only new, GEO-optimised sites use GEO vocabulary
- Definitional first paragraphs correlating with zero citation -- because only new sites explain what they are
- Content volume without citation (500+ pages, zero citations) -- those pages are days old
- FAQ schema count inversely correlated -- same confound as schema type count
- Social proof word count difference (14 vs 9) -- reflects site maturity, not a causal factor
- Statistics count difference (4 vs 2) -- established sites have more data to cite
Every one of these observations is genuine. None of them supports a causal claim.
4. What Remains Valid
Confounded data is not useless. Two types of logical operation remain valid because they require only single counterexamples, not population-level comparisons:
Necessary condition test: If Y (citation) occurs without X (a specific feature), then X is not necessary for Y. Finding a cited site without schema markup proves that schema is not necessary for citation. This is valid regardless of confounds.
Sufficient condition test: If X is present but Y is absent, then X is not sufficient for Y. Finding an uncited site with extensive schema proves that schema is not sufficient for citation. Also valid regardless of confounds.
From our dataset, we can validly state: schema markup is neither necessary nor sufficient for AI citation. Content volume is not sufficient. FAQ structure is not sufficient. Question-format H2s are not necessary for non-citation (cited sites also use them, albeit rarely). These are logically rigorous conclusions that survive confound analysis.
5. A Proposed Disclosure Standard
We propose that all GEO research publications include a mandatory confound matrix listing every variable that differs between compared groups. The matrix format is simple: Variable | Group A value | Group B value | Isolated? (Yes/No). If the "Isolated?" column contains any "No" entries, the publication must explicitly state that causal claims about individual variables are not supported.
This disclosure does not prevent publication of observational research. It simply requires that the limitations be visible to the reader rather than buried in a footnote or absent entirely.
6. Conclusions
We identify a systematic pattern of confounded observational comparisons being presented as causal findings in GEO research, including in our own programme's observational data. We catalogue 10 specific confounded claims from our dataset and demonstrate that while the correlations are genuine, they cannot support causal attribution. We propose a mandatory confound matrix disclosure standard and identify the logical operations (necessary/sufficient tests) that remain valid even in maximally confounded data.
The field of GEO will mature more rapidly if it adopts the evidentiary standards that established sciences take for granted. Confounded observations are a starting point for research, not an endpoint.
Confidence: HIGH for the methodological critique. The confound analysis applies standard causal inference principles to GEO research.
References
- The Scientific Institute for Generative Intelligence. "The Confound Problem in Observational GEO Research." SIGI-2026-050. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "A 60-Variable Comparative Dataset." SIGI-2026-036. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Limitations of the SIGI Research Program." SIGI-2026-091. generativeintelligence.institute, March 2026.
- Aggarwal, P., Murahari, V., Rajpurohit, T., et al. (2024). "GEO: Generative Engine Optimization." KDD 2024. ACM.