The Evidence Hierarchy for Generative Engine Optimization: Mapping Research Methods to Confidence Levels
Abstract
We propose an evidence hierarchy specifically adapted for GEO research, accounting for the unique challenges of studying AI system behaviour. The hierarchy maps seven evidence levels from anecdote (Level 1) to meta-analysis (Level 7) with GEO-specific method classifications: introspective self-report maps to Level 1-2, observational comparisons to Level 2-3, single-variable isolation probes to Level 4, and cross-model replicated experiments to Level 5. We identify five AI-specific replication dimensions (model versions, time points, query types, industry domains, and temperature settings) that distinguish AI behaviour research from traditional social science. We address the non-determinism problem: even at temperature=0, GPU floating-point operations produce variable outputs, requiring minimum repetition counts per condition. The critical distinction between universal claims ("X always causes Y") and probabilistic claims ("X increases the probability of Y") is mapped to evidence levels, with universal claims requiring Level 5+ evidence and probabilistic claims permitted at Level 4 with appropriate scoping.
Keywords
evidence hierarchy, GEO methodology, AI-specific replication, non-determinism, confidence levels, universal vs probabilistic claims, research methods mapping
1. Introduction
Evidence hierarchies are well-established in medical research, where they rank study designs from case reports to randomised controlled trials to meta-analyses. However, AI behaviour research presents unique challenges that make direct transplantation of existing hierarchies insufficient.
AI systems differ from human subjects in ways that affect research design: they can be queried repeatedly with identical inputs (enabling replication), their behaviour changes with model updates (creating temporal instability), they operate non-deterministically even at low temperature settings (requiring multiple runs), and they may behave differently across model architectures (requiring cross-model validation). An evidence hierarchy for GEO must account for these characteristics.
2. The GEO Evidence Hierarchy
2.1 Method-to-Level Mapping
| Level | Type | GEO Research Method | Example from SIGI Programme |
|---|---|---|---|
| 1 | Anecdote | Single observation, one query, one model | Noticing a specific citation in one AI response |
| 2 | Case Study | LLM introspective self-report, documented pattern | Trust signal taxonomy from model self-description |
| 3 | Observational Correlation | Multi-site comparison without variable isolation | Comparing cited vs. uncited sites across 50 variables |
| 4 | Controlled Experiment | Single-variable isolation probe, single model | Rating probe: 19 variations, one variable, held constants |
| 5 | Randomised Controlled | Cross-model replicated experiment (3+ models) | Rating probe replicated on 3 AI systems (upgrade target) |
| 6 | Systematic Review | Synthesis of all studies on a GEO factor | All evidence on entity density across all studies |
| 7 | Meta-Analysis | Statistical synthesis with effect sizes | Pooled effect size of entity density across all experiments |
2.2 AI-Specific Replication Dimensions
| Dimension | Rationale | Minimum for Level 5 |
|---|---|---|
| Model versions | Different architectures may implement different evaluation logic | 3+ distinct AI systems |
| Time points | Model updates and retrieval changes affect behaviour over time | 3+ time points spanning 3+ months |
| Query types | Findings may be specific to certain query structures | 3+ query type variations |
| Industry domains | Effects may be domain-specific (e.g., services vs. products) | 3+ industry verticals |
| Temperature settings | Even temperature=0 produces non-deterministic outputs on GPUs | Minimum 3 runs per condition at temperature=0 |
2.3 The Non-Determinism Problem
A critical challenge in AI behaviour research is that identical inputs do not guarantee identical outputs, even at temperature=0. GPU floating-point operations involve parallel computation where the order of operations can vary between runs, producing slightly different results. This means every experiment must include multiple repetitions per condition, and findings must be stated in terms of tendencies rather than absolutes unless consistency is verified across all repetitions.
2.4 Universal vs. Probabilistic Claims
| Claim Type | Example | Minimum Evidence Level |
|---|---|---|
| Universal ("always") | "Entity density always increases citation" | Level 5+ (cross-model replication) |
| Probabilistic ("tends to") | "Higher entity density tends to increase citation under these conditions" | Level 4 (controlled experiment) |
| Correlational ("associated with") | "Entity density is associated with citation" | Level 3 (observational) |
| Observational ("we observed") | "We observed higher citation in entity-dense content" | Level 2 (case study) |
3. Application: Classifying SIGI Findings
The SIGI research programme demonstrates the hierarchy in practice. The 10 probe experiments (129 data points across 10 variables) are classified at Level 4: they use single-variable isolation, produce replicable results within a single model, but have not yet been cross-model validated. All probe findings are stated using Level 4 language: "under these controlled conditions, X causes Y."
The observational studies (baseline search analysis, paid placement study, subscriber feedback loop) are classified at Level 2-3. They identify patterns and associations but cannot isolate causal variables. Findings are stated as "we observe" or "X is associated with Y."
No SIGI finding currently meets Level 5 requirements. The upgrade path for all Level 4 findings is identical: replicate across 3+ AI models, 3+ time points, and 3+ industry domains.
4. Limitations
- Adapted hierarchy: The 7-level structure is adapted from medical research. Whether these specific levels are optimal for AI behaviour research is an open question.
- Replication thresholds: The minimum of 3 models/time points/domains is pragmatic rather than theoretically derived. Different thresholds could be justified.
- Rapidly evolving field: AI systems change faster than the research cycle, potentially requiring time-stamped findings rather than persistent claims.
5. Conclusions
We propose an evidence hierarchy specifically adapted for GEO research, accounting for the unique challenges of studying AI system behaviour. The hierarchy maps standard research methods to GEO-specific implementations, defines five AI-specific replication dimensions, addresses the non-determinism problem, and distinguishes universal from probabilistic claims with corresponding minimum evidence levels.
The practical contribution is a shared vocabulary for evaluating GEO research quality. When a practitioner encounters a claim about AI citation behaviour, the hierarchy provides a framework for asking: at what evidence level was this established, and what would be required to strengthen it?
Note: This is a methodological framework paper presenting a standard for evidence evaluation in GEO research.
References
- The Scientific Institute for Generative Intelligence. "The Logic-First Research Methodology: An Evidentiary Standard for Generative Engine Optimization Claims." SIGI-2026-066. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Seven Logic Gates for GEO Research: A Framework for Validating Claims About AI Citation Behavior." SIGI-2026-067. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Single-Variable Isolation in AI Behavior Research: The Probe Experiment Design Template." SIGI-2026-069. generativeintelligence.institute, March 2026.