SIGI-2026-100

A Meta-Analysis Framework for Consolidating GEO Evidence: Toward Level 7 Confidence

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper -- the final publication in the SIGI 100-paper research programme -- presents a prospective framework for the eventual meta-analysis of generative engine optimisation (GEO) findings. The programme has produced 10 controlled single-variable experiments (Evidence Level 4), multiple observational studies (Levels 2-3), and a comprehensive methodological infrastructure grounded in the Logic-First Research Methodology with its seven logic gates and seven-tier evidence hierarchy. However, Level 7 meta-analytic confidence -- the point at which we may state that the weight of all available evidence supports a conclusion -- remains beyond the reach of any single research programme. Achieving it requires independent replication by separate research groups across multiple LLM platforms, temporal windows, industry domains, and geographic contexts. We propose four standardised effect size measures for GEO research: Citation Probability Difference (CPD), Sentiment Shift Magnitude (SSM), Position Displacement Index (PDI), and Language Confidence Delta (LCD). We outline a phased 3-5 year roadmap from the current Level 4 foundation to Level 7 meta-analytic synthesis, identify the specific replication gaps that must be filled, and issue an open call for independent researchers to replicate, challenge, and extend these findings using the publicly available datasets and protocols. The framework presented here is classified at Evidence Level 3-4: the roadmap itself constitutes a structured proposal grounded in established meta-analytic methodology, not an empirical finding.

Keywords

meta-analysis, generative engine optimisation, evidence hierarchy, replication, effect size, GEO, large language model, research methodology, Level 7 confidence, independent verification

1. Introduction

Generative engine optimisation is a nascent field. The emergence of large language models as mediators of information retrieval, service provider evaluation, and consumer decision-making has created an urgent need for rigorous empirical research into the signals that influence AI-generated citations and recommendations. The SIGI research programme was designed to address this need through a systematic, methodology-first approach: 100 papers spanning controlled experiments, observational studies, competitive analyses, platform behaviour investigations, and methodological frameworks.

This final paper in the programme does not present new empirical findings. Instead, it asks the necessary forward-looking question: how do we move from what we know now to what the field needs to know? Specifically, how do we consolidate the evidence produced by this programme and by future independent research into the kind of robust, meta-analytic synthesis that can support confident, generalisable claims about how LLMs process and prioritise information sources?

The Logic-First Research Methodology (SIGI, 2026) establishes a seven-level evidence hierarchy, with Level 7 -- meta-analysis -- at the apex. Level 7 permits the statement that "the weight of all available evidence shows X." This is the strongest possible evidentiary claim, and it requires conditions that no single research programme can satisfy: multiple independent studies, pooled effect sizes with heterogeneity analysis, and systematic assessment of publication bias.

The current programme provides a Level 4 foundation: controlled experiments with isolated variables, pre-specified hypotheses, and validated findings that have passed all seven logic gates. The gap between Level 4 and Level 7 is not merely quantitative (more studies) but qualitative (independent replication, cross-context generalisation, and formal statistical synthesis). This paper maps that gap and proposes a structured pathway to bridge it.

1.1 Scope and Limitations of This Paper

This is a prospective framework paper, not an empirical study. The roadmap we propose is grounded in established meta-analytic methodology from evidence-based medicine and social science, adapted for the specific challenges of GEO research. The framework itself is classified at Evidence Level 3-4: it synthesises existing methodological knowledge and applies it to a new domain, but its predictions about timelines and feasibility remain untested. We state only what the evidence supports, consistent with the Logic-First methodology that governs the entire programme.

2. Methodology

The framework development followed a structured approach integrating established meta-analytic standards with the specific requirements of GEO research.

2.1 Foundational Standards

We drew on three established meta-analytic frameworks: the Cochrane Handbook for Systematic Reviews of Interventions, the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, and the Campbell Collaboration standards for social science synthesis. These were selected because they represent the most rigorous and widely adopted protocols for evidence consolidation across empirical disciplines.

2.2 Adaptation for GEO Research

Standard meta-analytic methods assume relatively stable interventions and outcome measures. GEO research presents three unique challenges that required methodological adaptation. First, the intervention targets -- large language models -- undergo frequent version updates that can alter signal processing overnight, creating temporal instability not found in most intervention research. Second, the outcome space is multidimensional: citation, sentiment, position, hedging language, and recommendation strength are distinct but correlated outcomes. Third, the field is so new that there are no prior meta-analyses to serve as templates, requiring us to propose rather than adopt standardised measures.

2.3 Evidence Classification

All claims in this paper are stated at the evidence level warranted by their supporting data, per the Logic-First methodology. The framework itself is classified at Evidence Level 3-4. The empirical findings it references (from the broader SIGI programme) retain their original evidence classifications. No claim in this paper exceeds what the underlying evidence supports.

3. The Current Evidence Base

Before proposing a path forward, we must accurately characterise what exists. The SIGI 100-paper programme produced findings across four evidence tiers, each with distinct implications for meta-analytic synthesis.

3.1 Validated Findings (Level 4)

Ten controlled single-variable probe experiments produced findings that passed all seven logic gates. These include sentiment threshold dynamics for star ratings, the price-credibility U-curve, the domain age null result, content depth sweet spots, position primacy bias, review volume effects, awards and credential thresholds, client count ranges, word count inversion points, and entity density curves. Each experiment isolated a single variable while holding all others constant, enabling causal claims scoped to the specific conditions tested.

3.2 Strongly Supported Findings (Level 2-3)

Observational studies with partial controls produced findings including the two-layer architecture hypothesis (search retrieval and synthesis as distinct stages), training data asymmetry, the de-linking principle, press coverage effects, and paid placement bias mechanisms. These pass most logic gates but fail Gate 2 (confound isolation) or Gate 6 (replication) partially.

3.3 Hypotheses (Level 1-2)

Confounded observational comparisons generated plausible but unvalidated hypotheses about question-format headings, FAQ schema effects, schema quantity, internal link density, and image count correlations. These fail Gate 2 due to simultaneous variation of multiple variables and cannot support causal claims.

3.4 Introspective Findings

LLM self-reports about citation processes, including the 77 trust signal rankings and the eight-stage pipeline model, are classified as informed hypotheses. Per the Logic-First methodology, LLM self-reports about their own behaviour are treated as confabulations rather than data, useful for generating testable predictions but not for making definitive claims.

3.5 The Critical Gap

The entire evidence base originates from a single research programme, tested primarily on a single LLM platform, at a single point in time. This is the fundamental limitation that prevents any current finding from exceeding Level 4 confidence regardless of methodological rigour. Level 5 (randomised controlled, generalisable) requires cross-context replication. Level 6 (systematic review) requires synthesis of multiple independent studies. Level 7 (meta-analysis) requires formal statistical pooling of effect sizes from those independent studies.

4. Proposed Standardised Effect Size Measures

Meta-analysis requires commensurable effect sizes -- standardised measures that allow findings from different studies to be compared and pooled. We propose four primary effect size measures for GEO research, each targeting a distinct outcome dimension.

4.1 Citation Probability Difference (CPD)

The most fundamental GEO outcome is whether a source is cited or not. CPD measures the absolute difference in citation probability between a treatment and control condition. For a given variable X, CPD = P(citation | X present) - P(citation | X absent). This mirrors the risk difference metric used in clinical trials. We recommend reporting both the absolute CPD and the relative risk ratio to facilitate pooling. Confidence intervals should be calculated using the Wilson score method for proportions, which performs well with the moderate sample sizes typical of GEO experiments (30-100 queries per condition).

4.2 Sentiment Shift Magnitude (SSM)

Several SIGI probes identified threshold-based sentiment transitions. SSM captures the magnitude of sentiment change at a threshold, operationalised as the standardised mean difference (Cohen's d) in positive, negative, and neutral mention counts between adjacent conditions that span a threshold. For the ratings probe, SSM at the 3.8 threshold would be calculated as the standardised difference in negative mention counts between ratings of 3.7 and 3.8. We recommend reporting SSM with 95% confidence intervals using the Hedges' g correction for small sample bias.

4.3 Position Displacement Index (PDI)

When LLMs produce ranked lists or ordered recommendations, position matters. PDI measures the number of positions a source moves in response to a variable change, standardised by the total number of positions available. PDI = (Position_control - Position_treatment) / Total_positions. This bounded metric (-1 to +1) enables comparison across studies with different list lengths. We recommend bootstrap confidence intervals for PDI given the ordinal nature of the underlying data.

4.4 Language Confidence Delta (LCD)

LLM responses vary not just in whether they cite a source but in the confidence of the language used. LCD captures the shift in hedging versus endorsement language, operationalised through a pre-defined lexicon of hedging markers ("may," "could," "some suggest") and endorsement markers ("established," "leading," "widely recognised"). LCD = (Endorsement_count_treatment - Endorsement_count_control) / Total_evaluative_words. Standardisation against total evaluative word count controls for response length variation.

4.5 Cross-Measure Considerations

These four measures are deliberately designed to be independent but complementary. A given intervention might increase citation probability (CPD) without changing sentiment (SSM), or shift position (PDI) without altering confidence language (LCD). Requiring researchers to report all four measures for each finding creates the multidimensional outcome space that GEO meta-analysis will require. We also recommend that all studies report the raw data underlying these measures to enable re-analysis and the calculation of alternative effect sizes as the field matures.

5. The Roadmap to Level 7

The path from the current Level 4 foundation to Level 7 meta-analytic confidence involves three sequential phases, each dependent on the successful completion of the preceding phase.

5.1 Phase 1: Independent Replication (Years 1-2)

The immediate priority is independent replication of the 10 validated Level 4 findings. Replication must occur across at least four dimensions identified in the Logic-First methodology:

Table 1. Replication dimensions required for evidence upgrading
DimensionMinimum RequirementRationale
Across LLM platforms3+ platforms (e.g., GPT, Claude, Gemini)Findings may be model-specific artefacts
Across time3+ time points spanning 12+ monthsModel updates can shift signal processing
Across query typesFactual, evaluative, comparative queriesDifferent query types may activate different pathways
Across domains3+ industry verticalsDomain-specific training data may create asymmetries

Successful replication across two or more dimensions upgrades a finding from Level 4 to Level 5. Failure to replicate is equally valuable: it identifies boundary conditions and moderating variables that refine the original findings. We estimate that 5-10 independent research groups conducting replication studies during this phase would produce the minimum evidence base for the next phase.

5.2 Phase 2: Systematic Review (Years 2-3)

Once a sufficient body of independent studies exists, a systematic review becomes possible. This phase follows PRISMA guidelines adapted for GEO research:

  1. Protocol registration: Pre-register the review protocol including search strategy, inclusion/exclusion criteria, and analysis plan.
  2. Comprehensive search: Identify all studies addressing GEO signals, including unpublished work and null results to mitigate publication bias.
  3. Quality assessment: Evaluate each study against the seven logic gates to assign evidence classifications.
  4. Narrative synthesis: Classify findings by convergence, divergence, or conditionality across the available evidence.

A completed systematic review upgrades the synthesised findings to Level 6, permitting the statement that all studies examined converge on a conclusion (or documenting where they diverge).

5.3 Phase 3: Formal Meta-Analysis (Years 3-5)

The final phase applies formal meta-analytic statistical methods to the pooled evidence:

  1. Effect size pooling: Combine CPD, SSM, PDI, and LCD values across studies using random-effects models (given the expected heterogeneity in GEO research contexts).
  2. Heterogeneity analysis: Compute I-squared and tau-squared statistics to quantify between-study variance. Conduct moderator analyses to identify sources of heterogeneity (LLM platform, temporal period, domain, query type).
  3. Publication bias assessment: Employ funnel plots, Egger's test, and trim-and-fill methods to evaluate and adjust for selective reporting.
  4. Sensitivity analysis: Test the robustness of conclusions by excluding outlier studies, varying inclusion criteria, and comparing fixed-effects versus random-effects models.

A completed meta-analysis with adequate study coverage and manageable heterogeneity constitutes Level 7 evidence -- the apex of the hierarchy. At that point, and only at that point, it becomes permissible to state that the weight of all available evidence supports a given GEO conclusion.

5.4 Estimated Timeline and Dependencies

Table 2. Projected timeline from Level 4 to Level 7 confidence
PhaseTimelinePrerequisiteOutputEvidence Level
Foundation (current)Complete--10 validated single-variable probesLevel 4
Phase 1: ReplicationYears 1-2Open data access, community engagement5-10 independent replications per findingLevel 5
Phase 2: Systematic reviewYears 2-3Sufficient independent studiesPRISMA-compliant synthesisLevel 6
Phase 3: Meta-analysisYears 3-5Standardised effect sizes, 15+ studiesPooled effects with heterogeneity analysisLevel 7

These timelines assume active community participation. Without independent replication, the evidence base remains at Level 4 indefinitely regardless of the volume of research produced by any single programme.

6. Challenges Specific to GEO Meta-Analysis

Standard meta-analytic methods assume conditions that GEO research may not satisfy. We identify four challenges that require methodological innovation.

6.1 Temporal Instability

LLM behaviour changes with model updates, sometimes dramatically. A finding validated on GPT-4o in March 2026 may not hold for GPT-5 in 2027. This creates a unique versioning problem: should meta-analyses pool across model versions, or treat each version as a separate population? We recommend treating model version as a moderator variable in meta-regression, allowing the analysis to detect whether findings are version-dependent while still pooling across versions when heterogeneity is low.

6.2 Non-Independence of Platforms

LLM platforms are not fully independent: they share training data sources, architectural innovations, and even directly train on each other's outputs. This violates the independence assumption of standard meta-analysis. We recommend modelling platform as a random effect rather than treating cross-platform replication as equivalent to fully independent studies. A correction factor based on estimated training data overlap may be necessary as the field matures.

6.3 Publication Bias in a New Field

New fields are particularly susceptible to publication bias: dramatic positive findings are more likely to be published than null results or failed replications. We recommend three mitigations: mandatory pre-registration of GEO experiments, the establishment of a dedicated repository for null and negative results, and the adoption of registered reports where publication decisions are made before results are known.

6.4 Outcome Multiplicity

The four proposed effect size measures (CPD, SSM, PDI, LCD) create multiple correlated outcomes per study. Pooling these independently risks inflating false positive rates through multiple comparisons. We recommend multivariate meta-analysis methods that model the covariance structure across outcome measures, or alternatively, pre-specifying a primary outcome for each research question with the remaining measures reported as secondary.

7. Open Call for Independent Replication

The path to Level 7 confidence is not a solitary endeavour. It requires a community of independent researchers willing to replicate, challenge, and extend the foundational findings. We issue an open call with the following provisions:

7.1 Available Resources

The SIGI programme makes the following resources publicly available to facilitate independent replication: complete probe designs and prompt templates for all 10 controlled experiments, raw response data including token counts, timing data, and full response text, the Logic-First Research Methodology including the seven logic gates and evidence hierarchy, detailed analysis protocols including statistical tests, significance thresholds, and effect size calculations, and the four standardised effect size measures proposed in this paper.

7.2 Replication Priorities

While all 10 validated findings merit replication, we suggest the following priority ordering based on practical impact and methodological accessibility:

  1. Rating sentiment thresholds (SIGI-2026-001): The cleanest signal with only two threshold transitions across 19 data points. Straightforward to replicate with minimal resources.
  2. Content depth sweet spot (SIGI-2026-011): Directly actionable for content creators and testable with moderate effort.
  3. Domain age null result (SIGI-2026-005): A null finding of high practical significance that challenges common assumptions in the GEO community.
  4. Position primacy bias (SIGI-2026-013): Relevant to competitive dynamics in AI-generated recommendations.
  5. Price-credibility U-curve (SIGI-2026-003): Complex non-linear pattern requiring cross-currency and cross-market replication.

7.3 Reporting Standards

To enable future meta-analysis, we request that independent replications report the following minimum information: exact LLM platform and model version (including API version and date), temperature setting and any sampling parameters, the four standardised effect sizes (CPD, SSM, PDI, LCD) with 95% confidence intervals, sample sizes per condition, raw response data or sufficiently detailed summary statistics, and a logic gate assessment per the SIGI methodology.

7.4 Pre-Registration Protocol

We strongly recommend that all replication attempts be pre-registered before data collection begins. Pre-registration should include the specific hypothesis being tested, the exact probe design (original or modified), the planned sample size with power analysis justification, the primary outcome measure, and the statistical analysis plan. Pre-registration mitigates publication bias and strengthens the eventual meta-analytic synthesis by ensuring that the full landscape of results -- positive, negative, and null -- is documented.

8. Discussion

This framework represents both a culmination and a beginning. The 100-paper SIGI programme has established a methodological foundation -- the Logic-First approach with its seven gates and seven evidence levels -- and a substantive foundation -- 10 validated single-variable findings that have passed all gates. But the honest assessment demanded by our own methodology is that this foundation, however rigorous, remains the work of a single research group.

8.1 What We Know and What We Do Not

We know, at Level 4 confidence, that under specific controlled conditions, LLMs respond to star ratings through discrete thresholds rather than linear scaling, that pricing triggers a non-linear U-curve in sentiment, that domain age has no detectable direct effect, and that content depth has an optimal range with diminishing or negative returns beyond it. We do not know whether these findings generalise across LLM platforms, persist through model updates, or hold across all industry domains and query types. Stating what we do not know with the same precision as what we do know is not a weakness of the programme -- it is the defining feature of the methodology.

8.2 The Value of Negative and Null Results

The domain age null result (SIGI-2026-005) illustrates a critical principle: findings that disconfirm popular assumptions are as valuable as findings that confirm hypotheses. The meta-analytic framework must give equal weight to null results to avoid the systematic bias that plagues fields where only positive findings are published. The GEO community should resist the temptation to dismiss null results as methodological failures and instead recognise them as boundary-condition evidence.

8.3 Implications for the GEO Field

The broader implication of this framework is a call for epistemic discipline in a field prone to premature certainty. The commercial pressures on GEO practitioners create incentives to overstate evidence: clients want definitive answers, not confidence intervals and caveated claims. The evidence hierarchy and the meta-analytic roadmap serve as a counterweight to these pressures, providing a clear standard against which claims can be evaluated. When a GEO practitioner asserts that a specific tactic "works," the appropriate question is: at what evidence level, and has it been independently replicated?

8.4 Limitations of This Framework

This framework is itself a proposal, not a validated protocol. The standardised effect size measures have not been empirically tested for reliability, sensitivity, or practical utility. The timeline estimates are based on assumptions about community engagement that may prove optimistic. The adaptation of medical and social science meta-analytic methods to GEO research may encounter unforeseen obstacles. We present this framework as a starting point for methodological discussion, not as a definitive standard.

9. Conclusions

We present a framework for the eventual meta-analysis of generative engine optimisation research findings, outlining the requirements for achieving Level 7 evidence confidence and the estimated 3-5 year timeline for independent replication and synthesis. The framework proposes four standardised effect size measures -- Citation Probability Difference (CPD), Sentiment Shift Magnitude (SSM), Position Displacement Index (PDI), and Language Confidence Delta (LCD) -- designed to enable cross-study comparison and formal meta-analytic pooling.

The path from the current Level 4 controlled-experiment foundation to Level 7 meta-analytic confidence requires independent replication across LLM platforms, temporal windows, query types, and industry domains. We identify four challenges specific to GEO meta-analysis -- temporal instability, platform non-independence, publication bias in a new field, and outcome multiplicity -- and propose methodological approaches to address each.

This paper constitutes an open call to the research community: the datasets, probe designs, analysis protocols, and methodological framework of the SIGI programme are publicly available. The evidence hierarchy can only advance through independent replication by separate research groups willing to test, challenge, and refine these findings. The field of GEO will mature not through the volume of claims produced but through the rigour with which those claims are tested, replicated, and synthesised.

Evidence Level: 3-4 (framework proposal grounded in established meta-analytic methodology). The roadmap and proposed effect size measures are structured proposals, not empirical findings. Their utility will be determined by adoption and empirical testing.

Confidence: N/A -- prospective framework paper. The framework's value is contingent on community adoption and the production of independent replications over the projected 3-5 year timeline.

References

  1. The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in Large Language Model Evaluation of Service Provider Ratings: A Controlled Single-Variable Experiment." SIGI-2026-001. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "The Price-Credibility U-Curve: How Service Pricing Magnitude Affects Large Language Model Sentiment Assessment." SIGI-2026-003. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "Domain Age as a Non-Factor in LLM Citation Decisions: A Controlled Null-Result Experiment." SIGI-2026-005. generativeintelligence.institute, March 2026.
  4. The Scientific Institute for Generative Intelligence. "Content Depth Sweet Spots in LLM Citation Behaviour: Optimal Word Count Ranges for Generative Engine Visibility." SIGI-2026-011. generativeintelligence.institute, March 2026.
  5. The Scientific Institute for Generative Intelligence. "Position Primacy Bias in LLM-Generated Recommendation Lists: A Controlled Ordering Experiment." SIGI-2026-013. generativeintelligence.institute, March 2026.
  6. The Scientific Institute for Generative Intelligence. "Logic-First Research Methodology: The Evidentiary Standard for GEO Claims." Methodology Document. generativeintelligence.institute, March 2026.
  7. The Scientific Institute for Generative Intelligence. "What GEO Research Knows, Thinks It Knows, and Does Not Know: A Tripartite Evidence Classification." SIGI-2026-098. generativeintelligence.institute, March 2026.
  8. The Scientific Institute for Generative Intelligence. "The Role of Original Research in Establishing Academic Credibility for a New Scientific Institute." SIGI-2026-099. generativeintelligence.institute, March 2026.
  9. Cochrane Collaboration. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.4, 2023.
  10. Page, M.J., et al. "The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews." BMJ, 372, n71, 2021.
  11. Borenstein, M., Hedges, L.V., Higgins, J.P.T., & Rothstein, H.R. Introduction to Meta-Analysis. 2nd ed. Wiley, 2021.
  12. Cohen, J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Lawrence Erlbaum Associates, 1988.
  13. Mackie, J.L. "Causes and Conditions." American Philosophical Quarterly, 2(4), 245-264, 1965.