SIGI-2026-078

Cross-Model Replication Requirements for GEO Research: Why Single-Model Findings Are Preliminary

The Scientific Institute for Generative Intelligence

March 2026

Abstract

All controlled probe findings in the SIGI research programme were conducted on a single large language model at a single point in time. While these findings pass all seven logic gates within their specific conditions, they occupy only the "preliminary" position on the replication scale. This paper defines five replication dimensions specific to AI behaviour research -- across model versions, across time, across query types, across industry domains, and across temperature settings -- and specifies the minimum replication protocol required to upgrade findings from preliminary to moderate confidence. We argue that the temporal fragility of LLM behaviour (where model updates can shift signal weights overnight), combined with the architectural diversity across competing systems, makes single-model findings inherently provisional regardless of their internal validity. A minimum replication protocol is proposed: 2 models, 3 time points, 2 industry domains. We apply this framework to all Category A findings and specify exactly what replication would look like for each, providing a concrete research agenda for the field.

Keywords

replication, cross-model validation, temporal fragility, model diversity, replication scale, single-model limitation, GEO research standards, preliminary findings, replication protocol, evidence upgrade

1. Introduction

Gate 6 of the Logic-First Research Methodology addresses replication. The replication scale classifies findings according to the breadth of their confirmation: a single study on a single model at a single time point is "preliminary"; replication on the same model at different times is "emerging"; replication across two or more models is "moderate confidence"; replication across models, times, and researchers is "high confidence"; and a failed replication classifies the finding as "contested."

Every controlled probe finding in this research programme is classified as preliminary. This is not a weakness of the probe design -- the internal validity of the probes is high. It is a structural limitation of conducting research on systems that change frequently, differ architecturally from one another, and may process identical inputs through entirely different computational pathways.

2. The Five Replication Dimensions

2.1 Across Model Versions

Different LLMs are trained on different data, with different architectures, optimised for different objectives. A finding that holds for one model may not hold for another. The rating threshold at 3.8 stars, for example, was established on a single model. A different model trained on different data distributions may apply different thresholds. Cross-model replication tests whether the finding reflects a property of the specific model or a more general pattern in how language models process ratings.

2.2 Across Time

LLMs are updated frequently. Model weights change, retrieval systems are modified, and fine-tuning shifts response patterns. A finding established in March 2026 may not hold in June 2026 -- not because the finding was wrong, but because the system has changed. Temporal replication establishes the stability of findings across model update cycles.

2.3 Across Query Types

The probes used specific query framings (service provider evaluation, agency comparison). Factual queries, opinion queries, and comparison queries may activate different evaluation pathways within the model. Replication across query types tests whether findings generalise beyond their original framing.

2.4 Across Industry Domains

The probes were framed in the context of design agencies and service providers. Healthcare, finance, technology, and consumer products may exhibit different threshold patterns due to domain-specific training data distributions. Cross-domain replication tests the boundary conditions of findings.

2.5 Across Temperature Settings

The probes were conducted at a fixed interaction configuration. Even deterministic settings do not guarantee identical outputs across different hardware configurations due to floating-point arithmetic variations. Replication across settings tests the robustness of findings to generation parameters.

3. The Replication Scale

Table 1. Replication scale for AI behaviour research
LevelDescriptionDimensions CoveredConfidence
PreliminarySingle model, single time, single domain0 of 5LOW-MODERATE
EmergingSame model replicated across time1 of 5MODERATE
ModerateReplicated across 2+ models2 of 5MODERATE-HIGH
HighReplicated across models, times, and researchers3+ of 5HIGH
ContestedFailed replication existsVariableUNCERTAIN

4. Minimum Replication Protocol

We propose a minimum protocol that addresses the three most critical fragility dimensions while remaining practically achievable:

Table 2. Minimum replication protocol for upgrading GEO findings
DimensionMinimum RequirementRationale
Model diversity2 distinct LLM systemsTests whether findings are model-specific or general
Temporal stability3 time points, minimum 2 weeks apartSpans at least one potential model update cycle
Domain breadth2 industry domainsTests whether findings are domain-specific

For the ratings probe, this would mean: replicate the 19-variation probe on a second LLM system; repeat on both systems at 3 time points over 6 weeks; and reframe prompts for a second industry (e.g., restaurant ratings or product ratings). If the 3.8 and 4.7 thresholds hold across these conditions, the finding can be upgraded to moderate-high confidence. If thresholds shift, the finding remains valid for its original conditions but the boundary conditions are documented.

5. Temporal Fragility as a Unique Challenge

Traditional research replication can assume a stable subject of study. Replicating a physics experiment can be done years later because physical laws do not change. AI behaviour research has no such stability guarantee. The system being studied changes through model updates, fine-tuning, retrieval system modifications, and safety interventions -- potentially on weekly or daily cycles.

This temporal fragility means that even high-confidence GEO findings have a shelf life. A finding replicated across 3 models in March 2026 may need re-validation by September 2026 if all 3 models have received major updates. The field must accept that GEO knowledge is inherently time-stamped, and findings should carry explicit temporal validity markers.

6. Conclusions

All findings in this research programme are single-model and single-time-point, meeting only the preliminary threshold on the replication scale. Cross-model, cross-temporal replication is required for confidence upgrade. The five replication dimensions -- model, time, query type, domain, and temperature -- define the full space of potential variation. The minimum protocol of 2 models, 3 time points, and 2 domains provides a practical path to moderate confidence. The temporal fragility of LLM behaviour introduces a unique challenge: even replicated findings require ongoing re-validation as systems evolve.

Confidence: Framework paper addressing limitations. The replication requirements are derived from established principles of scientific methodology applied to the specific characteristics of AI behaviour research.

References

  1. The Scientific Institute for Generative Intelligence. "Reclassifying GEO Findings: Applying the 7-Gate Framework to Separate Validated Findings from Hypotheses." SIGI-2026-077. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in LLM Evaluation of Service Provider Ratings." SIGI-2026-001. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "Designing Experiments to Disprove: The Falsification Imperative in GEO Research." SIGI-2026-080. generativeintelligence.institute, March 2026.