SIGI-2026-093

Replication Challenges in AI Behaviour Research: Why Findings May Not Hold Across Models, Time, or Domains

The Scientific Institute for Generative Intelligence

March 2026

Abstract

Replication is the foundation of scientific confidence. In traditional experimental sciences, findings gain credibility through independent reproduction under similar conditions. AI behaviour research faces five distinct barriers that make replication fundamentally more challenging than in established fields: model version differences that alter behaviour between testing sessions, temporal instability from continuous model updates on 6-12 week cycles, query type sensitivity where factual, opinion, and comparison queries may invoke different processing pathways, domain specificity where findings in one industry vertical may not transfer to another, and temperature non-determinism where even purportedly deterministic settings produce variable outputs. This paper analyses each barrier, quantifies its impact where possible using data from the SIGI research programme's 129 controlled probes, and proposes the establishment of a GEO Findings Registry -- a standardised database tracking replication status across models, time points, and domains. We argue that without systematic replication infrastructure, the GEO field risks accumulating a body of findings that appears scientific but lacks the confirmation necessary for reliable practical application.

Keywords

replication, AI behaviour, model versioning, temporal instability, GEO research, findings registry, non-determinism, cross-model validation

1. Introduction

The replication crisis in psychology and biomedical research has demonstrated that even well-established fields can accumulate bodies of knowledge that partially collapse under systematic replication attempts. AI behaviour research -- and generative engine optimisation research specifically -- faces structural conditions that make replication even more challenging than in these already-troubled fields.

In traditional experimental sciences, the phenomenon under study does not receive software updates. A chemical reaction proceeds the same way regardless of when it is tested. A psychological effect, while potentially variable across cultures and time periods, does not change because the human brain received a firmware update. AI behaviour is fundamentally different: the system under study is actively modified by its creators, often without public notice, on cycles as short as six weeks.

This paper systematically examines each barrier to replication, draws on data from the SIGI research programme to illustrate the challenges, and proposes institutional infrastructure to track replication status across the emerging GEO research field.

2. Barrier 1: Model Version Differences

The SIGI programme conducted controlled probes on a single LLM and introspective research on a different LLM. Even this minimal degree of cross-model comparison reveals significant differences in response patterns. Different models may apply different threshold values for the same variables, weight trust signals differently, and employ structurally different retrieval-augmented generation architectures.

A threshold identified at 3.8 stars on one model could occur at 3.5 or 4.2 on another. The U-curve pricing pattern observed in our probes may manifest as a different shape entirely on a model with different training data composition. These are not marginal variations -- they could represent qualitatively different evaluation strategies.

Furthermore, apparent replication success across models could be misleading. If two models were trained on substantially overlapping corpora, they may have learned similar biases. A finding that replicates across such models would appear confirmed while actually reflecting shared data artefacts rather than genuine properties of LLM evaluation.

3. Barrier 2: Temporal Instability

Major LLM providers update their models on cycles ranging from 6 to 12 weeks. These updates can modify model weights, alter retrieval pipelines, adjust safety filtering, and change response generation patterns. A finding validated in March 2026 may no longer hold by May 2026.

The SIGI programme's entire probe dataset was collected within a narrow window in March 2026. This temporal concentration is both a strength (minimising within-study temporal confounds) and a limitation (providing no information about temporal stability). We do not know whether the rating threshold at 3.8 stars was present before our testing window or will persist after it.

The practical implication is severe: GEO practitioners implementing strategies based on research findings may be optimising for model behaviour that has already changed by the time their implementations are complete.

4. Barrier 3: Query Type Sensitivity

The SIGI programme's controlled probes used specific query framings within the service provider evaluation domain. Different query types -- factual queries, opinion queries, comparison queries, transactional queries -- may invoke different processing pathways within the same model. Our findings about how LLMs evaluate star ratings may apply only to evaluative queries and may not generalise to informational or transactional contexts.

Data from the SIGI research programme's framing analysis (Category D) demonstrates that changing a single word in a query can alter which entities are surfaced, which source types are prioritised, and how evaluation criteria are weighted. This sensitivity suggests that findings established under one query framing may not transfer to another, even within the same model and time period.

5. Barrier 4: Domain Specificity

All SIGI programme data comes from creative service industries. The degree to which findings generalise to healthcare, finance, technology, legal services, or other domains is entirely unknown. LLMs may process trust signals differently across domains -- a methodology section may carry more weight in medical research contexts than in creative services, and star ratings may have different threshold patterns for restaurants than for design agencies.

Domain specificity is particularly concerning because GEO research findings are often presented as universal principles. Claims such as "content depth of 3,000 words is optimal" may be accurate for the tested domain but misleading if applied to medical information (where brevity and precision may be more valued) or legal content (where exhaustive treatment may be expected).

6. Barrier 5: Temperature Non-Determinism

Even when LLMs are configured with temperature=0, outputs are not strictly deterministic. GPU floating-point operations can introduce micro-variations that propagate through the generation process, producing different outputs from identical inputs. The SIGI programme mitigated this through the probe methodology (testing multiple variations to establish patterns rather than relying on single outputs), but the fundamental non-determinism means that exact replication of individual responses is not possible.

This barrier is less severe than the others -- it affects individual responses rather than aggregate patterns -- but it complicates replication at the level of specific outputs and requires statistical approaches rather than exact matching.

7. A Replication Status Framework

Table 1. Proposed replication status levels for GEO findings
StatusDefinitionRequirements
PreliminarySingle study, single model, single time pointPublished methodology and data
EmergingReplicated on same model at different timeMinimum 30-day gap between tests
ModerateReplicated across 2+ modelsDifferent model families (not just versions)
StrongReplicated across models, times, and researchersOriginal research team confirmation
ContestedFailed replication existsPublished non-replication with methodology

Under this framework, all SIGI programme findings currently hold "Preliminary" status. This is not a weakness unique to our programme -- it is the status that applies to essentially all published GEO research as of March 2026.

8. Proposal: The GEO Findings Registry

We propose the establishment of a GEO Findings Registry -- a public, standardised database that tracks every published GEO finding alongside its replication metadata. Each entry would record: the finding, the model tested, the exact model version identifier, the date of testing, the methodology, the replication status level, links to any replication or non-replication attempts, and the domain in which the finding was established.

Such a registry would serve three purposes. First, it would allow practitioners to quickly assess which findings have been validated across multiple conditions. Second, it would incentivise replication by making replication studies discoverable and citable. Third, it would make the fragility of single-study findings visible, counteracting the tendency to treat preliminary results as established facts.

9. Conclusions

We identify five distinct barriers to replication in AI behaviour research and propose a findings registry to track replication status across the GEO field. All SIGI programme findings, along with essentially all published GEO research, currently hold "Preliminary" replication status. Upgrading findings to higher confidence levels requires systematic cross-model, cross-temporal, and cross-domain replication that the field has not yet attempted at scale.

The field of GEO must choose between two paths: one in which findings accumulate without replication, creating an illusion of knowledge that may collapse under scrutiny; and one in which replication infrastructure is built alongside primary research, producing a smaller but more reliable body of knowledge. We advocate for the latter.

Confidence: HIGH for the barrier identification and framework proposal. The five barriers are logically derived and empirically grounded in the challenges encountered during the SIGI programme.

References

  1. The Scientific Institute for Generative Intelligence. "Limitations of the SIGI Research Program." SIGI-2026-091. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "Future Research Directions: A Priority-Ranked Agenda." SIGI-2026-096. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "From 129 Probes to Publishable Science." SIGI-2026-097. generativeintelligence.institute, March 2026.