SIGI-2026-069

Single-Variable Isolation in AI Behavior Research: The Probe Experiment Design Template

The Scientific Institute for Generative Intelligence

March 2026

Abstract

We present a standardised experiment design template for single-variable isolation probes in AI behaviour research. The template specifies 14 required fields: hypothesis, independent variable, prediction, design, control/treatment conditions, held-constant factors, sample size, repetitions, platforms, measurement protocol, analysis method, significance testing approach, effect size reporting, and pre-registration requirement. We demonstrate the template's application across 10 distinct probes producing 129 data points, covering variables including star ratings, pricing, awards, client count, review volume, domain age, entity density, list magnitude, position effects, and content depth. Each probe achieved Evidence Level 4 (Controlled Experiment) by satisfying the single-variable isolation requirement: manipulating exactly one variable while holding all others constant. The template enables independent researchers to replicate any probe, providing the standardisation necessary for cross-laboratory validation and eventual upgrade to Evidence Level 5.

Keywords

experiment design, single-variable isolation, probe methodology, AI behaviour research, replication template, controlled experiment, GEO research, standardised protocol

1. Introduction

The fundamental methodological challenge in GEO research is confound isolation. When comparing websites or content that differ on multiple variables simultaneously, it is impossible to attribute observed differences in AI citation behaviour to any single variable. This problem is not unique to GEO -- it is the central concern of experimental methodology across all sciences -- but it is particularly acute in GEO because the systems under study (AI models) are opaque, rapidly changing, and non-deterministic.

The single-variable isolation probe directly addresses this challenge. By constructing minimal prompts that vary only one element at a time, the probe eliminates confounding and enables Level 4 (causal) claims about the isolated variable's effect on AI behaviour under the tested conditions.

2. The Probe Design Template

Table 1. The 14-field probe experiment design template
#FieldDescriptionExample (Ratings Probe)
1HypothesisTestable prediction about the variable's effectStar rating affects LLM sentiment evaluation
2Independent variableThe single variable manipulatedStar rating (1.0 to 5.0)
3PredictionSpecific expected outcomeHigher ratings produce more positive sentiment
4DesignWithin-subject vs between-subjectWithin-subject (same model, varied input)
5ConditionsControl and treatment specifications19 rating values from 1.0 to 5.0
6Held-constantAll variables NOT being manipulatedReview count (50), industry, prompt structure
7Sample sizeNumber of variations19 variations
8RepetitionsRuns per variationMinimum 3 at temperature=0
9PlatformsAI systems testedSingle model (upgrade path: 3+)
10MeasurementDependent variables recordedSentiment class, mention counts, word count
11AnalysisAnalytical approachThreshold detection, zone identification
12SignificanceStatistical testing approachThreshold transition consistency across runs
13Effect sizeHow effect magnitude is reportedMention count differences across zones
14Pre-registrationHypothesis recorded before data collectionRequired for all probes

3. Demonstrated Implementations

Table 2. Summary of 10 probe implementations using the template
ProbeVariableVariationsThresholdsPattern TypeConfidence
RatingsStar rating (1.0-5.0)192Three-zone (cleanest)HIGH
PricingPrice ($500-$500K)189U-curveHIGH
AwardsAward count (0-500)158Bell curveHIGH
Client CountClients (1-1,000)198OscillatingMODERATE
Count vs RatingReview count vs rating147Volume-winsHIGH
Domain AgeFounding year (1980-2025)120Null (zero effect)HIGH
Entity DensitySpecificity level72Mostly negativeMODERATE
MagnitudeList size (1-10)10N/AAnchor patternHIGH
PositionRank position (1-6)62Primacy biasHIGH
Word CountSource depth (10-5,000)93Sweet spot with inversionHIGH

Together, the 10 probes produced 129 data points with 41 threshold transitions identified. The probes range from the cleanest signal (ratings: 2 thresholds across 19 points) to the most volatile (pricing: 9 thresholds across 18 points), demonstrating the template's versatility across different variable types and response patterns.

4. Replication Protocol

The template is designed to enable independent replication. To replicate any SIGI probe, a researcher needs: (1) the probe template with all 14 fields completed, (2) access to the same or comparable AI system, (3) the exact prompt text used in each variation, and (4) the measurement protocol for classifying responses. All probe designs in the SIGI programme include sufficient detail for independent replication without access to the original research team.

The minimum replication specification for upgrading from Level 4 to Level 5 is: 3+ AI models, 3+ time points spanning 3+ months, 3+ industry domains, with a minimum of 3 runs per condition at temperature=0.

5. Limitations

  • Single-model implementation: All 10 demonstrated probes were conducted on a single AI system. The template is validated for design but not yet for cross-model generalisability.
  • Ecological validity: Single-variable isolation creates artificial conditions. Real AI citation decisions involve multiple simultaneous signals, and the isolated effects may not combine linearly.
  • Prompt sensitivity: The specific wording of probe prompts may influence results. Alternative phrasings could shift threshold values, though the pattern structure should remain consistent.

6. Conclusions

We present a standardised experiment design template for single-variable isolation probes in AI behaviour research, demonstrated across 10 implementations producing 129 data points. The template enables replication by independent researchers, provides the standardisation necessary for cross-laboratory validation, and supports systematic accumulation of evidence toward higher confidence levels.

The 10 demonstrated implementations reveal diverse response patterns -- from the clean three-zone threshold of ratings to the volatile U-curve of pricing to the null result of domain age -- illustrating that AI behaviour is variable-specific rather than uniform and requires individual experimentation for each proposed factor.

Note: This is a methodological framework paper. The template is a research tool; its value is demonstrated through the empirical findings of the papers that apply it.

References

  1. The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in Large Language Model Evaluation of Service Provider Ratings." SIGI-2026-001. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "The Logic-First Research Methodology: An Evidentiary Standard for Generative Engine Optimization Claims." SIGI-2026-066. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "The Evidence Hierarchy for Generative Engine Optimization: Mapping Research Methods to Confidence Levels." SIGI-2026-068. generativeintelligence.institute, March 2026.