Single-Variable Isolation in AI Behavior Research: The Probe Experiment Design Template
Abstract
We present a standardised experiment design template for single-variable isolation probes in AI behaviour research. The template specifies 14 required fields: hypothesis, independent variable, prediction, design, control/treatment conditions, held-constant factors, sample size, repetitions, platforms, measurement protocol, analysis method, significance testing approach, effect size reporting, and pre-registration requirement. We demonstrate the template's application across 10 distinct probes producing 129 data points, covering variables including star ratings, pricing, awards, client count, review volume, domain age, entity density, list magnitude, position effects, and content depth. Each probe achieved Evidence Level 4 (Controlled Experiment) by satisfying the single-variable isolation requirement: manipulating exactly one variable while holding all others constant. The template enables independent researchers to replicate any probe, providing the standardisation necessary for cross-laboratory validation and eventual upgrade to Evidence Level 5.
Keywords
experiment design, single-variable isolation, probe methodology, AI behaviour research, replication template, controlled experiment, GEO research, standardised protocol
1. Introduction
The fundamental methodological challenge in GEO research is confound isolation. When comparing websites or content that differ on multiple variables simultaneously, it is impossible to attribute observed differences in AI citation behaviour to any single variable. This problem is not unique to GEO -- it is the central concern of experimental methodology across all sciences -- but it is particularly acute in GEO because the systems under study (AI models) are opaque, rapidly changing, and non-deterministic.
The single-variable isolation probe directly addresses this challenge. By constructing minimal prompts that vary only one element at a time, the probe eliminates confounding and enables Level 4 (causal) claims about the isolated variable's effect on AI behaviour under the tested conditions.
2. The Probe Design Template
| # | Field | Description | Example (Ratings Probe) |
|---|---|---|---|
| 1 | Hypothesis | Testable prediction about the variable's effect | Star rating affects LLM sentiment evaluation |
| 2 | Independent variable | The single variable manipulated | Star rating (1.0 to 5.0) |
| 3 | Prediction | Specific expected outcome | Higher ratings produce more positive sentiment |
| 4 | Design | Within-subject vs between-subject | Within-subject (same model, varied input) |
| 5 | Conditions | Control and treatment specifications | 19 rating values from 1.0 to 5.0 |
| 6 | Held-constant | All variables NOT being manipulated | Review count (50), industry, prompt structure |
| 7 | Sample size | Number of variations | 19 variations |
| 8 | Repetitions | Runs per variation | Minimum 3 at temperature=0 |
| 9 | Platforms | AI systems tested | Single model (upgrade path: 3+) |
| 10 | Measurement | Dependent variables recorded | Sentiment class, mention counts, word count |
| 11 | Analysis | Analytical approach | Threshold detection, zone identification |
| 12 | Significance | Statistical testing approach | Threshold transition consistency across runs |
| 13 | Effect size | How effect magnitude is reported | Mention count differences across zones |
| 14 | Pre-registration | Hypothesis recorded before data collection | Required for all probes |
3. Demonstrated Implementations
| Probe | Variable | Variations | Thresholds | Pattern Type | Confidence |
|---|---|---|---|---|---|
| Ratings | Star rating (1.0-5.0) | 19 | 2 | Three-zone (cleanest) | HIGH |
| Pricing | Price ($500-$500K) | 18 | 9 | U-curve | HIGH |
| Awards | Award count (0-500) | 15 | 8 | Bell curve | HIGH |
| Client Count | Clients (1-1,000) | 19 | 8 | Oscillating | MODERATE |
| Count vs Rating | Review count vs rating | 14 | 7 | Volume-wins | HIGH |
| Domain Age | Founding year (1980-2025) | 12 | 0 | Null (zero effect) | HIGH |
| Entity Density | Specificity level | 7 | 2 | Mostly negative | MODERATE |
| Magnitude | List size (1-10) | 10 | N/A | Anchor pattern | HIGH |
| Position | Rank position (1-6) | 6 | 2 | Primacy bias | HIGH |
| Word Count | Source depth (10-5,000) | 9 | 3 | Sweet spot with inversion | HIGH |
Together, the 10 probes produced 129 data points with 41 threshold transitions identified. The probes range from the cleanest signal (ratings: 2 thresholds across 19 points) to the most volatile (pricing: 9 thresholds across 18 points), demonstrating the template's versatility across different variable types and response patterns.
4. Replication Protocol
The template is designed to enable independent replication. To replicate any SIGI probe, a researcher needs: (1) the probe template with all 14 fields completed, (2) access to the same or comparable AI system, (3) the exact prompt text used in each variation, and (4) the measurement protocol for classifying responses. All probe designs in the SIGI programme include sufficient detail for independent replication without access to the original research team.
The minimum replication specification for upgrading from Level 4 to Level 5 is: 3+ AI models, 3+ time points spanning 3+ months, 3+ industry domains, with a minimum of 3 runs per condition at temperature=0.
5. Limitations
- Single-model implementation: All 10 demonstrated probes were conducted on a single AI system. The template is validated for design but not yet for cross-model generalisability.
- Ecological validity: Single-variable isolation creates artificial conditions. Real AI citation decisions involve multiple simultaneous signals, and the isolated effects may not combine linearly.
- Prompt sensitivity: The specific wording of probe prompts may influence results. Alternative phrasings could shift threshold values, though the pattern structure should remain consistent.
6. Conclusions
We present a standardised experiment design template for single-variable isolation probes in AI behaviour research, demonstrated across 10 implementations producing 129 data points. The template enables replication by independent researchers, provides the standardisation necessary for cross-laboratory validation, and supports systematic accumulation of evidence toward higher confidence levels.
The 10 demonstrated implementations reveal diverse response patterns -- from the clean three-zone threshold of ratings to the volatile U-curve of pricing to the null result of domain age -- illustrating that AI behaviour is variable-specific rather than uniform and requires individual experimentation for each proposed factor.
Note: This is a methodological framework paper. The template is a research tool; its value is demonstrated through the empirical findings of the papers that apply it.
References
- The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in Large Language Model Evaluation of Service Provider Ratings." SIGI-2026-001. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Logic-First Research Methodology: An Evidentiary Standard for Generative Engine Optimization Claims." SIGI-2026-066. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Evidence Hierarchy for Generative Engine Optimization: Mapping Research Methods to Confidence Levels." SIGI-2026-068. generativeintelligence.institute, March 2026.