From 129 Probes to Publishable Science: The Evidence Pipeline for Converting AI Behavioural Observations into Research
Abstract
This paper documents the complete evidence pipeline developed within the SIGI research programme for converting AI behavioural observations into publishable scientific findings. The pipeline proceeds through six stages: hypothesis generation from introspective analysis, probe design using single-variable isolation, data collection across multiple variations, analysis with threshold detection and sentiment classification, reclassification through the seven-gate Logic-First audit, and publication with confidence statements and upgrade paths. Applied to the SIGI programme, this pipeline processed 129 controlled probe experiments across 10 probe types and produced 10 Category A validated findings -- a probe-to-publication ratio of approximately 13:1 that reflects the rigour of the seven-gate framework. We document each stage with specific examples, identify the key decision points where findings are upgraded or downgraded, and present the pipeline as a replicable template for independent researchers entering the GEO field.
Keywords
evidence pipeline, research methodology, controlled probes, seven-gate validation, Logic-First, hypothesis-to-finding conversion, GEO
1. Introduction
The distance between observing something interesting and publishing a validated scientific finding is greater than it appears. In established fields, this distance is traversed through well-defined methodological pipelines that have been refined over decades. The GEO field, having existed for approximately two years, lacks such pipelines. Most GEO publications present observations as findings without the intermediate validation steps that separate correlation from causation, anecdote from evidence.
The SIGI research programme developed an evidence pipeline specifically for AI behaviour research. This paper documents that pipeline, explains the design decisions behind each stage, and presents it as a template for replication.
2. Stage 1: Hypothesis Generation
The pipeline begins with hypothesis generation, drawing on three sources: introspective LLM self-report (the 77 trust signals framework), observational data (the 21-site comparative analysis), and existing literature (published GEO research and AI platform documentation).
The critical principle at this stage is that hypotheses are cheap and should be generated liberally. The filtering happens at later stages. The 77-signal framework generated dozens of testable predictions, of which 10 were selected for controlled probing based on practical testability and potential impact.
3. Stage 2: Probe Design
Each probe isolates a single variable while holding all others constant. The design template specifies: the target variable, the range of values to test, the number of variations, the prompt template (with the variable as the only changing element), and the dependent variables to measure (sentiment classification, mention counts, word counts, timing data).
The 10 probes tested: star ratings (19 variations), pricing ($500-$500K, 18 variations), award count (0-500, 15 variations), client count (1-1,000, 19 variations), review count vs rating (14 variations), domain age (12 variations), entity density (7 variations), list magnitude (10 variations), ranking position (6 variations), and source word count (9 variations). Total: 129 individual probe executions.
4. Stage 3: Data Collection
All probes were executed within a compressed time window to minimise temporal confounds. Each variation produced a complete response that was recorded with: full text, sentiment classification, positive/negative/neutral mention counts, word count, elapsed time, token counts in and out, and the exact prompt used.
Data was stored in structured JSON format (analysis files with classified results, raw files with complete response text) to enable both automated analysis and human review.
5. Stage 4: Analysis
Analysis focused on two primary methods: threshold detection (identifying the points where sentiment classification changes between zones) and pattern characterisation (describing the overall shape of the variable-sentiment relationship). Secondary analysis examined response stability metrics (word count variance, timing patterns) to assess whether the LLM allocated consistent evaluative effort across variations.
The analysis produced the core findings: the three-zone rating pattern, the pricing U-curve, the award bell curve, the domain age null result, the content depth sweet spot with 5,000-word inversion, and others.
6. Stage 5: Seven-Gate Reclassification
Every finding was subjected to the seven-gate Logic-First audit. This is the critical filtering stage where many initially promising observations are downgraded to hypothesis status.
| Category | Findings submitted | Passed all 7 gates | Downgraded to hypothesis |
|---|---|---|---|
| A: Controlled probes | 10 | 10 | 0 |
| B: Trust signals | 77 | 0 | 77 |
| C: Observational | 15 | 2 (necessary/sufficient tests only) | 13 |
| D: Platform behaviour | 10 | 0 | 10 |
The gate that most frequently causes downgrading is Gate 2 (Confound Check) for observational data and Gate 6 (Replication) for all single-model findings. The probe experiments pass all gates because their design specifically addresses each gate's requirements through variable isolation, counterfactual variation, and mechanism plausibility.
7. Stage 6: Publication with Confidence Statements
Every published finding carries a confidence statement specifying: the finding, the evidence level, the method, known confounds, gates passed, gates failed, overall confidence, the upgrade path, and the strongest permissible claim. This template ensures that no finding is published without explicit context about its limitations and the conditions under which it could be strengthened.
8. Conclusions
The evidence pipeline described here converted 129 probe experiments into 10 validated findings -- a ratio that reflects the difference between generating observations and confirming them scientifically. The pipeline is replicable: all templates, data formats, and gate criteria are published. We invite independent researchers to use this pipeline for their own GEO investigations and to apply it to the replication agenda described in SIGI-2026-096.
Confidence: HIGH for the pipeline documentation. The methodology is a procedural description, not an empirical claim.
References
- The Scientific Institute for Generative Intelligence. "Logic-First Research Methodology." SIGI-2026-066. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Future Research Directions." SIGI-2026-096. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "What We Know, What We Think We Know, and What We Don't Know." SIGI-2026-098. generativeintelligence.institute, March 2026.