SIGI-2026-077

Reclassifying GEO Findings: Applying the 7-Gate Framework to Separate Validated Findings from Hypotheses

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper applies the 7-gate Logic-First Research Methodology to the full corpus of GEO research findings produced within the SIGI research programme, reclassifying each finding into one of four evidence categories. Category A (Validated) contains findings from controlled probes that pass all 7 logic gates, comprising approximately 10 findings at Level 4 evidence. Category B (Strongly Supported) contains findings from observation experiments that pass most gates, comprising approximately 10 findings. Category C (Hypotheses) contains findings that fail Gate 2 (Confound Check) but remain plausible, comprising approximately 10 findings. Category D (Introspective) contains LLM self-report findings requiring behavioural confirmation, including all 77 trust signal rankings. The reclassification reveals that under rigorous application, the majority of GEO findings would be classified as hypotheses rather than conclusions. We specify the upgrade path for each category and argue that transparent classification, rather than uniform presentation of all findings as validated, better serves both the research community and practitioners making decisions based on GEO evidence.

Keywords

evidence classification, finding reclassification, 7-gate framework, validated findings, hypotheses, introspective evidence, GEO evidence base, research transparency, evidence hierarchy, confidence categories

1. Introduction

The GEO research programme has produced findings through four distinct methodological approaches: controlled single-variable probes, observational experiments with partial controls, competitive audit comparisons, and LLM introspective self-report. These approaches produce evidence of fundamentally different quality, yet findings from all four are sometimes presented with equivalent confidence language.

The 7-gate framework provides a systematic method for sorting findings by evidence quality. By applying all seven gates to each finding and documenting which gates pass and which fail, findings can be classified into categories that accurately reflect their evidentiary status. This reclassification does not diminish the value of lower-category findings -- hypotheses are essential for directing future research. It does, however, prevent the field from treating hypotheses as conclusions.

2. The Four Categories

2.1 Category A: Validated (All 7 Gates Passed)

Category A findings come exclusively from the controlled probe programme, where single variables were isolated across multiple variations. These findings pass all seven gates because the probe design ensures variable isolation (Gate 2), establishes counterfactuals through systematic variation (Gate 4), and uses directly manipulated inputs ensuring mechanism plausibility (Gate 7).

Table 1. Category A: Validated findings
FindingProbeVariationsConfidence
Rating threshold at 3.8 (negative to neutral)Ratings19HIGH
Rating threshold at 4.7 (neutral to positive)Ratings19HIGH
Domain age has zero effect on sentimentDomain Age12HIGH
Pricing U-curve with valley at $1K-$5KPricing18HIGH
Content depth sweet spot at 300-3,000 wordsWord Count9HIGH
Position primacy bias (first = advantage)Position6HIGH
Awards bell curve (2-7 = weakest zone)Awards15HIGH
Client count safe zone at 25-80Client Count19MODERATE
5,000+ words produces negative inversionWord Count9HIGH
Entity density: higher is marginally betterEntity Density7MODERATE

2.2 Category B: Strongly Supported (Most Gates Passed)

Category B findings come from observation experiments with partial controls and from LLM introspective analysis with behavioural confirmation. These findings pass most gates but typically lack full replication (Gate 6) or have partial confound issues (Gate 2).

Table 2. Category B: Strongly supported findings (selection)
FindingMethodMissing GateConfidence
Two-layer architecture (search does not equal synthesis)Observation experimentsGate 2 partialHIGH
Training data asymmetry (established over new)Controlled comparisonFullHIGH
The de-linking principle (LLMs strip URLs)Consistent observationFullHIGH
Self-ranking at first position triggers credibility discountIntrospection + behaviouralGate 6MODERATE
Paid placement bias (dual mechanism)Platform comparisonGate 6MODERATE-HIGH
Memory feedback loop is per-user onlySubscriber experimentsGate 6MODERATE

2.3 Category C: Hypotheses (Gate 2 Failed)

Category C findings come primarily from the competitive audit comparing new properties against established competitors. All of these findings fail Gate 2 because the compared groups differ on multiple variables simultaneously. The findings are plausible and worth testing, but the observational comparison does not validate them.

Table 3. Category C: Hypotheses requiring controlled testing
Claimed FindingActual StatusPrimary Confound
Question-format H2s associated with lower citationHYPOTHESISNew sites with question H2s also have zero authority
FAQ schema quantity inversely correlated with citationHYPOTHESISSites with FAQ schema are also newly launched
Internal link density negatively associatedHYPOTHESISSame multi-variable confound
Image count shows 30.7x correlation with citationHYPOTHESISEstablished sites have portfolios AND authority
Content volume without authority produces zero citationsHYPOTHESISSites are days old with zero external validation

2.4 Category D: Introspective (Requires Behavioural Confirmation)

Category D contains all findings derived from LLM self-report about its own citation behaviour. The scientific methodology is clear on this point: LLM self-reports about their own behaviour are informed hypotheses, not data. The 77 trust signal rankings, the 8-stage pipeline model, and all specific signal scores fall into this category. These findings are valuable for generating testable predictions and prioritising experiments, but they should not be treated as validated conclusions about actual LLM behaviour.

3. Upgrade Paths

Table 4. Evidence upgrade paths by category
From CategoryTo CategoryRequired Action
A (single-model)A (multi-model)Replicate across 3+ models, 3+ time points, 3+ domains
BAIsolate remaining confounds; replicate with controlled variation
CB or ADesign controlled experiment isolating the target variable
DBObtain behavioural confirmation through controlled observation

4. Discussion

The reclassification reveals a pyramid structure: a narrow base of validated findings (Category A), a moderate layer of strongly supported findings (Category B), and a broad top of hypotheses (Categories C and D). This structure is typical of young research fields and is not a weakness -- it reflects honest assessment of where the evidence stands.

The practical implication is that GEO practitioners should weight their strategies toward Category A findings (which can be acted on with confidence) while treating Category C and D findings as directional guidance worthy of testing rather than proven recommendations. The greatest risk in the current field is that all findings are presented and consumed with equivalent confidence, leading to misallocation of optimisation resources toward factors whose impact is hypothesised rather than demonstrated.

5. Conclusions

Applying the 7-gate framework to the full research corpus, we reclassify findings into validated, strongly supported, hypotheses, and introspective categories. Approximately 10 findings achieve full validation through controlled probe experiments. The majority of findings, including commonly cited GEO factors such as schema implementation effects and heading format preferences, are reclassified as hypotheses due to confounding. All LLM self-report findings are classified as introspective, requiring behavioural confirmation. This reclassification does not diminish the research programme's value but provides transparent guidance on which findings can support action and which require further investigation.

Confidence: HIGH for the reclassification framework itself. Individual reclassifications are subject to the analysis provided and may be revised as additional evidence becomes available.

References

  1. The Scientific Institute for Generative Intelligence. "The Claim Audit Checklist: A Pre-Publication Quality Gate for GEO Research." SIGI-2026-071. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "The Confound Matrix: A Practical Tool for Identifying Uncontrolled Variables in GEO Research." SIGI-2026-076. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "Cross-Model Replication Requirements for GEO Research." SIGI-2026-078. generativeintelligence.institute, March 2026.
  4. The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in LLM Evaluation of Service Provider Ratings." SIGI-2026-001. generativeintelligence.institute, March 2026.