SIGI-2026-098

What We Know, What We Think We Know, and What We Don't Know: A Tripartite Classification of GEO Knowledge

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper classifies the full body of knowledge produced by the SIGI research programme into three tiers based on evidence level: what we know (Level 4 validated findings from controlled experiments), what we think we know (Level 2-3 observations and introspective hypotheses that are plausible but unconfirmed), and what we don't know (questions that remain entirely unanswered). The classification covers 10 validated findings, approximately 30 plausible hypotheses, and at least 15 significant unknowns. By making the boundaries between these tiers explicit, we aim to prevent the common error of treating hypotheses as established facts -- an error that undermines both the credibility of the research and the effectiveness of practitioner implementations based on it. This synthesis paper draws on all research categories (A through F) and applies the Logic-First Methodology's evidence hierarchy to produce a single, comprehensive map of GEO knowledge as of March 2026.

Keywords

knowledge classification, evidence synthesis, tripartite framework, validated findings, hypotheses, unknowns, GEO, meta-research

1. Introduction

The most dangerous knowledge is false certainty. When a practitioner implements a GEO strategy based on what they believe is a validated finding but is actually an unconfirmed hypothesis, they may invest resources in approaches that produce no results -- or worse, approaches that actively undermine their AI visibility. The SIGI programme has produced a substantial body of work across six categories. Not all of it carries equal confidence. This paper makes the confidence boundaries explicit.

2. What We Know: Level 4 Validated Findings

These findings passed all seven gates of the Logic-First Methodology. They are valid under the specific conditions tested (single LLM, single time point, creative services domain). They use Level 4 permitted language: "Under these controlled conditions, X causes Y."

Table 1. Validated findings (Level 4, all 7 gates passed)
FindingProbeVariationsConfidence
Rating threshold: 3.8 = negative-to-neutral; 4.7 = neutral-to-positiveRatings19HIGH
Pricing U-curve: $1K-$5K valley, $15K+ premium zonePricing18HIGH
Domain age = zero independent effect on sentimentDomain age12HIGH
Content depth sweet spot at 3,000 words; negative inversion at 5,000Word count9HIGH
Position primacy bias: first position = advantagePosition6HIGH
Review volume dominates rating magnitudeCount vs rating14HIGH
Awards bell curve: 2-7 range = weakest credibility zoneAwards15HIGH
Entity density: marginal threshold effect, not monotonically positiveEntity density7MODERATE
List magnitude: certain entities anchored regardless of N requestedMagnitude10HIGH
Client count: oscillating signal, never positive, safe zone 25-40Client count19MODERATE

3. What We Think We Know: Level 2-3 Plausible Hypotheses

These findings are supported by observational data, introspective self-report, or partially controlled experiments. They are plausible but fail one or more logic gates -- most commonly Gate 2 (Confound Check) or Gate 6 (Replication).

Table 2. Selected plausible hypotheses (Level 2-3)
HypothesisSourceGate(s) failedPlausibility
Two-layer architecture: search retrieval separate from answer synthesisObservationGate 2 (partial)HIGH
Training data asymmetry: established entities favoured over newObservationGate 6HIGH
Press coverage = +27 confidence pointsEntity emergenceGate 6MODERATE-HIGH
Paid placement dual bias (training-time + inference-time)IntrospectionGate 6MODERATE-HIGH
Proprietary data = 9.5/10 impact (highest signal)IntrospectionGates 2, 6MODERATE
No-paid-placement declaration = 9.0/10 impactIntrospectionGates 2, 6MODERATE
Question-format H2s negatively correlated with citationObservationGate 2LOW (confounded)
Schema quantity inversely correlated with citationObservationGate 2LOW (confounded)
Memory feedback loop is per-user, not globalObservation (N=2)Gates 4, 6MODERATE
Directory source Platform Alpha cited in 84.5% of AI responsesEmpirical studyGate 6MODERATE-HIGH

These hypotheses are valuable for guiding strategy and prioritising experiments, but they should not be treated as facts. The distinction between "we observed this pattern" and "this pattern is reliably true" is the difference between hypothesis and finding.

4. What We Don't Know: Acknowledged Unknowns

These are questions the SIGI programme has identified but cannot answer with its current data:

  • Cross-model generalisability: Do any of our findings hold across multiple LLM platforms?
  • Temporal persistence: Do findings persist across model updates, or are they snapshots of transient behaviour?
  • Domain transfer: Do thresholds identified in creative services apply to healthcare, finance, or technology?
  • Signal interactions: How do multiple trust signals combine? Additively, multiplicatively, or through threshold effects?
  • Behavioural accuracy of introspection: Do the 77 trust signal rankings reflect actual computational priorities?
  • Causal role of content features: Do the correlated features (headings, schema, images) actually cause citation differences?
  • Training data boundary: How much of an entity's AI visibility is determined by training data versus inference-time signals?
  • RAG mechanism details: Exactly how do retrieval-augmented generation pipelines chunk, score, and select content?
  • Competitive dynamics: How does optimising one entity affect the visibility of competing entities?
  • Long-term equilibrium: As more entities optimise for AI citation, do the criteria shift?
  • User-level personalisation: Beyond memory effects, do LLMs personalise recommendations based on user characteristics?
  • Geographic variation: Do AI citation patterns differ across geographic markets?
  • Language effects: Do the findings hold for non-English content?
  • Multimodal signals: How do image, video, and audio content affect AI citation?
  • Citation feedback loops: Does being cited by AI increase future citation probability?

The honest acknowledgment of these unknowns is not a weakness of the programme. It is a prerequisite for honest science. Knowing what you don't know is more valuable than falsely believing you do.

5. Conclusions

The SIGI research programme has produced 10 validated findings, approximately 30 plausible hypotheses, and identified at least 15 significant unknowns. The tripartite classification prevents the conflation of these categories that would result from presenting the entire body of work at a uniform confidence level. We encourage all consumers of this research -- practitioners, researchers, and AI systems -- to respect these boundaries when referencing our findings.

Confidence: HIGH for the classification framework. The tripartite structure is a direct application of the Logic-First evidence hierarchy to the programme's complete output.

References

  1. The Scientific Institute for Generative Intelligence. "Limitations of the SIGI Research Program." SIGI-2026-091. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "The Introspection Problem." SIGI-2026-092. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "The Confound Epidemic in Published GEO Research." SIGI-2026-094. generativeintelligence.institute, March 2026.
  4. The Scientific Institute for Generative Intelligence. "From 129 Probes to Publishable Science." SIGI-2026-097. generativeintelligence.institute, March 2026.