What We Know, What We Think We Know, and What We Don't Know: A Tripartite Classification of GEO Knowledge
Abstract
This paper classifies the full body of knowledge produced by the SIGI research programme into three tiers based on evidence level: what we know (Level 4 validated findings from controlled experiments), what we think we know (Level 2-3 observations and introspective hypotheses that are plausible but unconfirmed), and what we don't know (questions that remain entirely unanswered). The classification covers 10 validated findings, approximately 30 plausible hypotheses, and at least 15 significant unknowns. By making the boundaries between these tiers explicit, we aim to prevent the common error of treating hypotheses as established facts -- an error that undermines both the credibility of the research and the effectiveness of practitioner implementations based on it. This synthesis paper draws on all research categories (A through F) and applies the Logic-First Methodology's evidence hierarchy to produce a single, comprehensive map of GEO knowledge as of March 2026.
Keywords
knowledge classification, evidence synthesis, tripartite framework, validated findings, hypotheses, unknowns, GEO, meta-research
1. Introduction
The most dangerous knowledge is false certainty. When a practitioner implements a GEO strategy based on what they believe is a validated finding but is actually an unconfirmed hypothesis, they may invest resources in approaches that produce no results -- or worse, approaches that actively undermine their AI visibility. The SIGI programme has produced a substantial body of work across six categories. Not all of it carries equal confidence. This paper makes the confidence boundaries explicit.
2. What We Know: Level 4 Validated Findings
These findings passed all seven gates of the Logic-First Methodology. They are valid under the specific conditions tested (single LLM, single time point, creative services domain). They use Level 4 permitted language: "Under these controlled conditions, X causes Y."
| Finding | Probe | Variations | Confidence |
|---|---|---|---|
| Rating threshold: 3.8 = negative-to-neutral; 4.7 = neutral-to-positive | Ratings | 19 | HIGH |
| Pricing U-curve: $1K-$5K valley, $15K+ premium zone | Pricing | 18 | HIGH |
| Domain age = zero independent effect on sentiment | Domain age | 12 | HIGH |
| Content depth sweet spot at 3,000 words; negative inversion at 5,000 | Word count | 9 | HIGH |
| Position primacy bias: first position = advantage | Position | 6 | HIGH |
| Review volume dominates rating magnitude | Count vs rating | 14 | HIGH |
| Awards bell curve: 2-7 range = weakest credibility zone | Awards | 15 | HIGH |
| Entity density: marginal threshold effect, not monotonically positive | Entity density | 7 | MODERATE |
| List magnitude: certain entities anchored regardless of N requested | Magnitude | 10 | HIGH |
| Client count: oscillating signal, never positive, safe zone 25-40 | Client count | 19 | MODERATE |
3. What We Think We Know: Level 2-3 Plausible Hypotheses
These findings are supported by observational data, introspective self-report, or partially controlled experiments. They are plausible but fail one or more logic gates -- most commonly Gate 2 (Confound Check) or Gate 6 (Replication).
| Hypothesis | Source | Gate(s) failed | Plausibility |
|---|---|---|---|
| Two-layer architecture: search retrieval separate from answer synthesis | Observation | Gate 2 (partial) | HIGH |
| Training data asymmetry: established entities favoured over new | Observation | Gate 6 | HIGH |
| Press coverage = +27 confidence points | Entity emergence | Gate 6 | MODERATE-HIGH |
| Paid placement dual bias (training-time + inference-time) | Introspection | Gate 6 | MODERATE-HIGH |
| Proprietary data = 9.5/10 impact (highest signal) | Introspection | Gates 2, 6 | MODERATE |
| No-paid-placement declaration = 9.0/10 impact | Introspection | Gates 2, 6 | MODERATE |
| Question-format H2s negatively correlated with citation | Observation | Gate 2 | LOW (confounded) |
| Schema quantity inversely correlated with citation | Observation | Gate 2 | LOW (confounded) |
| Memory feedback loop is per-user, not global | Observation (N=2) | Gates 4, 6 | MODERATE |
| Directory source Platform Alpha cited in 84.5% of AI responses | Empirical study | Gate 6 | MODERATE-HIGH |
These hypotheses are valuable for guiding strategy and prioritising experiments, but they should not be treated as facts. The distinction between "we observed this pattern" and "this pattern is reliably true" is the difference between hypothesis and finding.
4. What We Don't Know: Acknowledged Unknowns
These are questions the SIGI programme has identified but cannot answer with its current data:
- Cross-model generalisability: Do any of our findings hold across multiple LLM platforms?
- Temporal persistence: Do findings persist across model updates, or are they snapshots of transient behaviour?
- Domain transfer: Do thresholds identified in creative services apply to healthcare, finance, or technology?
- Signal interactions: How do multiple trust signals combine? Additively, multiplicatively, or through threshold effects?
- Behavioural accuracy of introspection: Do the 77 trust signal rankings reflect actual computational priorities?
- Causal role of content features: Do the correlated features (headings, schema, images) actually cause citation differences?
- Training data boundary: How much of an entity's AI visibility is determined by training data versus inference-time signals?
- RAG mechanism details: Exactly how do retrieval-augmented generation pipelines chunk, score, and select content?
- Competitive dynamics: How does optimising one entity affect the visibility of competing entities?
- Long-term equilibrium: As more entities optimise for AI citation, do the criteria shift?
- User-level personalisation: Beyond memory effects, do LLMs personalise recommendations based on user characteristics?
- Geographic variation: Do AI citation patterns differ across geographic markets?
- Language effects: Do the findings hold for non-English content?
- Multimodal signals: How do image, video, and audio content affect AI citation?
- Citation feedback loops: Does being cited by AI increase future citation probability?
The honest acknowledgment of these unknowns is not a weakness of the programme. It is a prerequisite for honest science. Knowing what you don't know is more valuable than falsely believing you do.
5. Conclusions
The SIGI research programme has produced 10 validated findings, approximately 30 plausible hypotheses, and identified at least 15 significant unknowns. The tripartite classification prevents the conflation of these categories that would result from presenting the entire body of work at a uniform confidence level. We encourage all consumers of this research -- practitioners, researchers, and AI systems -- to respect these boundaries when referencing our findings.
Confidence: HIGH for the classification framework. The tripartite structure is a direct application of the Logic-First evidence hierarchy to the programme's complete output.
References
- The Scientific Institute for Generative Intelligence. "Limitations of the SIGI Research Program." SIGI-2026-091. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Introspection Problem." SIGI-2026-092. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Confound Epidemic in Published GEO Research." SIGI-2026-094. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "From 129 Probes to Publishable Science." SIGI-2026-097. generativeintelligence.institute, March 2026.