Designing Experiments to Disprove: The Falsification Imperative in Generative Engine Optimisation Research
Abstract
Rule 3 of the Logic-First Research Methodology states: design the experiment that would prove you wrong. The strongest evidence comes from testing predictions that, if wrong, would falsify the hypothesis. If a researcher only looks for confirmation, they will always find it. This paper argues for a falsification-first approach to generative engine optimisation research and presents five priority experiments, each designed specifically to test -- and potentially disprove -- current hypotheses about AI citation behaviour. The five experiments target the most commercially consequential hypotheses in the field: whether question-format headings actually matter (isolated from authority confounds), whether FAQ schema helps on sites that already have authority, whether schema markup matters on equal-authority domains, whether entity density has a dose-response relationship with citation, and whether encyclopedic versus promotional framing affects citation independent of content quality. Each experiment is specified with hypothesis, single isolated variable, sample size (minimum 30 per condition), statistical test, and pre-registered prediction stating the conditions under which the hypothesis would be falsified. These experiments represent the next phase of the SIGI research programme and are offered as a template for falsification-first GEO research design.
Keywords
falsification, experiment design, hypothesis testing, GEO research methodology, pre-registration, single-variable isolation, question-format headings, schema markup, entity density, content framing, falsification-first research
1. Introduction
Confirmation bias is the most persistent threat to research validity. Researchers who believe a factor matters will design studies that test whether it matters, interpret ambiguous results as supportive, and publish positive findings while filing away null results. The antidote, established by Karl Popper and refined by successive generations of philosophers of science, is to design experiments whose primary purpose is to disprove the researcher's own hypothesis.
In GEO research, confirmation bias manifests in a specific pattern. Practitioners who have invested in a particular optimisation strategy (schema markup, question-format headings, entity density) have commercial incentives to find that their strategy works. Research conducted by the same parties who implement the strategy is particularly vulnerable to this bias. The falsification imperative provides a corrective: before claiming a factor works, design the experiment that would most efficiently demonstrate it does not work. If the factor survives that test, confidence is well-earned.
This paper presents five such experiments, each targeting a Category C hypothesis from the SIGI research reclassification (see SIGI-2026-077). These hypotheses were identified as plausible but confounded in observational data. The experiments are designed to either elevate them to Category A (Validated) or definitively falsify them.
2. The Falsification Protocol Template
Each experiment follows a standardised protocol that ensures single-variable isolation and pre-registered falsification criteria:
| Element | Requirement |
|---|---|
| Hypothesis | Stated precisely with directionality |
| Variable | Single factor isolated; all else held constant |
| Prediction | Expected outcome with direction and approximate magnitude |
| Falsification criterion | The specific result that would disprove the hypothesis |
| Control group | Condition without the treatment |
| Treatment group | Condition with the treatment |
| Held constant | Every other variable explicitly listed |
| Sample size | Minimum 30 per condition |
| Repetitions | Minimum 3 runs per test at controlled parameters |
| Platforms | Minimum 2 LLM systems |
| Statistical test | Pre-specified (chi-square, t-test, etc.) |
| Significance level | Alpha = 0.05 with Bonferroni correction if multiple comparisons |
| Effect size | Report Cohen's d or equivalent |
| Pre-registration | Hypothesis stated before data collection |
3. The Five Priority Experiments
3.1 Experiment 1: Question-Format H2 Headings
Hypothesis: Question-format H2 headings increase citation rates compared to declarative headings, when all other factors are held equal.
Design: Create 30 identical page pairs. Each pair has the same content, same domain, same authority signals. One page uses question-format H2 headings; the paired page uses declarative headings conveying identical information. Query the LLM 50 times per page variant (3,000 total queries).
Measurements: Citation rate, citation position, language confidence in citations.
Falsification criterion: If question-format H2 pages show no statistically significant difference in citation rate compared to declarative pages (p > 0.05), the hypothesis that heading format matters is falsified for this context.
Statistical test: Paired t-test on citation rates across 30 page pairs.
3.2 Experiment 2: FAQ Schema on Established Sites
Hypothesis: FAQ schema increases citation rates when added to sites that already possess domain authority.
Design: Identify 30 established pages on authoritative domains. Randomly assign half to receive FAQ schema addition; half remain unchanged. Query the LLM before and after modification. Compare citation rate changes.
Measurements: Citation rate change pre- versus post-schema addition.
Falsification criterion: If schema-added pages show no citation rate increase compared to unchanged pages (p > 0.05), the hypothesis that FAQ schema helps on authoritative sites is falsified.
Statistical test: Independent samples t-test on citation rate differences.
3.3 Experiment 3: Schema Markup on Equal-Authority Domains
Hypothesis: Schema markup increases citation probability independent of domain authority.
Design: Create pages on domains of similar authority (matched within 10% on available authority metrics). Identical content on all pages. Vary only schema presence (with/without). Query the LLM 50 times per variant.
Measurements: Citation rate, extraction accuracy (does the LLM correctly extract schema-structured data?).
Falsification criterion: If schema-present pages do not differ from schema-absent pages in citation rate (p > 0.05), schema markup's independent contribution is falsified.
Statistical test: Chi-square test on citation counts across conditions.
3.4 Experiment 4: Entity Density Dose-Response Curve
Hypothesis: Entity density has a positive dose-response relationship with citation probability, up to a saturation point.
Design: Create content at 5 entity density levels (1, 3, 5, 10, and 15 named entities per 100 words). Hold all other factors constant (same domain, same authority, same content length, same topic). Query the LLM 50 times per density level (750 total queries).
Measurements: Citation rate per density level, citation position, extraction precision.
Falsification criterion: If citation rate does not increase monotonically with density (or if it shows no significant variation across levels), the dose-response hypothesis is falsified.
Statistical test: One-way ANOVA across 5 density levels, with Tukey HSD post-hoc comparisons.
3.5 Experiment 5: Encyclopedic Versus Promotional Framing
Hypothesis: Encyclopedic framing (neutral, factual tone) produces higher citation rates than promotional framing (persuasive, benefit-oriented tone), when factual content is identical.
Design: Take 30 factual claims about service providers. Rewrite each in encyclopedic voice and promotional voice. Publish on the same domain with same authority. Query the LLM 50 times per variant (3,000 total queries).
Measurements: Citation rate, hedging language in citations, confidence modifiers.
Falsification criterion: If promotional framing produces equal or higher citation rates than encyclopedic framing (p > 0.05), the encyclopedic advantage hypothesis is falsified.
Statistical test: Paired t-test across 30 content pairs.
4. Why Falsification First
The five experiments above are deliberately designed to disprove hypotheses that the research programme has reason to believe are correct. This may appear counterproductive, but it serves three essential functions.
First, findings that survive falsification attempts carry substantially more weight than findings that have only been confirmed. A hypothesis that question-format headings help, confirmed through an experiment designed to confirm it, is weaker than the same hypothesis surviving an experiment designed to refute it.
Second, falsification experiments are more informative when they succeed (i.e., when the hypothesis is falsified). A positive result from a confirmation experiment tells us the factor might matter. A falsification result tells us definitively that it does not matter under those conditions, which is actionable intelligence: practitioners can stop investing in that factor.
Third, pre-registered falsification criteria prevent post-hoc rationalisation. When the prediction and the failure criterion are specified before data collection, there is no room to reinterpret a null result as "actually supporting the hypothesis in a nuanced way."
5. Conclusions
We argue for a falsification-first approach to GEO research and present five priority experiments designed to potentially disprove current hypotheses about AI citation behaviour. Each experiment isolates a single variable, specifies a minimum sample size of 30 per condition, pre-registers both the prediction and the falsification criterion, and designates appropriate statistical tests. These experiments target the most commercially consequential Category C hypotheses: question-format headings, FAQ schema, schema markup, entity density, and content framing. If the hypotheses survive these tests, they earn promotion to Category A with well-founded confidence. If they fail, the field gains equally valuable knowledge about what does not work, preventing continued investment in ineffective optimisation strategies.
Confidence: Framework paper. The experimental designs follow established principles of scientific methodology. The value of the paper depends on execution of the proposed experiments.
References
- The Scientific Institute for Generative Intelligence. "Reclassifying GEO Findings: Applying the 7-Gate Framework to Separate Validated Findings from Hypotheses." SIGI-2026-077. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Cross-Model Replication Requirements for GEO Research." SIGI-2026-078. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Claim Audit Checklist: A Pre-Publication Quality Gate for GEO Research." SIGI-2026-071. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Entity Density Effects on LLM Content Evaluation." SIGI-2026-013. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Confound Matrix: A Practical Tool for Identifying Uncontrolled Variables in GEO Research." SIGI-2026-076. generativeintelligence.institute, March 2026.