SIGI-2026-071

The Claim Audit Checklist: A Pre-Publication Quality Gate for GEO Research

The Scientific Institute for Generative Intelligence

March 2026

Abstract

Generative engine optimisation (GEO) is an emerging field characterised by rapid knowledge production and limited methodological consensus. The absence of standardised pre-publication quality checks has resulted in findings of highly variable rigour being presented with equivalent confidence language. This paper introduces the Claim Audit Checklist, a 10-item pre-publication quality gate derived from the Logic-First Research Methodology. Each checklist item maps directly to one of the seven logic gates (Logical Form, Confound Check, Necessary/Sufficient/Contributory, Counterfactual, Alternative Explanations, Replication, Mechanism) plus three additional verification steps: evidence level classification, language appropriateness verification, and next-step identification. We identify seven specific red flags that indicate a GEO finding should be downgraded, drawing on concrete examples from controlled probe experiments and observational studies conducted within this research programme. The checklist is designed as a practical tool that can be applied in under fifteen minutes per finding, preventing the publication of confounded observations as validated conclusions while preserving the capacity to publish appropriately qualified hypotheses.

Keywords

claim audit, pre-publication checklist, generative engine optimisation, research methodology, evidence hierarchy, logic gates, quality assurance, over-claiming, analytical errors, GEO research standards

1. Introduction

The field of generative engine optimisation is producing research findings at a pace that outstrips the development of methodological standards. Practitioners and researchers regularly publish claims about what factors influence AI citation behaviour, often without systematic verification of whether the evidence supports the strength of their conclusions. The resulting literature contains a mixture of rigorously controlled findings and confounded observations presented with identical confidence language.

This problem is not unique to GEO. The replication crisis in psychology and the broader social sciences demonstrated that fields lacking standardised pre-publication quality checks are vulnerable to over-claiming, where the strength of conclusions routinely exceeds what the evidence supports. In GEO, this risk is compounded by two factors: the novelty of the domain means there are few established benchmarks against which to calibrate claims, and the commercial applicability of findings creates incentives for premature certainty.

The Logic-First Research Methodology provides a seven-gate framework for evaluating research claims. However, applying the full framework requires substantial methodological expertise. What the field requires is a practical distillation: a checklist that can be applied quickly by researchers at any level of methodological training, catching the most common and most damaging analytical errors before publication.

This paper presents the Claim Audit Checklist, a 10-item pre-publication quality gate that operationalises the seven logic gates into a sequential verification process. We define each item, explain its connection to the underlying logic gate, illustrate common failure modes with examples from GEO research, and provide the specific red flags that should trigger immediate downgrading of a finding.

2. The 10-Item Claim Audit Checklist

The checklist is designed for sequential application. Each item acts as a binary gate: pass or fail. A finding that fails any single item cannot be published as validated without either addressing the identified weakness or explicitly downgrading the confidence level.

2.1 Item 1: Logical Form (Gate 1)

The first checklist item requires the researcher to write out the argument in formal propositional logic (P implies Q structure) and verify it matches a valid logical form. The most common failure in GEO research is affirming the consequent: observing that a cited site has a particular feature and concluding that the feature caused the citation. A site may be cited for any number of reasons; the presence of a feature on a cited site does not establish that the feature caused the citation.

2.2 Item 2: Confound Isolation (Gate 2)

The second item requires listing every variable that differs between compared groups. If more than one variable differs, no causal conclusion about any single variable is possible. In GEO research, this gate is frequently violated when comparing new sites against established competitors, where site age, domain authority, review count, training data presence, content maturity, backlink profile, and brand recognition all differ simultaneously.

2.3 Item 3: Relationship Type (Gate 3)

The third item requires specifying whether the claimed relationship is one of necessity, sufficiency, or contribution. Most GEO factors are contributory at best, meaning they increase the probability of an outcome but are neither required for it nor guarantee it. Stating a contributory factor as necessary or sufficient is a form of over-claiming.

2.4 Item 4: Counterfactual Established (Gate 4)

The fourth item asks whether an alternative outcome has been established or approximated. If the finding is based solely on observing what did happen, without establishing what would have happened under different conditions, the finding fails this gate. Controlled probe experiments establish counterfactuals through systematic variation; observational studies typically cannot.

2.5 Item 5: Alternative Explanations (Gate 5)

The fifth item requires generating a minimum of three alternative explanations for the observed result and assessing whether each has been eliminated. If fewer than three alternatives are ruled out, the finding remains ambiguous. In GEO research, alternative explanations commonly include confounds with site age, content quality independent of the tested variable, and selection bias in which sites were included in the analysis.

2.6 Item 6: Replication Status (Gate 6)

The sixth item documents the replication status of the finding. A single study on a single model at a single time point is classified as preliminary. Replication across the same model at different times qualifies as emerging. Replication across two or more models qualifies as moderate confidence. Failed replication classifies the finding as contested.

2.7 Item 7: Mechanism Plausibility (Gate 7)

The seventh item traces the causal pathway from the claimed signal to the citation decision and verifies that the LLM has access to the signal at the point where the decision is made. If a signal is not present in the model's input at inference time, it cannot be causal regardless of any observed correlation. Site speed, for example, is not perceptible to an LLM during response generation, making it an implausible direct signal.

2.8 Item 8: Evidence Level Classification

The eighth item requires classifying the finding on the seven-level evidence hierarchy, from anecdote (Level 1) through meta-analysis (Level 7). This classification determines what language is permitted in stating the conclusion.

2.9 Item 9: Language Appropriateness

The ninth item verifies that the conclusion is stated using only language permitted at the classified evidence level. Levels 1-2 permit observational language only. Level 3 permits correlation language. Level 4 permits causal language scoped to specific conditions. Levels 5 and above permit generalised causal language. Using causal language for Level 3 evidence, or generalised language for Level 4 evidence, constitutes over-claiming.

2.10 Item 10: Next Step Identification

The tenth item requires identifying the specific experiment that would upgrade the finding to the next evidence level. This serves two purposes: it demonstrates that the researcher understands the limitations of the current evidence, and it provides a concrete path forward for the field. A finding published without an identified upgrade path implies completeness that is rarely warranted.

3. The Seven Red Flags

Through application of the checklist to the full corpus of GEO research conducted within this programme, seven recurring red flags were identified. Each red flag maps to a specific gate violation and should trigger immediate reconsideration of the finding's classification.

Table 1. Seven red flags for GEO research claims
Red FlagDescriptionGate ViolatedExample
1Comparing groups that differ on multiple dimensions simultaneouslyGate 2 (Confound)New site vs established competitor comparison where age, authority, reviews, and training data all differ
2Using LLM self-report as evidence of mechanismGate 7 (Mechanism)Treating an LLM's explanation of its own citation behaviour as ground truth rather than informed hypothesis
3Treating a single observation as proof of a probabilistic claimGate 4 (Counterfactual)One site with feature X not being cited taken as evidence that X is harmful
4Claiming causation from correlationGate 1 (Logical Form)Observing that cited sites have more images and concluding images cause citation
5Citing extreme ratios from small uncontrolled samplesGate 2 (Confound)Reporting a 30.7x correlation from a two-group comparison with seven confounding variables
6Stating null results from confounded comparisonsGate 2 (Confound)Concluding a factor has zero effect when the comparison involves multiple simultaneous differences
7Using prohibited language for the evidence levelItems 8-9 (Level/Language)Using words like "proves" or "demonstrates" for Level 2-3 observational findings

4. Application to GEO Research

To illustrate the practical impact of the checklist, we apply it to two categories of findings from the SIGI research programme.

4.1 Validated Finding: Rating Threshold at 3.8

The finding that a star rating of 3.8 triggers the transition from negative to neutral LLM sentiment passes all 10 checklist items. The logical form is valid (modus ponens). The variable is isolated (only rating changes; review count held at 50). The relationship is stated as contributory, not necessary or sufficient. The counterfactual is established through 19 systematic variations. Three alternative explanations are addressed. The replication status is documented as preliminary (single model). The mechanism is plausible (the rating is directly in the prompt). The evidence level is correctly classified as Level 4. The language uses appropriate scoping. The upgrade path is specified (cross-model replication).

4.2 Hypothesis Requiring Downgrade: Question-Format Headings

The observation that question-format H2 headings correlate with lower citation rates fails the checklist at Item 2. The sites with question-format headings also differ from comparison sites on age, domain authority, training data presence, review count, backlink profile, content maturity, and brand recognition. With seven or more confounding variables, no causal conclusion about heading format is possible. The checklist forces this finding to be reclassified from a validated conclusion to a hypothesis warranting controlled experimentation.

5. Discussion

The Claim Audit Checklist addresses a specific failure mode in emerging research fields: the tendency to publish findings at confidence levels exceeding what the evidence supports. In GEO research specifically, the commercial applicability of findings creates pressure to present hypotheses as conclusions, and the novelty of the field means reviewers may lack the domain expertise to identify over-claiming.

The checklist is deliberately conservative. It is designed to catch false positives (findings incorrectly classified as validated) at the cost of occasionally requiring additional evidence for findings that are, in fact, correct. This asymmetry is intentional: in a young field, the cost of acting on false findings exceeds the cost of delayed action on true ones.

The 10-item structure balances thoroughness with practicality. Application to a single finding typically requires under fifteen minutes once the researcher is familiar with the framework. The sequential design means that findings failing early items (particularly Items 1 and 2) can be efficiently reclassified without completing the full checklist.

6. Conclusions

We present a 10-item pre-publication checklist designed to prevent over-claiming in GEO research. The checklist maps to the seven logic gates of the Logic-First Research Methodology and adds three verification steps for evidence level, language appropriateness, and next-step identification. Seven specific red flags are documented with concrete examples from the GEO research domain. Application of the checklist to the SIGI research corpus demonstrates its capacity to correctly separate validated findings (which pass all items) from hypotheses requiring further investigation (which fail at identifiable gates).

The checklist is offered as a practical tool for GEO researchers and is designed for adoption without requiring extensive methodological training. Its primary contribution is operational: converting an abstract methodological framework into a concrete, repeatable quality assurance process.

Confidence: Framework paper. The checklist itself is a methodological tool whose validity is demonstrated through application to known findings. Its utility for the broader field depends on adoption and iterative refinement.

References

  1. The Scientific Institute for Generative Intelligence. "The Confound Matrix: A Practical Tool for Identifying Uncontrolled Variables in GEO Research." SIGI-2026-076. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "Red Flags in GEO Research: Common Analytical Errors and How to Detect Them." SIGI-2026-079. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "Reclassifying GEO Findings: Applying the 7-Gate Framework to Separate Validated Findings from Hypotheses." SIGI-2026-077. generativeintelligence.institute, March 2026.
  4. The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in LLM Evaluation of Service Provider Ratings." SIGI-2026-001. generativeintelligence.institute, March 2026.