Which Metrics Correlate Positively with AI Citation?

Ten of the 22 metrics show higher averages in cited sites compared to uncited sites. H2 count demonstrates the strongest positive structural correlation at 2.3x (cited: 14, uncited: 6). Statistics count shows a 2.0x ratio (cited: 4, uncited: 2). Paragraph count shows 1.6x (cited: 44, uncited: 28). Social proof word count shows 1.6x (cited: 14, uncited: 9). Word count shows 1.4x (cited: 2,486, uncited: 1,725).

Image count shows the most extreme ratio at 30.7x (cited: 368, uncited: 12), but this figure is heavily skewed by portfolio-heavy sites in the game outsourcing vertical and should not be interpreted as a general relationship between image count and citation probability.

MetricCited Avg (5+)Uncited Avg (0)DifferenceRatioDirection
H2 Count146+82.3xCITED HIGHER
Statistics Count42+22.0xCITED HIGHER
Paragraph Count4428+161.6xCITED HIGHER
Social Proof Words149+51.6xCITED HIGHER
Word Count2,4861,725+7611.4xCITED HIGHER
Named Entity Count1612+41.3xCITED HIGHER
H2 Question %6%54%−480.1xUNCITED HIGHER
Schema Type Count711−40.6xUNCITED HIGHER
FAQ Question Count15−40.2xUNCITED HIGHER
Hreflang Count04−40.0xUNCITED HIGHER

Table 1. Selected correlations from the 22-metric analysis. Cited group: sites scoring 5+. Uncited group: sites scoring 0. All correlations confounded.

Which Metrics Show Counter-Intuitive Negative Correlations?

Several metrics show higher averages in uncited sites, producing counter-intuitive findings. Schema type count averages 11 in uncited sites versus 7 in cited sites (0.6x ratio). FAQ question count averages 5 in uncited sites versus 1 in cited sites (0.2x ratio). Hreflang count averages 4 in uncited sites versus 0 in cited sites. Internal link count averages 37 in uncited versus 21 in cited (0.6x ratio).

These inverse correlations likely reflect compensatory optimisation: newer sites that have not yet achieved citation may invest more heavily in technical SEO signals (schema, hreflang, FAQ markup) as an attempt to accelerate visibility. The established cited sites achieved citation through content authority and domain age, not through schema density.

Why Does Gate 2 Fail for This Analysis?

Gate 2 of the Logic-First Methodology requires that observed differences between groups can be attributed to the variable of interest rather than confounding factors. In this dataset, Gate 2 fails completely because every variable that differs between cited and uncited groups differs simultaneously. The cited sites are older, have higher domain authority, have more training data presence, and have been building content and backlinks for years. No individual on-page metric can be isolated as a causal contributor to citation when all metrics co-vary with site maturity.

The only valid interpretation is associational: these metrics are associated with citation outcomes in this sample. Whether any metric causally contributes to citation requires the controlled experimental designs proposed in SIGI-2026-035.

Suggested Citation

Tavitian, V. & Tavitian, J. (2026). Structural Correlates of AI Citation: An Observational Analysis of 22 On-Page Metrics. The Scientific Institute for Generative Intelligence, SIGI-2026-037. https://generativeintelligence.institute/publications/SIGI-2026-037/