SIGI-2026-046

Image Count Disparity and AI Citation: Deconstructing the 30.7x Ratio Observation

The Scientific Institute for Generative Intelligence

March 2026

Category C — Competitive Intelligence • Evidence Level 2 (Case Study)

Abstract

The correlation analysis of 21 service industry websites (SIGI-2026-037) identified image count as the most extreme ratio in the dataset: cited sites average 368 images versus 12 for uncited sites, a 30.7x disparity. This paper deconstructs this ratio to determine whether it represents a meaningful signal or a statistical artefact. We demonstrate that the ratio is driven almost entirely by a single outlier: one cited game outsourcing studio with 4,396 images (a portfolio-heavy homepage). Other cited sites range from 0 to 138 images, with a median of 27. Critically, both zero-image sites and high-image sites achieve citation in this sample: one site with zero homepage images scores 7/10, while a site with 4,396 images scores 8/10. Image count is therefore NOT predictive of citation: counterexamples exist in both directions. This finding serves as a methodological case study in why extreme correlations from small observational samples require scrutiny before being interpreted as actionable signals.

Keywords

image count, AI citation, outlier analysis, correlation artefact, visual content, portfolio sites, statistical methodology

1. Introduction

When analysing observational data for factors associated with AI citation, extreme ratios naturally attract attention. A 30.7x disparity in any metric between cited and uncited groups would, at face value, suggest a strong relationship. However, statistical ratios from small samples are highly susceptible to outlier effects, and the presence of extreme values can produce ratios that misrepresent the underlying distribution.

This paper uses the image count ratio as a case study in responsible interpretation of observational GEO data, demonstrating the analytical steps required to distinguish genuine signals from statistical artefacts.

2. Results

2.1 Image Count Distribution

Site (Anonymised)VerticalImage CountCitation Score
Outsource AlphaGame Outsourcing1388
Outsource BetaGame Outsourcing277
Outsource GammaGame Outsourcing155
Outsource DeltaGame Outsourcing226
Outsource EpsilonGame Outsourcing428
Outsource ZetaGame Outsourcing83
DaaS AlphaDaaS358
DaaS BetaDaaS187
DaaS GammaDaaS126
DaaS DeltaDaaS59
Design AlphaDesign Agency4,39610
Design BetaDesign Agency07
Design GammaDesign Agency455
Design DeltaDesign Agency82

2.2 Outlier Impact Analysis

CalculationMean Image CountMedian Image CountRatio vs Uncited
With outlier (Design Alpha)3682730.7x
Without outlier29222.4x

Removing a single site reduces the ratio from 30.7x to 2.4x — an 12.8-fold reduction. This demonstrates that 92% of the observed effect is attributable to one outlier.

2.3 Counterexample Analysis

The sufficiency and necessity tests both fail for image count as a predictor:

  • High images, cited: Design Alpha (4,396 images, score 10) — high images CAN coincide with citation
  • Zero images, cited: Design Beta (0 images, score 7) — images are NOT necessary for citation
  • Low images, uncited: Multiple uncited sites (8–15 images, score 0) — low images CAN coincide with non-citation

3. Discussion

The 30.7x ratio is a genuine mathematical calculation from the data, but it is not a meaningful signal for AI citation strategy. The ratio tells us that one cited site has an extremely image-heavy portfolio homepage, not that images drive citation.

This finding illustrates a broader principle: in small observational samples, a single outlier can dominate aggregate statistics and produce ratios that appear to indicate strong relationships where none exist. Responsible data analysis requires examining distributions rather than relying on ratios, and testing claims against counterexamples before accepting them as signals.

The image count metric is particularly susceptible to outlier effects because creative industry sites vary enormously in visual content density. A portfolio-focused game art studio naturally has thousands of images; a consulting firm may have none. This variation reflects business type, not citation strategy.

4. Limitations

  • Image count is a crude metric: It does not distinguish between portfolio images, decorative elements, icons, and informational graphics.
  • Homepage only: Image counts reflect homepage content, not total site image inventory.
  • Alt text quality not assessed: Image alt text quality, which may be more relevant to AI processing, was not included in this analysis.

5. Conclusions

The 30.7x image count ratio between cited and uncited sites is an artefact of portfolio-heavy outliers, not a predictive signal. Both zero-image and high-image sites achieve citation in this sample. Image count is neither necessary nor sufficient for AI citation. This finding serves as a methodological case study in why extreme correlations from small observational samples require scrutiny before being used to inform content strategy.

Confidence: LOW for any image-citation relationship. HIGH for the methodological demonstration of outlier effects on ratio-based analysis.

References

  1. The Scientific Institute for Generative Intelligence. “Structural Correlates of AI Citation: An Observational Analysis of 22 On-Page Metrics.” SIGI-2026-037. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. “A 60-Variable Comparative Dataset for Studying AI Citation Behavior Across Service Industry Websites.” SIGI-2026-036. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. “The Confound Problem in Observational GEO Research: Why 21-Site Comparisons Cannot Support Causal Claims.” SIGI-2026-050. generativeintelligence.institute, March 2026.