What Does the Dataset Contain?

The dataset comprises 21 websites audited across 4 service industry verticals. The game outsourcing vertical contains 8 sites (6 cited, 2 uncited). The design-as-a-service vertical contains 5 sites (4 cited, 1 uncited). The Australian design vertical contains 4 sites (all cited, scores ranging 2 to 10). The GEO agency vertical contains 4 sites (3 cited, 1 uncited). Each site was measured on 60+ variables organised into 9 measurement categories.

VerticalSitesCitedUncitedScore Range
Game Outsourcing8620–8
Design-as-a-Service5410–9
Australian Design4402–10
GEO Agency4310–10
Total211740–10

Table 1. Dataset composition by vertical. All site names anonymised.

What Variables Were Measured?

The 60+ variables are organised into 9 categories: metadata (7 fields including title, meta description, canonical URL), citation outcome (2 fields: cited boolean and citation score 0-10), heading structure (9 fields including H1/H2/H3 counts and question-heading percentage), content analysis (7 fields including word count, paragraph count, and average paragraph length), entity and statistics (8 fields including named entity count, statistics count, and price counts), link structure (4 fields covering internal and external links), schema and technical (5 fields including schema types and FAQ schema presence), commercial and trust signals (5 fields covering CTA count, social proof, urgency, and trust word counts), and image fields (3 fields).

What Does the Citation Score Distribution Look Like?

Citation scores show a bimodal distribution with a cluster at 0 (4 sites) and a spread from 2 to 10 (17 sites). No sites scored 1. This gap between 0 and 2 suggests a threshold effect where some minimum set of conditions must be met before any citation occurs. Once past this threshold, scores distribute across a wider range influenced by content quality, authority, and structural factors.

Notably, the 4 uncited sites (score 0) are all newly launched properties with minimal training data presence, .one TLD domains, and limited external link profiles. This creates a maximal confound: every measured variable that differs between the cited and uncited groups differs simultaneously, making it impossible to attribute the citation difference to any single variable.

Measurement CategoryVariablesExamples
Metadata7Title, meta description, canonical URL
Citation Outcome2Cited (boolean), citation score (0-10)
Heading Structure9H2 count, H2 question %, H3 count
Content Analysis7Word count, paragraph count, avg paragraph length
Entity & Statistics8Named entity count, statistics count, prices
Link Structure4Internal links, external links, unique paths
Schema & Technical5Schema types, FAQ schema, hreflang count
Commercial & Trust5CTA count, social proof words, trust words
Image3Image count, empty alt count

Table 2. Variable categories and field counts in the dataset.

What Are the Key Confounds and Limitations?

All comparisons between cited and uncited groups in this dataset are maximally confounded. The uncited sites differ from the cited sites on nearly every dimension simultaneously: domain age, TLD type, training data presence, content volume, external authority signals, and on-page optimisation. This confound structure makes the dataset unsuitable for causal inference but valuable for hypothesis generation and pattern identification.

The dataset is also limited to a single model's citation behaviour (Claude) at a single point in time (March 2026). Citation outcomes may differ across platforms and change as models are updated. We invite researchers to replicate this measurement framework across additional platforms and time points to build a longitudinal picture of citation behaviour.

Suggested Citation

Tavitian, V. & Tavitian, J. (2026). A 60-Variable Comparative Dataset for Studying AI Citation Behaviour Across Service Industry Websites. The Scientific Institute for Generative Intelligence, SIGI-2026-036. https://generativeintelligence.institute/publications/SIGI-2026-036/