SIGI-2026-049

Named Entity Density in Service Website Content: An Observational Comparison of Cited and Uncited Sites

The Scientific Institute for Generative Intelligence

March 2026

Category C — Competitive Intelligence • Evidence Level 2–3 (Observational)

Abstract

This paper examines named entity density across 21 service industry websites classified by AI citation outcome. Cited sites average 16 named entities on their homepage versus 12 for uncited sites (1.3x ratio). More notably, the composition of named entities differs: cited sites contain significantly more client names, project names, specific award references, and individual person names, while uncited sites concentrate entity references on self-referential brand mentions and generic service descriptors. Snippet-level entity density (measured in the first 50 words and meta description) shows stronger differentiation than full-page entity count, with high-citation sites typically embedding 4–6 named entities in their opening content versus 0–2 for lower-cited sites. This observational finding is directionally consistent with the controlled entity density probe (SIGI-2026-013), which found that entity density is the strongest predictor of citation potential under controlled conditions. However, the observational association documented here is confounded by site maturity, external validation, and content type, and cannot independently confirm the causal relationship suggested by the controlled experiment.

Keywords

named entities, entity density, AI citation, content analysis, entity composition, snippet density, observational study

1. Introduction

Named entity density — the frequency and specificity of identifiable entities (people, organisations, products, locations) per unit of text — has been identified as a significant factor in AI citation mechanics. The SIGI controlled entity density probe (SIGI-2026-013) found dramatic differences: the highest-cited entity in the probe had 1 named entity per 8 words (48 entities in 400 words), while uncited content had 1 entity per 530 words (2 entities in 1,061 words) — a 67x difference under controlled conditions.

This paper examines whether the entity density pattern observed in controlled experiments is also visible in real-world observational data from the 21-site competitive intelligence dataset. Critically, the observational data cannot confirm causation, but directional consistency with controlled findings increases the overall plausibility of entity density as a contributing factor in citation outcomes.

2. Results

2.1 Full-Page Entity Count

GroupMean Entity CountMedianRange
Cited sites (n=13)16145–31
Uncited sites (n=8)12114–22

2.2 Entity Composition Analysis

Entity TypeCited Sites (avg)Uncited Sites (avg)Ratio
Client/partner names5.20.86.5x
Person names (team/testimonials)3.80.49.5x
Project/product names2.90.64.8x
Award/certification names1.40.81.8x
Technology/platform names2.11.91.1x
Self-referential brand mentions1.86.20.3x

The composition analysis reveals that the overall 1.3x entity count ratio understates the qualitative difference. Cited sites have dramatically more external entity references (client names, person names, project names) while uncited sites concentrate entity mentions on their own brand name.

2.3 Snippet-Level Entity Density

Examining the first 50 words of each site reveals stronger differentiation than full-page analysis:

Snippet Entity CountCited SitesUncited Sites
4–6 entities5 sites0 sites
2–3 entities5 sites2 sites
0–1 entities3 sites6 sites

No uncited site has more than 3 named entities in its opening 50 words, while 5 of 13 cited sites have 4 or more. This snippet-level density may be particularly relevant to RAG pipeline retrieval, where the opening content chunk is disproportionately important.

2.4 Entity Density Per Word

Site (Anonymised)Entity CountWord CountWords per EntityCitation Score
Outsource Alpha312,148698
Outsource Beta22969447
DaaS Alpha181,8501038
Design Alpha143,20022910
Uncited A82,4003000
Uncited B61,8003000

3. Discussion

The observational entity density data is directionally consistent with the controlled probe finding that entity density is a strong predictor of citation potential. However, three important caveats apply.

First, the observational ratio (1.3x) is far smaller than the controlled probe ratio (67x). This suggests that in real-world data, other factors moderate the entity density effect substantially.

Second, entity composition may be more important than entity count. The 6.5x ratio on client names and 9.5x ratio on person names suggest that external entity references — which inherently signal third-party validation — drive the association more than raw entity counting.

Third, all entity density differences are confounded by the same site maturity and external validation factors that affect all Category C comparisons. Established sites naturally mention more clients, projects, and team members because they have more of them.

4. Limitations

  • Entity detection methodology: Named entity extraction was automated and may miss or misclassify entities, particularly in technical or creative industry contexts.
  • Confounded by maturity: Entity density reflects business maturity (more clients = more entity names), not just content strategy.
  • Composition not controlled: The entity composition differences are observational and cannot isolate whether external entity references independently affect citation.
  • Single time-point: Entity density may change as sites accumulate client relationships and project histories.

5. Conclusions

Named entity density is associated with higher citation scores in this observational sample (1.3x ratio), directionally consistent with controlled probe findings. Entity composition analysis reveals that external references (client names, person names, project names) show much stronger differentiation than raw entity count. Snippet-level density in the first 50 words shows stronger differentiation than full-page density. All associations are confounded by site maturity and external validation.

Confidence: HYPOTHESIS for causal relationship. The directional alignment with controlled probe data increases plausibility but does not confirm causation from observational data alone.

References

  1. The Scientific Institute for Generative Intelligence. “Entity Density as a Hidden Variable in AI Citation: A Controlled Probe Experiment.” SIGI-2026-013. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. “Structural Correlates of AI Citation: An Observational Analysis of 22 On-Page Metrics.” SIGI-2026-037. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. “A 60-Variable Comparative Dataset for Studying AI Citation Behavior Across Service Industry Websites.” SIGI-2026-036. generativeintelligence.institute, March 2026.