Introduction
The primary analysis of the entity density probe (SIGI-2026-013) established that the relationship between named entity density and LLM citability assessment is non-monotonic, with 85.7% of variations producing negative sentiment. This companion paper examines the qualitative dimension of these assessments: the specific citability ratings, diagnostic labels, and evaluative language the LLM used when rejecting or marginally accepting content at each density level.
Understanding why the LLM rejects entity-rich content as non-citable is strategically important for content creators. If the rejection is based on entity quantity, the remedy is different than if it is based on entity type, tone, or the distinction between verifiable facts and promotional claims.
Methodology
This paper conducts a qualitative analysis of the response content from the entity density probe (ED00 through ED30). We catalogue the explicit citability ratings provided by the LLM, the diagnostic labels applied to content at each density level, and the evaluative language patterns that distinguish accepted from rejected content. All entity references from the original probe have been anonymised (Provider Alpha, Client Beta, Award Body Gamma, etc.).
Results: Citability Ratings by Density Level
| Test ID | Density Level | Citability Rating | Diagnostic Label | Sentiment |
|---|---|---|---|---|
| ED00 | 0 (fully generic) | 1–2/10 | “Reads more like marketing copy” | negative |
| ED05 | 5 (location) | Very low | “Marketing copy rather than factual, citable content” | negative |
| ED10 | 10 (location + date + sectors) | Very low | “Marketing copy or a general description” | negative |
| ED15 | 15 (named provider + award) | Low | “Temporal impossibility” flagged for future award date | negative |
| ED20 | 20 (known provider + brands) | Low (poor reliability) | “Marketing copy” | negative |
| ED25 | 25 (full business details + ABN/ACN) | Low | “Promotional tone” | neutral |
| ED30 | 30 (maximum density) | Very low | “Not suitable” for academic/professional research | negative |
The Marketing Copy Classification Pattern
The most consistent pattern across all density levels is the LLM’s classification of content as “marketing copy.” This label appeared in diagnostic assessments at levels 0, 5, 10, and 20, and the related “promotional tone” label appeared at level 25. The model identifies a specific content archetype — promotional, self-descriptive business content — and applies a categorical rejection regardless of how many named entities are included.
Critically, at ED20, where the content referenced well-known entities (anonymised as “Provider Alpha”) and multiple recognisable brand clients, the LLM still rated citability as “poor reliability.” This demonstrates that brand recognition alone is insufficient to overcome the marketing-copy classification. The model appears to evaluate tone and intent rather than entity prestige when making citability determinations.
The Institutional Marker Distinction
The ED25 variation produced the only non-negative sentiment in the probe. This variation uniquely included business registration numbers (ABN/ACN), founder names with explicit attribution, and dual business locations with specific postal codes. These elements share a common characteristic: they are externally verifiable through independent registries. Unlike brand names and award claims, which can be fabricated, institutional registration numbers can be confirmed through government databases.
This observation suggests a potential taxonomy of entities for AI citability purposes: promotional entities (brand names, award claims, client lists) that the LLM treats as assertions, versus institutional entities (registration numbers, formal qualifications, government-registered details) that the LLM may treat as verifiable facts. The distinction between these entity types — rather than the total count of entities — may be the operative variable in citability assessment.
Discussion
These qualitative findings complement the quantitative results from SIGI-2026-013 and suggest a more nuanced model of LLM citability assessment than simple entity counting. The LLM appears to apply a multi-stage evaluation: first, it classifies content tone (informational versus promotional); second, it assesses entity verifiability; and third, it evaluates information density relative to promotional saturation.
The regression at ED30 is particularly instructive. After momentarily reaching neutral at ED25, the addition of further entities (qualifications, phone numbers, seven named clients, expanded service lists) pushed sentiment back to negative. The model classified this maximally dense content as “not suitable” for citation, suggesting that entity saturation triggers a promotional-intent heuristic even when individually verifiable facts are present.
For content architecture, these findings suggest that the path to AI citability does not run through entity accumulation. Instead, it may require a fundamental shift in content tone from promotional to informational, coupled with the inclusion of externally verifiable institutional markers rather than (or in addition to) brand entity claims.
Limitations
The distinction between institutional markers and brand entities is an emergent observation from a single probe with only 7 variations, not a pre-registered hypothesis. The specific effect of ABN/ACN numbers versus other institutional markers has not been isolated. Qualitative analysis of LLM responses is inherently interpretive. Cross-model replication is required to determine whether these citability patterns generalise beyond the tested LLM.
Conclusions
The LLM’s citability assessment operates through tone classification and entity verifiability evaluation rather than through entity counting. Content that reads as promotional receives consistently low citability ratings regardless of how many named entities it contains. The potential distinction between institutional verification markers and brand entity names represents a promising direction for further research into the mechanisms of AI citation eligibility.
Confidence Statement: MODERATE. The distinction between entity types (brand names versus registration numbers) is an emergent observation, not a pre-registered hypothesis. The pattern is consistent and interpretable but requires targeted replication to confirm.