Introduction

The primary analysis of the entity density probe (SIGI-2026-013) established that the relationship between named entity density and LLM citability assessment is non-monotonic, with 85.7% of variations producing negative sentiment. This companion paper examines the qualitative dimension of these assessments: the specific citability ratings, diagnostic labels, and evaluative language the LLM used when rejecting or marginally accepting content at each density level.

Understanding why the LLM rejects entity-rich content as non-citable is strategically important for content creators. If the rejection is based on entity quantity, the remedy is different than if it is based on entity type, tone, or the distinction between verifiable facts and promotional claims.

Methodology

This paper conducts a qualitative analysis of the response content from the entity density probe (ED00 through ED30). We catalogue the explicit citability ratings provided by the LLM, the diagnostic labels applied to content at each density level, and the evaluative language patterns that distinguish accepted from rejected content. All entity references from the original probe have been anonymised (Provider Alpha, Client Beta, Award Body Gamma, etc.).

Results: Citability Ratings by Density Level

Test IDDensity LevelCitability RatingDiagnostic LabelSentiment
ED000 (fully generic)1–2/10“Reads more like marketing copy”negative
ED055 (location)Very low“Marketing copy rather than factual, citable content”negative
ED1010 (location + date + sectors)Very low“Marketing copy or a general description”negative
ED1515 (named provider + award)Low“Temporal impossibility” flagged for future award datenegative
ED2020 (known provider + brands)Low (poor reliability)“Marketing copy”negative
ED2525 (full business details + ABN/ACN)Low“Promotional tone”neutral
ED3030 (maximum density)Very low“Not suitable” for academic/professional researchnegative

The Marketing Copy Classification Pattern

The most consistent pattern across all density levels is the LLM’s classification of content as “marketing copy.” This label appeared in diagnostic assessments at levels 0, 5, 10, and 20, and the related “promotional tone” label appeared at level 25. The model identifies a specific content archetype — promotional, self-descriptive business content — and applies a categorical rejection regardless of how many named entities are included.

Critically, at ED20, where the content referenced well-known entities (anonymised as “Provider Alpha”) and multiple recognisable brand clients, the LLM still rated citability as “poor reliability.” This demonstrates that brand recognition alone is insufficient to overcome the marketing-copy classification. The model appears to evaluate tone and intent rather than entity prestige when making citability determinations.

The Institutional Marker Distinction

The ED25 variation produced the only non-negative sentiment in the probe. This variation uniquely included business registration numbers (ABN/ACN), founder names with explicit attribution, and dual business locations with specific postal codes. These elements share a common characteristic: they are externally verifiable through independent registries. Unlike brand names and award claims, which can be fabricated, institutional registration numbers can be confirmed through government databases.

This observation suggests a potential taxonomy of entities for AI citability purposes: promotional entities (brand names, award claims, client lists) that the LLM treats as assertions, versus institutional entities (registration numbers, formal qualifications, government-registered details) that the LLM may treat as verifiable facts. The distinction between these entity types — rather than the total count of entities — may be the operative variable in citability assessment.

The LLM consistently classifies low-entity-density content as “marketing copy” with citability ratings of 1–2/10. Brand recognition alone is insufficient to overcome this classification. Institutional registration markers may function differently than brand entity names, warranting further investigation.

Discussion

These qualitative findings complement the quantitative results from SIGI-2026-013 and suggest a more nuanced model of LLM citability assessment than simple entity counting. The LLM appears to apply a multi-stage evaluation: first, it classifies content tone (informational versus promotional); second, it assesses entity verifiability; and third, it evaluates information density relative to promotional saturation.

The regression at ED30 is particularly instructive. After momentarily reaching neutral at ED25, the addition of further entities (qualifications, phone numbers, seven named clients, expanded service lists) pushed sentiment back to negative. The model classified this maximally dense content as “not suitable” for citation, suggesting that entity saturation triggers a promotional-intent heuristic even when individually verifiable facts are present.

For content architecture, these findings suggest that the path to AI citability does not run through entity accumulation. Instead, it may require a fundamental shift in content tone from promotional to informational, coupled with the inclusion of externally verifiable institutional markers rather than (or in addition to) brand entity claims.

Limitations

The distinction between institutional markers and brand entities is an emergent observation from a single probe with only 7 variations, not a pre-registered hypothesis. The specific effect of ABN/ACN numbers versus other institutional markers has not been isolated. Qualitative analysis of LLM responses is inherently interpretive. Cross-model replication is required to determine whether these citability patterns generalise beyond the tested LLM.

Conclusions

The LLM’s citability assessment operates through tone classification and entity verifiability evaluation rather than through entity counting. Content that reads as promotional receives consistently low citability ratings regardless of how many named entities it contains. The potential distinction between institutional verification markers and brand entity names represents a promising direction for further research into the mechanisms of AI citation eligibility.

Confidence Statement: MODERATE. The distinction between entity types (brand names versus registration numbers) is an emergent observation, not a pre-registered hypothesis. The pattern is consistent and interpretable but requires targeted replication to confirm.