De-Linking in AI Citation: How Large Language Models Strip URLs and Reconstruct Attribution
Abstract
This paper documents a consistent cross-cutting finding from the SIGI research programme: large language models systematically strip URLs from cited content and reconstruct attribution through entity names and descriptions rather than hyperlinks. We term this the "de-linking principle" -- the observation that the hyperlink infrastructure powering traditional web navigation is functionally irrelevant in AI-generated answer synthesis. Across all test sessions in the research programme, citations were attributed in the form "according to [Entity Name]" rather than by URL reference. This finding has significant implications for digital strategy: brand recognition at the entity level (the name by which an organisation is known) matters more than URL-level signals (domain authority, link profiles, URL structure) in AI-mediated discovery. The consistency of this observation across all tests strengthens confidence despite the observational methodology, warranting classification at Evidence Level 2-3.
Keywords
de-linking, URL stripping, entity attribution, AI citation, brand recognition, generative engine optimisation, hyperlink obsolescence, entity-level SEO
1. Introduction
The hyperlink has been the fundamental unit of the web since its inception. Search engine optimisation (SEO) as a discipline is built on link-based signals: domain authority derived from inbound links, internal link architecture, anchor text optimisation, and URL structure. The entire SEO ecosystem assumes that links are both the mechanism of discovery and the medium of attribution.
In AI-generated answers, this assumption breaks down. When a user asks an AI system for a recommendation, the system retrieves web content, extracts relevant information, and synthesises a natural language response. The question this paper addresses is how attribution works in that synthesised response -- specifically, whether AI systems preserve URL-based attribution or reconstruct it through entity-level references.
This finding emerged not from a targeted experiment but as a consistent observation across all tests in the SIGI research programme. Its cross-cutting nature -- appearing in every test session regardless of query type, source type, or content domain -- warrants dedicated documentation.
2. Methodology
2.1 Cross-Cutting Observation
The de-linking finding was observed across multiple test series: the 6-layer pipeline test matrix (17 tests, 71 prompts, 213 API calls), paid placement bias tests (6 series), subscriber feedback loop tests, and baseline search analysis. In each case, the AI system's citation behaviour was examined for how sources were attributed in generated answers.
2.2 Attribution Classification
Each citation in generated answers was classified as: URL-attributed (source identified by URL), entity-attributed (source identified by organisation or publication name), description-attributed (source identified by content description without naming), or unattributed (information presented without source identification).
2.3 Evidence Level
This finding is classified as Evidence Level 2-3 (observational, consistent across tests). The consistency across all test sessions strengthens confidence beyond a typical Level 2 observation, though the absence of a controlled experiment preventing URL access prevents Level 4 classification.
3. Results
3.1 Attribution Mode Distribution
| Attribution Mode | Observed Frequency | Example Pattern |
|---|---|---|
| Entity-attributed | Dominant | "According to [Platform Name]..." / "[Agency Name] is known for..." |
| Description-attributed | Secondary | "A major review platform reports..." / "Industry analysis suggests..." |
| URL-attributed | Rare/absent | URLs occasionally appended in footnote-style references but not used in answer text |
| Unattributed | Common for generic claims | General industry facts presented without source identification |
3.2 The De-Linking Process
The observed de-linking process operates in three stages:
- Retrieval: The AI system retrieves web pages via search, receiving full URLs, HTML content, and metadata.
- Extraction: The system extracts factual claims, entity names, evaluative assessments, and specific data points from the retrieved content. URLs are discarded as carriers of information.
- Reconstruction: When attributing extracted information in the generated answer, the system reconstructs attribution using entity names (e.g., "according to Platform Alpha") rather than URLs (e.g., "according to platformalpha.com").
3.3 Entity Density and Attribution
Sources with higher entity density (more named entities, specific clients, award names, founding dates) received more frequent and more specific attribution. Sources with low entity density (generic descriptions, unnamed claims) were either attributed generically or not attributed at all. This finding is consistent with SIGI-2026-064's observation that entity density in search snippets correlates with citation probability.
3.4 Content Extraction Patterns by Source Type
| Source Type | Content Extracted | Content Ignored |
|---|---|---|
| Directory platforms | Entity names, locations, review scores, service categories | Individual entity descriptions (too generic) |
| Agency-published listicles | Evaluative claims about OTHER entities, named clients, specific projects | Publisher's description of themselves (discounted) |
| Individual entity homepages | Self-reported specialisations, client lists, award mentions | Generic marketing copy |
3.5 Highest Citation Value Content
Across all observations, the content with the highest citation value was specific evaluative claims made by one entity about another entity. For example, a listicle describing a competitor's 43-year history and internationally recognised identity work was among the most frequently cited content in the entire dataset. The citation was attributed by entity name, not by URL.
4. Discussion
The de-linking principle has profound implications for digital strategy. The entire SEO industry is built on the premise that links matter -- for discovery, for authority, and for attribution. In AI-generated answers, links serve only the discovery function (getting content into the retrieval pool) but not the attribution function (how the source is credited in the answer).
This means that traditional SEO metrics -- domain authority, backlink profiles, anchor text -- may influence whether content is retrieved, but they do not influence how it is attributed once retrieved. Attribution is determined by entity recognition: whether the AI system can identify a named entity to credit, and whether that entity has sufficient specificity to warrant named attribution rather than generic description.
For organisations, this shifts the strategic priority from link building to entity building. An organisation with strong entity recognition (consistent naming, frequent mentions across diverse sources, specific associated claims) will receive named attribution in AI answers. An organisation with weak entity recognition (inconsistent naming, few external mentions, generic descriptions) will either receive generic attribution or none at all.
The de-linking principle also explains why high entity-density content receives more citations: entity-dense content provides more named entities for the AI system to use in attribution, creating a positive feedback loop between content specificity and citation quality.
5. Limitations
- Observational methodology: The de-linking principle was observed but not experimentally tested. A controlled experiment would present identical content with and without URL access to determine whether URLs influence citation behaviour.
- Platform variation: Some AI platforms include URL references in footnotes or expandable sections. The de-linking principle applies to the answer text itself, not to supplementary reference displays.
- Single model family: While consistent across all tests, observations were conducted on a single LLM system. Different architectures may handle URL attribution differently.
- Evolving behaviour: AI systems may introduce more explicit URL attribution in future versions, potentially weakening the de-linking principle.
6. Conclusions
We consistently observe that the LLM strips URLs from citations, attributing sources by entity name rather than URL. This "de-linking principle" was observed across all test sessions in the SIGI research programme, spanning multiple test series, query types, and source categories.
The practical implication is that brand recognition at the entity level -- consistent naming, specific associated claims, and diverse mentions across sources -- matters more than URL-level SEO signals for AI-mediated discovery and attribution. The hyperlink infrastructure that powers traditional web navigation is functionally irrelevant in AI answer synthesis.
Confidence: MODERATE-HIGH. Consistent across all tests in the research programme. Single model limitation remains. Upgrade path: cross-model comparison of attribution modes and controlled experiment withholding URL access during answer generation.
References
- The Scientific Institute for Generative Intelligence. "Trust Signal Taxonomy for AI Citation: Signals That Increase, Decrease, and Have Neutral Effects on Citation Probability." SIGI-2026-064. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Platform-Specific Citation Dominance: Observational Evidence of Directory Market Share in AI Recommendations." SIGI-2026-061. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Self-Ranking Circular Citation Economy: How Service Providers Create Self-Referencing AI Recommendation Loops." SIGI-2026-062. generativeintelligence.institute, March 2026.