Content Extraction by Source Type: What AI Systems Extract from Directories, Listicles, and Agency Homepages
Abstract
This paper documents how a large language model differentiates content extraction behaviour based on source type during retrieval-augmented generation. Through introspective analysis of the model's extraction process across four source categories -- directories, agency-published listicles, individual agency homepages, and editorial content -- we identify distinct extraction patterns for each. From directories, the model extracts primarily structural data (entity names, locations, review scores) while ignoring individual entity descriptions as too generic. From agency-published listicles, the model extracts specific evaluative claims about third-party entities while discounting the publisher's self-description. From individual agency homepages, all content is treated as self-reported with corroboration sought from third-party sources. The highest-value extractable content is identified as specific evaluative claims about third parties -- verifiable facts about external entities presented in a non-self-interested context. Generic, interchangeable descriptions are never extracted regardless of source type, confirming that specificity is a universal prerequisite for citation utility.
Keywords
content extraction, source type, AI citation, directories, listicles, agency homepages, third-party claims, self-reported content, generative engine optimisation
1. Introduction
When an AI system constructs an answer from multiple retrieved sources, it does not extract content uniformly. Different source types may be treated differently in terms of what content is extracted, how it is attributed, and what level of trust is assigned to extracted claims. Understanding these extraction patterns is essential for content strategists seeking to optimise for AI citation, as the same claim may be citation-worthy when it appears in one source type but ignored when it appears in another.
2. Methodology
The LLM was presented with search results containing all four source types (directories, agency-published listicles, individual agency homepages, and editorial content positions that were vacant in this query space). Through introspective probing, the model was asked to describe what it would extract from each source type, what it would ignore, and what criteria determined the difference. All sources and entities are anonymised.
3. Results
3.1 Extraction by Source Type
| Source Type | Content Extracted | Content Ignored | Trust Treatment |
|---|---|---|---|
| Directories | Entity names, locations, review scores, service categories, agency counts | Individual entity descriptions (too generic) | Structural data trusted; evaluative claims treated as commercial |
| Agency-published listicles | Specific claims about OTHER agencies: named clients, founding dates, specific awards, team sizes, named methodologies | Publisher's description of themselves | Third-party evaluative claims cited; self-description discounted |
| Individual agency homepages | Self-reported specialisations, client lists, award badges, team size, pricing | Generic marketing copy | ALL treated as self-reported; corroboration sought |
| Editorial content (hypothetical) | Independent methodology, comparative analysis, industry context | N/A (none available) | Would receive highest trust for evaluative claims |
3.2 Highest Citation Value Content
The content with the highest citation value across all source types was specific evaluative claims about third-party entities. For example, a listicle describing a competitor agency's multi-decade history and a specific internationally recognised identity project represented the most citable content in the entire result set. The claim was specific (named entity, named project, verifiable timeframe), verifiable (award databases and public records can confirm), and non-self-interested (the publisher was evaluating an external entity, not promoting itself).
3.3 Universal Extraction Filter
Regardless of source type, generic interchangeable descriptions were never extracted. Descriptions such as "full-service creative agency" or "strategic branding solutions" that could apply to any entity in the category were filtered out at the extraction stage. This suggests that specificity functions as a universal prerequisite for citation utility, independent of source type trust level.
4. Discussion
The source-type-dependent extraction pattern has important implications for content strategy. For publishers of listicle content, the most citation-worthy content is not their self-description but their evaluative descriptions of other entities. This creates a counter-intuitive incentive: the more specific and well-researched a publisher's descriptions of competitors, the more likely the publisher's content is to be cited overall, even though the publisher's self-description is discounted.
For agency homepages, the finding that all content is treated as self-reported underscores the importance of third-party validation. An agency can improve its citation prospects not by adding more content to its own site but by ensuring that specific verifiable claims about it appear on third-party sources.
The universal specificity filter suggests that content quality, as measured by factual specificity, is more important than source type for extraction. A specific verifiable claim from a lower-trust source may be more citable than a generic claim from a higher-trust source.
5. Limitations
- Introspective report: Extraction behaviour is based on the model's self-report, which may not fully reflect actual extraction patterns.
- Single query domain: Extraction patterns may vary across different service categories and markets.
- No editorial baseline: The absence of editorial content in the result set means the editorial extraction pattern is hypothetical.
- Single model: Different AI systems may implement different extraction strategies.
6. Conclusions
The LLM reports differentiating extraction behaviour by source type, preferentially citing specific evaluative claims about third parties over self-promotional content. Directories are mined for structural data, listicles for third-party evaluative claims, and homepages are treated as self-reported requiring corroboration. Generic descriptions are universally filtered regardless of source type trust level. These findings suggest that content specificity is the primary determinant of extraction utility, with source type governing trust attribution rather than extraction eligibility.
Confidence: MODERATE. Introspective report on extraction behaviour. Behavioural confirmation would require tracking actual extraction patterns across many queries.
References
- The Scientific Institute for Generative Intelligence. "Search Position Versus Citation Priority." SIGI-2026-056. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Entity Density in Search Snippets." SIGI-2026-060. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "A Five-Tier Trust Hierarchy for AI Citation Sources." SIGI-2026-053. generativeintelligence.institute, March 2026.