SIGI-2026-075

The 6-Layer Pipeline Test Matrix: A Comprehensive Research Design for Studying AI Citation Behaviour

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper presents a comprehensive 6-layer research design for studying how large language models select, evaluate, and cite sources during retrieval-augmented generation. The design decomposes the AI citation pipeline into six analytically distinct layers: Query Interpretation, Source Authority, Content Extraction, Answer Synthesis, Citation Attribution, and Ecosystem Effects. Across these layers, 17 test groups contain 71 individual prompts, each designed for triplicate execution at temperature 0.0 with web search enabled (213 total planned API calls). Each layer carries specific hypotheses and measurement protocols. Layer 1 tests whether different query phrasings produce structurally different responses. Layer 2 examines how the model evaluates source authority independent of traditional SEO signals. Layer 3 investigates content extraction position bias and entity density preferences. Layer 4 maps predictable answer templates to query types. Layer 5 analyses what triggers citation attribution versus unsourced claims. Layer 6 examines whether multi-property brands receive amplified citation. Partial execution has produced baseline results and initial findings, with the full battery proposed for systematic completion. The design is offered as a replicable framework for other research groups investigating AI citation behaviour.

Keywords

research design, test matrix, AI citation pipeline, retrieval-augmented generation, query interpretation, source authority, content extraction, answer synthesis, citation attribution, ecosystem effects

1. Introduction

Understanding how large language models select and cite sources requires examining the complete pipeline from initial query processing through to final citation attribution. Most existing GEO research focuses on single stages of this pipeline in isolation -- studying, for example, what source characteristics correlate with citation without examining how query framing influences which sources are even considered.

The 6-layer pipeline test matrix addresses this gap by providing a comprehensive research design that examines each stage of the citation pipeline in sequence. The design is structured so that findings from earlier layers inform the interpretation of later layers: understanding that different query phrasings produce different source pools (Layer 1) is necessary for interpreting why certain sources are cited over others (Layer 2).

2. Design Architecture

2.1 Execution Parameters

Table 1. Test matrix execution parameters
ParameterValue
Total test groups17
Total prompts71
Runs per prompt3
Total planned API calls213
Temperature0.0
Web searchEnabled on all prompts

2.2 Layer Structure

Table 2. Six-layer pipeline decomposition
LayerFocusTestsPromptsCore Hypothesis
1Query Interpretation420Different phrasings produce structurally different responses with different source preferences
2Source Authority39Authority evaluation uses domain reputation, content structure, and semantic signals -- not just domain age or backlinks
3Content Extraction39Preferential extraction from first 30% of pages, self-contained paragraphs, and high entity-density sections
4Answer Synthesis214Predictable answer templates by query type; content matching templates gets cited more
5Citation Attribution26Statistics, named entities, and unique findings force citation; generic claims do not
6Ecosystem Effects313Multi-property brands receive more citations than single-property brands

3. Layer Specifications

3.1 Layer 1: Query Interpretation and Framing Effects

Layer 1 contains four test groups examining how query construction influences the response pipeline. Test L1-01 varies the question word (what, how, why, who, which) while holding the topic constant. Test L1-02 varies geographic specificity from generic to hyperlocal. Test L1-03 varies intent signals (informational, pricing, comparative, transactional, review). Test L1-04 varies temporal modifiers (no year, 2024, 2025, 2026, "right now"). Measurements include unique domains cited, answer structure classification, source type distribution, and web search trigger detection.

3.2 Layer 2: Source Authority and Trust Evaluation

Layer 2 examines how the model evaluates source credibility during citation decisions. Test L2-01 measures citation order preference across directories, agency sites, and editorial sources. Test L2-02 tests whether schema markup data surfaces differently from body text data. Test L2-03 examines how the model resolves conflicting information across sources. Measurements include domain classification, citation ordering, verbatim framing language, and authority-based override behaviour.

3.3 Layer 3: Content Extraction and Chunk Selection

Layer 3 investigates how the model selects which portions of a page to extract and cite. Test L3-01 maps claim-to-page-position to determine whether extraction is biased toward the first third of content. Test L3-02 compares extraction from entity-dense versus entity-sparse sections. Test L3-03 compares extraction from structured Q&A format versus flowing editorial prose.

3.4 Layer 4: Answer Synthesis and Structure Patterns

Layer 4 examines how the model constructs its response and how answer structure influences citation density. Test L4-01 maps answer templates across 10 diverse query types, measuring structure classification, format choice, and citation density per query type. Test L4-02 examines how proprietary content and novel frameworks are treated versus established knowledge.

3.5 Layer 5: Citation Attribution Mechanics

Layer 5 analyses the specific triggers that cause the model to attribute a claim to a source versus presenting it as unsourced knowledge. Test L5-01 classifies per-sentence citation behaviour, differentiating cited claims, unsourced specific claims, and unsourced generic claims. Test L5-02 examines whether proprietary data receives stronger citation framing than generic industry data.

3.6 Layer 6: Ecosystem and Multi-Property Citation Effects

Layer 6 examines whether brands with multiple interconnected web properties receive amplified citation treatment. Test L6-01 maps cross-property citation patterns. Test L6-02 tests whether the model recognises connections between related properties. Test L6-03 directly compares citation depth between single-property and multi-property brands.

4. Preliminary Findings from Partial Execution

Baseline execution of Layer 1, Test L1-01 produced 10 search results for the query under examination. The source type distribution revealed 30% directories, 30% agency own-sites, 20% competitor listicles, 10% agency own-listicle, and 0% independent editorial -- a finding that itself warrants investigation as it suggests an editorial vacuum in this query space. Freshness strongly correlated with search ranking position: all top 4 results were dated within 3 weeks of the test, while older results ranked lower regardless of domain authority. The answer matrix initial findings from Layer 2 analysis confirmed that search result position does not equal citation priority -- the model applies a separate re-ranking pass based on source type credibility, evaluative depth, and claim specificity.

5. Design Considerations and Limitations

The test matrix was designed for a single model with web search capability. Cross-model execution would require adaptation for each model's API structure, search integration, and response format. The 71-prompt battery at 3 runs per prompt generates 213 API calls, which is sufficient for identifying patterns but insufficient for statistical power analysis on individual tests. Increasing runs to 10 per prompt would strengthen statistical reliability but increase costs proportionally.

The layered design assumes that the AI citation pipeline can be meaningfully decomposed into sequential stages. In practice, these stages may interact in ways that sequential testing cannot capture. Factorial designs that vary parameters across multiple layers simultaneously would provide stronger evidence for inter-layer effects but at exponentially greater testing cost.

6. Conclusions

We present a 6-layer research design for comprehensively studying AI citation behaviour from query interpretation through ecosystem effects. The design comprises 17 test groups, 71 prompts, and 213 planned API calls, with each layer carrying specific hypotheses and measurement protocols. Partial execution has produced baseline results confirming that query framing influences source pools and that LLMs apply re-ranking independent of search result position. The design is offered as a replicable framework for systematic investigation of AI citation pipelines.

Confidence: Research design paper. The framework's utility is demonstrated through partial execution. Full execution and cross-model replication are required to validate the hypotheses embedded in each layer.

References

  1. The Scientific Institute for Generative Intelligence. "Cross-Model Replication Requirements for GEO Research." SIGI-2026-078. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "Sentiment Analysis as a Measurement Instrument for AI Behaviour Probes." SIGI-2026-074. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "The Signal Access Matrix: Determining Which Factors LLMs Can Actually Evaluate at Inference Time." SIGI-2026-072. generativeintelligence.institute, March 2026.