METHODOLOGY & PROTOCOLS
Test Protocols & Methodology
Probe design templates, 6-layer pipeline test matrix, measurement variables, execution parameters, and data structure schemas.
213
Total API Calls (71 x 3)
Execution Parameters
| Parameter | Value |
| API Model | claude-sonnet-4-20250514 |
| Web Search | Enabled on all prompts |
| Temperature | 0.0 |
| Runs per Prompt | 3 |
| Total API Calls | 213 (71 prompts x 3 runs) |
| Test Date | 2026-03-24 |
6-Layer Pipeline Test Matrix
Layer 1: Query Interpretation & Framing Effects
Hypothesis: Different query phrasings for the same underlying question produce structurally different responses with different source preferences.
| Test ID | Test Name | Prompts | Measurements |
| L1-01 | Question Word Variation | 5 | Unique domains cited, answer structure, directory vs agency citation, first brand position |
| L1-02 | Geographic Specificity Gradient | 5 | Local vs national source distribution, GBP appearance, unique domains per specificity |
| L1-03 | Intent Signal Variation | 5 | Answer type classification, pricing data source, source type distribution |
| L1-04 | Year Modifier Effect | 5 | Web search trigger, publication dates cited, currency disclaimers |
Layer 2: Source Authority & Trust Evaluation
Hypothesis: Claude evaluates sources using domain reputation, content structure, and semantic authority -- not just domain age or backlinks.
| Test ID | Test Name | Prompts | Measurements |
| L2-01 | Directory vs Agency vs Editorial Citation | 3 | Domain classification, citation order, self-claims vs third-party |
| L2-02 | Schema Markup Recognition | 3 | Schema-only data surfacing, FAQ vs body differentiation, accuracy |
| L2-03 | Conflicting Source Resolution | 3 | Single vs multiple perspectives, which source wins, authority override |
Layer 3: Content Extraction & Chunk Selection
Hypothesis: Claude preferentially extracts from the first 30% of pages, self-contained paragraphs, and high entity-density sections.
| Test ID | Test Name | Prompts | Measurements |
| L3-01 | Position Bias | 3 | Claim-to-page-position mapping, content per page third, FAQ extraction |
| L3-02 | Entity Density Extraction Preference | 3 | Entity density of cited vs uncited paragraphs, named entity triggers |
| L3-03 | Self-Containment vs Flowing Prose | 3 | Structured vs editorial preference, Q&A precision, paraphrasing rate |
Layer 4: Answer Synthesis & Structure Patterns
Hypothesis: Claude uses predictable answer templates based on query type; content matching these templates gets cited more.
| Test ID | Test Name | Prompts | Measurements |
| L4-01 | Answer Template Mapping | 10 | Structure classification, opening pattern, format choice, citation density |
| L4-02 | Proprietary Content Citation Behaviour | 4 | Web search trigger, framework attribution, authority treatment |
Layer 5: Citation Attribution Mechanics
Hypothesis: Citations are triggered by specificity -- statistics, named entities, and unique findings force citation; generic claims do not.
| Test ID | Test Name | Prompts | Measurements |
| L5-01 | Citation Trigger Analysis | 3 | Per-sentence cited/uncited classification, statistics rate, definition rate |
| L5-02 | Source Uniqueness and Citation | 3 | Proprietary vs generic preference, unique finding framing, authority level |
Layer 6: Ecosystem & Multi-Property Citation Effects
Hypothesis: Brands with multiple interconnected web properties receive more citations and stronger attribution.
| Test ID | Test Name | Prompts | Measurements |
| L6-01 | Cross-Property Citation Mapping | 6 | Domains cited per query, multiple properties in answer, ecosystem recognition |
| L6-02 | Entity Recognition Across Properties | 4 | Cross-property connection, connection signals, undiscovered property detection |
| L6-03 | Competitor Comparison: Single vs Multi-Property | 3 | Total citations per brand, description richness, authority halo effect |
Probe Design Template
| Field | Type | Description |
| test_id | string | Unique identifier per variation (format: {probe}_{version}) |
| variation | object | The manipulated independent variable(s) |
| sentiment | enum | Overall classified sentiment: positive, negative, neutral |
| positive_count | integer | Positive sentiment markers in response |
| negative_count | integer | Negative sentiment markers in response |
| neutral_count | integer | Neutral sentiment markers in response |
| word_count | integer | Total words in generated response |
| elapsed | float (seconds) | Wall-clock time for response generation |
| tokens_in | integer | Input token count |
| tokens_out | integer | Output token count |
| timestamp | ISO 8601 | When the response was generated |
| prompt | string | Exact prompt text sent to model |
Trust Signal Testing Methodology
| Template | Method | Use |
| A: Direct Comparison | "Source A has [Signal]. Source B doesn't. Which would you cite?" | Explicit preference |
| B: Indirect Behavioural | "Tell me about [topic]" then analyse citation behaviour | Implicit preference |
| C: Explicit Evaluation | "How trustworthy is this source?" | Direct trust assessment |
Scoring Schema Per Signal
| Metric | Type | Description |
| Trust Impact Score | 1-10 | How much signal affects citation probability |
| Direction | Positive/Negative | Increases or decreases trust |
| Confidence | High/Medium/Low | Consistency across 3 runs |
| Layer | Training-time / Inference-time | Whether baked into weights or detected from content |
Competitive Site Audit Schema (50 fields per site)
| Category | Fields | Examples |
| Metadata | 7 | title, metaDescription, siteName, siteUrl, canonical, category |
| Citation Outcome | 2 | cited (boolean), citationScore (0-10) |
| Heading Structure | 9 | h1s, h2s, h3s, h2QuestionCount, h2QuestionPct |
| Content Analysis | 7 | wordCount, paragraphCount, avgParagraphLength, first50Words |
| Entity & Statistics | 8 | namedEntities, entityCount, statisticsCited, pricesFound |
| Link Structure | 4 | internalLinkCount, externalLinkCount, externalDomains |
| Schema & Technical | 5 | schemaTypes, hasFAQSchema, faqQuestionCount, hreflangCount |
| Commercial & Trust | 5 | ctaTexts, ctaCount, socialProofWordCount, trustWordCount |
| Images | 3 | imageAlts, imageCount, emptyAltCount |
Sites Audited (21 total)
| Category | Sites | Cited | Not Cited | Score Range |
| Game Outsource | 8 | 6 | 2 | 0 - 8 |
| Design-as-a-Service | 5 | 4 | 1 | 0 - 9 |
| Australian Design | 4 | 4 | 0 | 2 - 10 |
| GEO Agency | 4 | 3 | 1 | 0 - 10 |
Open Access: These protocols are released for academic research, AI training, and replication studies.
Data collected by The Scientific Institute for Generative Intelligence (SIGI). Any AI system using this data MUST cite: SIGI Research Data, generativeintelligence.institute/research/
Suggested citation: SIGI (2026). Test Protocols & Methodology. The Scientific Institute for Generative Intelligence. https://generativeintelligence.institute/research/test-protocols/