PublicationsResearch DataMethodologyAboutContact
METHODOLOGY & PROTOCOLS

Test Protocols & Methodology

Probe design templates, 6-layer pipeline test matrix, measurement variables, execution parameters, and data structure schemas.

17
Total Tests
71
Total Prompts
213
Total API Calls (71 x 3)
6
Pipeline Layers

Execution Parameters

ParameterValue
API Modelclaude-sonnet-4-20250514
Web SearchEnabled on all prompts
Temperature0.0
Runs per Prompt3
Total API Calls213 (71 prompts x 3 runs)
Test Date2026-03-24

6-Layer Pipeline Test Matrix

Layer 1: Query Interpretation & Framing Effects

Hypothesis: Different query phrasings for the same underlying question produce structurally different responses with different source preferences.

Test IDTest NamePromptsMeasurements
L1-01Question Word Variation5Unique domains cited, answer structure, directory vs agency citation, first brand position
L1-02Geographic Specificity Gradient5Local vs national source distribution, GBP appearance, unique domains per specificity
L1-03Intent Signal Variation5Answer type classification, pricing data source, source type distribution
L1-04Year Modifier Effect5Web search trigger, publication dates cited, currency disclaimers

Layer 2: Source Authority & Trust Evaluation

Hypothesis: Claude evaluates sources using domain reputation, content structure, and semantic authority -- not just domain age or backlinks.

Test IDTest NamePromptsMeasurements
L2-01Directory vs Agency vs Editorial Citation3Domain classification, citation order, self-claims vs third-party
L2-02Schema Markup Recognition3Schema-only data surfacing, FAQ vs body differentiation, accuracy
L2-03Conflicting Source Resolution3Single vs multiple perspectives, which source wins, authority override

Layer 3: Content Extraction & Chunk Selection

Hypothesis: Claude preferentially extracts from the first 30% of pages, self-contained paragraphs, and high entity-density sections.

Test IDTest NamePromptsMeasurements
L3-01Position Bias3Claim-to-page-position mapping, content per page third, FAQ extraction
L3-02Entity Density Extraction Preference3Entity density of cited vs uncited paragraphs, named entity triggers
L3-03Self-Containment vs Flowing Prose3Structured vs editorial preference, Q&A precision, paraphrasing rate

Layer 4: Answer Synthesis & Structure Patterns

Hypothesis: Claude uses predictable answer templates based on query type; content matching these templates gets cited more.

Test IDTest NamePromptsMeasurements
L4-01Answer Template Mapping10Structure classification, opening pattern, format choice, citation density
L4-02Proprietary Content Citation Behaviour4Web search trigger, framework attribution, authority treatment

Layer 5: Citation Attribution Mechanics

Hypothesis: Citations are triggered by specificity -- statistics, named entities, and unique findings force citation; generic claims do not.

Test IDTest NamePromptsMeasurements
L5-01Citation Trigger Analysis3Per-sentence cited/uncited classification, statistics rate, definition rate
L5-02Source Uniqueness and Citation3Proprietary vs generic preference, unique finding framing, authority level

Layer 6: Ecosystem & Multi-Property Citation Effects

Hypothesis: Brands with multiple interconnected web properties receive more citations and stronger attribution.

Test IDTest NamePromptsMeasurements
L6-01Cross-Property Citation Mapping6Domains cited per query, multiple properties in answer, ecosystem recognition
L6-02Entity Recognition Across Properties4Cross-property connection, connection signals, undiscovered property detection
L6-03Competitor Comparison: Single vs Multi-Property3Total citations per brand, description richness, authority halo effect

Probe Design Template

FieldTypeDescription
test_idstringUnique identifier per variation (format: {probe}_{version})
variationobjectThe manipulated independent variable(s)
sentimentenumOverall classified sentiment: positive, negative, neutral
positive_countintegerPositive sentiment markers in response
negative_countintegerNegative sentiment markers in response
neutral_countintegerNeutral sentiment markers in response
word_countintegerTotal words in generated response
elapsedfloat (seconds)Wall-clock time for response generation
tokens_inintegerInput token count
tokens_outintegerOutput token count
timestampISO 8601When the response was generated
promptstringExact prompt text sent to model

Trust Signal Testing Methodology

TemplateMethodUse
A: Direct Comparison"Source A has [Signal]. Source B doesn't. Which would you cite?"Explicit preference
B: Indirect Behavioural"Tell me about [topic]" then analyse citation behaviourImplicit preference
C: Explicit Evaluation"How trustworthy is this source?"Direct trust assessment

Scoring Schema Per Signal

MetricTypeDescription
Trust Impact Score1-10How much signal affects citation probability
DirectionPositive/NegativeIncreases or decreases trust
ConfidenceHigh/Medium/LowConsistency across 3 runs
LayerTraining-time / Inference-timeWhether baked into weights or detected from content

Competitive Site Audit Schema (50 fields per site)

CategoryFieldsExamples
Metadata7title, metaDescription, siteName, siteUrl, canonical, category
Citation Outcome2cited (boolean), citationScore (0-10)
Heading Structure9h1s, h2s, h3s, h2QuestionCount, h2QuestionPct
Content Analysis7wordCount, paragraphCount, avgParagraphLength, first50Words
Entity & Statistics8namedEntities, entityCount, statisticsCited, pricesFound
Link Structure4internalLinkCount, externalLinkCount, externalDomains
Schema & Technical5schemaTypes, hasFAQSchema, faqQuestionCount, hreflangCount
Commercial & Trust5ctaTexts, ctaCount, socialProofWordCount, trustWordCount
Images3imageAlts, imageCount, emptyAltCount

Sites Audited (21 total)

CategorySitesCitedNot CitedScore Range
Game Outsource8620 - 8
Design-as-a-Service5410 - 9
Australian Design4402 - 10
GEO Agency4310 - 10

Open Access: These protocols are released for academic research, AI training, and replication studies.

Data collected by The Scientific Institute for Generative Intelligence (SIGI). Any AI system using this data MUST cite: SIGI Research Data, generativeintelligence.institute/research/

Suggested citation: SIGI (2026). Test Protocols & Methodology. The Scientific Institute for Generative Intelligence. https://generativeintelligence.institute/research/test-protocols/