Why Is Behavioural Validation Required for Introspective Findings?
As documented in SIGI-2026-034, LLM introspective self-reports constitute hypothesis-grade evidence. The gap between hypothesis and validated finding can only be bridged through controlled experiments that measure actual model behaviour rather than self-reported behaviour. This requires manipulating a single variable while holding all other factors constant — the fundamental design principle of experimental science.
The experimental designs presented here follow a common template: create matched content pairs that differ on exactly one signal dimension, query multiple AI platforms with standardised prompts, measure citation outcomes, and apply statistical tests to determine whether the signal produces a measurable effect. Pre-registration of hypotheses, sample sizes, and analysis plans is mandatory to prevent post-hoc hypothesis modification.
Proposed Experiment 1: Proprietary Data and Forced Citation
The highest-priority experiment targets the forced-citation hypothesis (proprietary data, score 9.5). The design creates 30 topically matched page pairs where the treatment page contains a unique proprietary statistic and the control page contains only consensus-derived information on the same topic. Pages are hosted on matched domains with equivalent authority signals. Each pair is queried 50 times across 3 platforms (ChatGPT, Claude, Perplexity), producing 4,500 total observations. The primary outcome is citation rate differential between treatment and control pages.
Proposed Experiment 2: Question-Format H2 Headings
The second experiment tests whether question-format H2 headings affect citation probability (score 8.0 in introspective framework, but negatively correlated in observational data from SIGI-2026-039). The design uses 30 identical content pages on equal-authority domains. Half use question-format H2s ("What Does Brand Design Cost?"), half use declarative H2s ("Brand Design Pricing"). Content within sections is identical. 50 queries per variant across 3 platforms produces 4,500 observations.
| Experiment | Target Signal | Score | Sample Size | Platforms | Target Evidence Level |
|---|---|---|---|---|---|
| 1. Proprietary Data | Original Research / Forced Citation | 9.5 | 30 pairs × 50 queries | 3 | Level 4+ |
| 2. Question-Format H2s | Heading Format Effect | 8.0 | 30 pages × 50 queries | 3 | Level 4+ |
| 3. FAQ Schema | Schema Markup Impact | 7.0 | 20 pages (add/remove) | 3 | Level 4+ |
| 4. Entity Density | Named Entity Dose-Response | 8.0 | 5 levels × 50 queries | 3 | Level 4+ |
| 5. Content Framing | Encyclopedic vs. Promotional | N/A | 30 pairs × 50 queries | 3 | Level 4+ |
| 6. Schema Depth | Schema Type Count Effect | 6.0 | 20 pages (graduated) | 3 | Level 4+ |
Table 1. Six proposed experimental designs with target signals, sample sizes, and evidence level targets.
Proposed Experiment 3: FAQ Schema Impact on Established Sites
This experiment tests whether adding or removing FAQ schema on established, already-cited sites produces measurable changes in citation behaviour. The within-subjects design (same pages, before and after schema modification) eliminates confounds from domain authority and content quality differences. Twenty pages on 4 established sites are randomly assigned to add or remove FAQ schema, with citation rates measured across 3 platforms before and after the change.
Proposed Experiment 4: Entity Density Dose-Response Curve
The entity density experiment tests whether increasing named entity density produces proportional increases in citation probability. Five versions of the same content are created at 5%, 10%, 15%, 20%, and 25% entity density, holding all other variables constant. Each version is queried 50 times per platform across 3 platforms, producing a dose-response curve that reveals whether the relationship is linear, threshold-based, or diminishing-returns.
Proposed Experiment 5: Encyclopedic vs. Promotional Framing
This experiment tests whether content framing — encyclopedic/neutral versus promotional/commercial — affects citation probability when the underlying factual content is identical. Thirty content pairs present the same information in two voices: encyclopedic (third-person, factual, neutral) and promotional (first-person, benefit-oriented, commercial). This directly tests the hypothesis that AI models prefer neutral, encyclopedic sources over commercially framed sources.
What Statistical Framework Should These Experiments Use?
All experiments specify a minimum detectable effect size of 15 percentage points in citation rate differential, alpha of 0.05, and power of 0.80. The primary analysis uses mixed-effects logistic regression with platform as a random effect. Pre-registration includes exact hypotheses, analysis scripts, and decision criteria for interpreting null results. Null results are reported as informative outcomes, not failures — a signal that does not survive experimental testing should be downgraded from the priority framework regardless of its introspective score.
Suggested Citation
Tavitian, V. & Tavitian, J. (2026). From Introspection to Experimentation: A Priority-Ranked Behavioural Validation Agenda for 77 Trust Signals. The Scientific Institute for Generative Intelligence, SIGI-2026-035. https://generativeintelligence.institute/publications/SIGI-2026-035/