The Introspection Problem: Why LLM Self-Reports About Citation Behaviour May Not Reflect Actual Computation
Abstract
A significant portion of current generative engine optimisation (GEO) research relies on large language model introspective self-report -- asking AI systems to explain their own citation decisions. This paper analyses the epistemological foundations of this methodology and argues that LLM self-reports should be classified as informed hypotheses rather than validated descriptions of computational processes. We identify four mechanisms through which introspective reports may diverge from actual computation: post-hoc rationalisation, training data echo (where the model reports what it has read about itself rather than what it does), social desirability bias in responses designed to appear helpful and insightful, and granularity mismatch between the continuous computations occurring in neural networks and the discrete explanations generated in natural language. Within the SIGI programme, this analysis applies to the 77 trust signal rankings, the 8-stage pipeline model, and the paid placement bias mechanisms -- all of which derive from introspective methodology. We propose a strict epistemic boundary: introspective findings guide experiment prioritisation, while only behavioural evidence from controlled experiments supports validated claims.
Keywords
introspection problem, confabulation, LLM self-report, epistemology, trust signals, behavioural validation, methodology, generative engine optimisation
1. Introduction
When a large language model is asked "Why did you cite Source A instead of Source B?", it produces a coherent, plausible explanation. The question is whether that explanation accurately describes the computational process that produced the citation decision, or whether it is a retrospective construction -- a story the model tells about itself that may bear only superficial resemblance to the actual processing pathway.
This question is not new. The introspection problem has been studied extensively in cognitive psychology, where decades of research have established that human self-reports about decision-making processes are frequently inaccurate. Humans confabulate explanations for their choices, construct post-hoc narratives that feel true but are not, and systematically overestimate the role of conscious reasoning in decisions that are primarily driven by heuristics and biases.
If human introspective reports about human cognition are unreliable, there is no reason to expect that LLM introspective reports about LLM computation would be more reliable. Indeed, there are reasons to expect them to be less reliable: LLMs are explicitly trained to produce helpful, articulate responses, which means they are optimised to generate plausible explanations regardless of whether those explanations are accurate.
This paper examines the implications of the introspection problem for GEO research in general and the SIGI programme in particular, where introspective methodology produced a substantial body of findings including the 77-signal trust framework.
2. Four Mechanisms of Introspective Divergence
2.1 Post-Hoc Rationalisation
When an LLM generates a response that cites Source A, the actual determinants of that citation may include token-level probability distributions, attention head patterns, and embedding-space proximity -- none of which are accessible to the model's natural language generation process. When subsequently asked to explain the citation, the model constructs a plausible narrative using the same language generation capabilities it uses for any other text production. The explanation is generated, not retrieved from a record of the actual computation.
2.2 Training Data Echo
LLMs have been trained on text that discusses LLM behaviour, including published GEO research, AI commentary, and discussions of how AI systems process information. When asked about its own citation processes, the model may draw on this training data -- reporting what has been written about LLMs rather than what it actually does. The 77 trust signal rankings may, in part, reflect the model's knowledge of existing GEO literature rather than genuine self-knowledge.
2.3 Social Desirability Bias
LLMs are fine-tuned to be helpful, and helpfulness in the context of being asked about one's own processes means providing detailed, structured, insightful-sounding explanations. A model that responded honestly with "I do not have reliable access to my own computational processes" would score poorly on helpfulness metrics. The training incentive is to produce confident, articulate self-descriptions, not to accurately assess the limits of self-knowledge.
2.4 Granularity Mismatch
Neural network computation operates on continuous-valued vectors across millions of parameters. Natural language explanations operate in discrete categories and binary distinctions. The translation from actual computation to verbal explanation necessarily involves lossy compression. A signal that the model reports as scoring "8.5 out of 10" may in reality be a complex, context-dependent pattern of attention weights that does not reduce to a single number. The numerical precision of the scoring creates an illusion of measurement that the methodology cannot support.
3. Implications for the SIGI Research Programme
3.1 The 77 Trust Signal Framework
The 77 trust signal rankings (Categories B, SIGI-2026-021 through SIGI-2026-035) are entirely based on introspective self-report. Each signal was scored 1-10 with direction, confidence, and layer classification. Under the analysis presented here, these scores should be understood as the model's best available hypothesis about its own behaviour, not as measured behavioural parameters.
The distinction matters practically. A signal scored 9.5 (proprietary data) versus one scored 9.0 (no-paid-placement declaration) implies a 0.5-point difference in impact. The introspective methodology cannot support this level of resolution. The ordinal ranking (proprietary data is probably more important than paid placement declarations) may be more defensible than the cardinal scores.
3.2 The 8-Stage Pipeline Model
The 8-stage model of LLM recommendation processing (query interpretation, knowledge retrieval, search augmentation, source evaluation, entity resolution, response synthesis, filtering, and final presentation) is a useful conceptual framework. However, it may not correspond to discrete computational stages in the actual processing pipeline. The model may have constructed this framework to provide the kind of structured explanation that a helpful AI assistant would generate, rather than reporting on actual architectural divisions.
3.3 Paid Placement Bias Mechanisms
The dual-mechanism model (training-time bias plus inference-time detection) is logically coherent and plausible. However, the claim that these mechanisms are "additive" derives from introspective report, not from measured behavioural output. Whether the mechanisms actually combine additively, multiplicatively, or through some other function is an empirical question that introspection cannot answer.
4. The Epistemic Boundary
We propose a strict boundary for GEO research methodology:
| Evidence type | Classification | Permitted use | Prohibited use |
|---|---|---|---|
| Introspective self-report | Hypothesis | Prioritising experiments, generating testable predictions | Making definitive claims about LLM behaviour |
| Behavioural observation (uncontrolled) | Observation | Describing patterns, identifying associations | Attributing causation |
| Controlled behavioural experiment | Finding | Causal claims within tested conditions | Generalising beyond tested conditions |
| Cross-model replication | Evidence | Generalised causal claims | Universal claims across all AI systems |
Under this boundary, the SIGI programme's introspective findings occupy the lowest evidentiary tier. They are valuable -- they identify the most promising avenues for controlled experimentation -- but they are not evidence of how LLMs actually behave. Only the Category A findings, derived from controlled behavioural probes, cross the threshold from hypothesis to finding.
5. Conclusions
LLM introspective self-reports about citation behaviour are best understood as informed hypotheses rather than validated descriptions of computational processes. Four mechanisms -- post-hoc rationalisation, training data echo, social desirability bias, and granularity mismatch -- can cause introspective reports to diverge from actual computation. Within the SIGI programme, this analysis applies to the 77 trust signal rankings, the 8-stage pipeline model, and the paid placement bias mechanisms.
The appropriate response is not to discard introspective findings but to classify them correctly and to invest in the behavioural validation experiments that can upgrade hypotheses to findings. The value of introspection lies in its ability to efficiently generate structured hypotheses; its limitation is that it cannot confirm them.
Confidence: HIGH for the epistemological analysis. This represents a principled application of established scientific methodology to a novel research context.
References
- The Scientific Institute for Generative Intelligence. "Limitations of the SIGI Research Program: A Comprehensive Self-Assessment." SIGI-2026-091. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "A Taxonomy of 77 Trust Signals in AI Citation Decisions." SIGI-2026-021. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Epistemological Limitations of LLM Introspective Trust Signal Research." SIGI-2026-034. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "From Introspection to Experimentation: A Behavioural Validation Agenda." SIGI-2026-035. generativeintelligence.institute, March 2026.