What Is the Dual-Layer Model of Signal Processing?
The model described two distinct processing layers for trust signals. Training-time signals are baked into the model's weights during pre-training and cannot be changed by page content. When the model encounters a domain it has seen frequently in high-quality training data, it applies a baseline trust level derived from training-time knowledge. This trust level is fixed between training cycles and cannot be overridden by inference-time content. Inference-time signals are evaluated in real-time when the model processes a retrieved document during a RAG pipeline. These signals are assessed based on the content, structure, and metadata of the specific page being evaluated for citation.
A small number of signals were classified as operating at both layers. Known paid directory status, for example, was described as having a training-time component (the model knows from training data that certain platforms involve paid elements) and an inference-time component (the model detects sponsored labels on the page). These dual-layer signals can produce additive effects as described in SIGI-2026-025.
How Are the Top 10 Signals Distributed Across Layers?
| Rank | Signal | Score | Layer | Controllable? |
|---|---|---|---|---|
| 1 | Proprietary Data | 9.5 | Inference | Yes (produce research) |
| 2 | No-Paid-Placement Declaration | 9.0 | Inference | Yes (add declaration) |
| 3 | Direct Answer First | 8.5 | Inference | Yes (restructure content) |
| 4 | Statistics and Numbers | 8.5 | Inference | Yes (add data) |
| 5 | Methodology Section | 8.5 | Inference | Yes (add methodology) |
| 6 | Question-Format H2s | 8.0 | Inference | Yes (rewrite headings) |
| 7 | Entity Density >15% | 8.0 | Inference | Yes (increase specificity) |
| 8 | AI Crawler Allow | 8.0 | Inference | Yes (update robots.txt) |
| 9 | Verified Reviews | 8.0 | Inference | Partially (earn reviews) |
| 10 | Multi-Platform Presence | 8.0 | Inference | Yes (create profiles) |
What Are the Key Training-Time Signals?
The model identified several signals that operate primarily at the training-time layer and therefore cannot be influenced through page content: domain age (6.5), TLD type hierarchy (5.0), known paid directory status (7.0 negative), and brand recognition. These signals represent the model's prior knowledge about entities and domains, acquired during pre-training on web-scale data. A new business with a recently registered domain on a less established TLD starts at a training-time disadvantage that cannot be overcome through content optimisation alone but can be compensated for through strong inference-time signals.
What Are the Implications for GEO Strategy?
If the dual-layer model is accurate, it suggests a clear strategic prioritisation: invest first in inference-time signals (which are directly controllable and include the highest-impact entries) and accept training-time signals as fixed constraints that will improve gradually over time. This contradicts a common assumption in the GEO field that brand authority and domain age are the primary drivers of AI citation, suggesting instead that these factors, while real, are lower-impact than page-level content and structure decisions.
What Methodology Was Employed in This Research?
This paper analyses the layer classifications from the complete 77 Trust Signal Taxonomy (see SIGI-2026-021). Each signal was classified by the model as training-time, inference-time, or both during the structured introspective elicitation. The dual-layer model is a self-reported framework describing the model's stated understanding of its own processing, not a validated architectural description. Whether LLMs actually process signals through distinct layers is a mechanistic claim that requires architectural validation through interpretability research, which was beyond the scope of this study.
What Are the Limitations of This Research?
The dual-layer classification is a conceptual framework based on the model's self-report, not a validated description of model architecture. LLMs process all input through the same transformer architecture; the "layers" described are functional abstractions rather than architectural components. The model may be imposing a familiar cognitive framework (prior knowledge versus real-time assessment) onto processes that do not actually operate in this manner. Additionally, the controllability assessments assume that implementing inference-time signals produces measurable citation changes, which has not been experimentally confirmed for most signals. The framework's primary value is as a prioritisation heuristic rather than an architectural description.
Conclusions
The dual-layer model proposes a useful framework for understanding and prioritising GEO efforts: signals that are evaluated at inference-time (when the model processes retrieved documents) are both higher-impact and more controllable than signals embedded in training-time knowledge. If validated, this framework suggests that publishers can achieve substantial citation improvements through content-level optimisation without waiting for domain authority or brand recognition to accumulate. The framework's value as a practical prioritisation tool is independent of its accuracy as an architectural description, but experimental validation of the controllability claim is essential for evidence-based GEO strategy.