What Is the Dual-Layer Model of Signal Processing?

The model described two distinct processing layers for trust signals. Training-time signals are baked into the model's weights during pre-training and cannot be changed by page content. When the model encounters a domain it has seen frequently in high-quality training data, it applies a baseline trust level derived from training-time knowledge. This trust level is fixed between training cycles and cannot be overridden by inference-time content. Inference-time signals are evaluated in real-time when the model processes a retrieved document during a RAG pipeline. These signals are assessed based on the content, structure, and metadata of the specific page being evaluated for citation.

A small number of signals were classified as operating at both layers. Known paid directory status, for example, was described as having a training-time component (the model knows from training data that certain platforms involve paid elements) and an inference-time component (the model detects sponsored labels on the page). These dual-layer signals can produce additive effects as described in SIGI-2026-025.

How Are the Top 10 Signals Distributed Across Layers?

RankSignalScoreLayerControllable?
1Proprietary Data9.5InferenceYes (produce research)
2No-Paid-Placement Declaration9.0InferenceYes (add declaration)
3Direct Answer First8.5InferenceYes (restructure content)
4Statistics and Numbers8.5InferenceYes (add data)
5Methodology Section8.5InferenceYes (add methodology)
6Question-Format H2s8.0InferenceYes (rewrite headings)
7Entity Density >15%8.0InferenceYes (increase specificity)
8AI Crawler Allow8.0InferenceYes (update robots.txt)
9Verified Reviews8.0InferencePartially (earn reviews)
10Multi-Platform Presence8.0InferenceYes (create profiles)
We observed that 9 of 10 highest-impact signals operate at the inference-time layer and are directly controllable through page-level content and configuration decisions. This suggests that AI citation behaviour may be substantially more controllable than current GEO practice assumes.

What Are the Key Training-Time Signals?

The model identified several signals that operate primarily at the training-time layer and therefore cannot be influenced through page content: domain age (6.5), TLD type hierarchy (5.0), known paid directory status (7.0 negative), and brand recognition. These signals represent the model's prior knowledge about entities and domains, acquired during pre-training on web-scale data. A new business with a recently registered domain on a less established TLD starts at a training-time disadvantage that cannot be overcome through content optimisation alone but can be compensated for through strong inference-time signals.

What Are the Implications for GEO Strategy?

If the dual-layer model is accurate, it suggests a clear strategic prioritisation: invest first in inference-time signals (which are directly controllable and include the highest-impact entries) and accept training-time signals as fixed constraints that will improve gradually over time. This contradicts a common assumption in the GEO field that brand authority and domain age are the primary drivers of AI citation, suggesting instead that these factors, while real, are lower-impact than page-level content and structure decisions.

What Methodology Was Employed in This Research?

This paper analyses the layer classifications from the complete 77 Trust Signal Taxonomy (see SIGI-2026-021). Each signal was classified by the model as training-time, inference-time, or both during the structured introspective elicitation. The dual-layer model is a self-reported framework describing the model's stated understanding of its own processing, not a validated architectural description. Whether LLMs actually process signals through distinct layers is a mechanistic claim that requires architectural validation through interpretability research, which was beyond the scope of this study.

What Are the Limitations of This Research?

The dual-layer classification is a conceptual framework based on the model's self-report, not a validated description of model architecture. LLMs process all input through the same transformer architecture; the "layers" described are functional abstractions rather than architectural components. The model may be imposing a familiar cognitive framework (prior knowledge versus real-time assessment) onto processes that do not actually operate in this manner. Additionally, the controllability assessments assume that implementing inference-time signals produces measurable citation changes, which has not been experimentally confirmed for most signals. The framework's primary value is as a prioritisation heuristic rather than an architectural description.

Conclusions

The dual-layer model proposes a useful framework for understanding and prioritising GEO efforts: signals that are evaluated at inference-time (when the model processes retrieved documents) are both higher-impact and more controllable than signals embedded in training-time knowledge. If validated, this framework suggests that publishers can achieve substantial citation improvements through content-level optimisation without waiting for domain authority or brand recognition to accumulate. The framework's value as a practical prioritisation tool is independent of its accuracy as an architectural description, but experimental validation of the controllability claim is essential for evidence-based GEO strategy.