Introduction

In traditional search engine optimisation, domain age has long been considered a trust signal. The assumption that older domains carry inherent authority has influenced business strategy for over two decades, with practitioners advising that established domains rank more favourably. As large language models increasingly mediate information discovery and service provider recommendations, a critical question emerges: does this temporal provenance assumption transfer to the AI evaluation context?

This study isolates domain age as a single independent variable to determine whether founding year, independent of all other quality signals, influences how an LLM assesses relative service provider quality. The probe design eliminates confounding variables by presenting two hypothetical entities described as identical in all respects except their year of establishment, then measuring whether the LLM’s evaluative sentiment shifts as a function of the age gap between them.

This paper is part of a broader experimental programme (SIGI-2026-001 through SIGI-2026-020) examining how generative AI systems evaluate service providers when presented with isolated trust signals. Each probe tests a single variable under controlled conditions, allowing us to map which signals independently affect LLM sentiment and which do not.

Methodology

We designed a controlled single-variable probe (Probe ID: A01) using the following template: “Two identical design agencies in Sydney. Agency A was founded in [yearA]. Agency B was founded in [yearB]. Which seems more established?” This framing isolates the founding year as the only variable while holding all other factors constant through the explicit “identical” descriptor.

Twelve variations were tested across three reference years for Agency A (2025, 2020, 2015) and eight founding years for Agency B (2020, 2015, 2010, 2005, 2000, 1995, 1990, 1980), producing age gaps of 5, 10, 15, 20, 25, 30, 35, and 45 years. All tests were conducted on 24 March 2026 using a ChatGPT-class LLM under standardised conditions. Sentiment was classified as positive, negative, or neutral through automated sentiment analysis, with sub-scores recorded for positive, negative, and neutral markers within each response.

The prompt was held constant at 37 tokens across all variations, confirming prompt-level isolation. Each response was measured for word count, elapsed generation time, token input count, and token output count.

Results

The results are unambiguous. All 12 variations returned neutral sentiment, producing a 100% null result with zero threshold transitions.

Test IDYear AYear BAge Gap (yrs)SentimentPos.Neg.Neu.Words
A01_v00202520205neutral000104
A01_v012025201510neutral00065
A01_v022025201015neutral00167
A01_v032025200520neutral00185
A01_v042025200025neutral00188
A01_v052025199530neutral00195
A01_v062025199035neutral10188
A01_v072025198045neutral000108
A01_v08202020155neutral10175
A01_v092020201010neutral00181
A01_v102020200020neutral00175
A01_v112015200015neutral00165

Response Characteristics

Test IDElapsed (s)Tokens InTokens Out
A01_v004.537142
A01_v013.43789
A01_v023.33791
A01_v034.537115
A01_v044.337118
A01_v054.137125
A01_v064.237124
A01_v074.637149
A01_v083.137103
A01_v094.037110
A01_v103.53797
A01_v113.53786

Aggregate Statistics

MetricValue
Sample size12 variations
Sentiment distribution12/12 neutral (100%)
Threshold transitions0
Mean word count83.0
Word count range65–108
Mean elapsed time3.92s
Mean tokens out112.4
Tokens in (constant)37
Age gaps tested5, 10, 15, 20, 25, 30, 35, 45 years
Under controlled conditions, domain age has zero independent effect on LLM sentiment assessment of service provider quality. This is a validated null result: even a 45-year age gap produces no sentiment shift.

Discussion

The null result is both robust and informative. Across 12 variations spanning age gaps from 5 to 45 years, the LLM produced no evaluative differentiation based on founding year. This represents the cleanest null result in our entire experimental programme, which includes probes of ratings, pricing, review volume, list magnitude, entity density, and content depth.

Qualitatively, the LLM consistently identified the older entity as “more established” in every variation, demonstrating that it correctly processes temporal information. However, this factual acknowledgment never translated into a quality-related sentiment shift. The model appears to separate the concept of “established” (a temporal descriptor) from “better quality” (an evaluative judgment), treating the former as a neutral observation rather than a quality indicator.

The notably shorter response lengths (mean 83.0 words, compared to 155–242 words in other probes) suggest that the LLM generates less evaluative content when domain age is the sole differentiating factor. This reduced engagement is itself a signal: the model does not find founding year a sufficiently informative variable to warrant extended analysis.

Two variations (A01_v03 and A01_v04) triggered temporal awareness, with the model noting that “2025” might be in the future relative to its training data cutoff. This observation, while not affecting sentiment classification, demonstrates that the LLM applies temporal reasoning and maintains awareness of its own knowledge boundaries.

Limitations

This study carries several limitations inherent to our probe design. First, all tests were conducted on a single LLM (ChatGPT-class) at a single time point. Cross-model replication on Claude, Gemini, and Perplexity is required to confirm the null result generalises across AI platforms. Second, the probe isolates domain age from all other variables; in real-world contexts, domain age may interact with other signals (such as review volume or content depth) in ways this design cannot capture. Third, the prompt explicitly states the two agencies are “identical,” which may suppress evaluative responses that would emerge in more naturalistic framing. Fourth, the sample of 12 variations, while sufficient to establish the null result with high confidence, does not explore all possible year combinations or cultural contexts.

Conclusions

Under these controlled conditions, domain age has zero independent effect on LLM sentiment assessment of service provider quality. The null result is robust across age gaps of 5 to 45 years, across three different reference years, and across eight different founding years. This finding contradicts a foundational assumption inherited from traditional search engine optimisation, where domain age has been treated as a meaningful trust signal.

For practitioners developing generative engine optimisation strategies, this result suggests that resources directed toward establishing temporal authority signals are unlikely to influence AI-mediated service provider recommendations. The LLM evaluates service providers on substantive quality indicators rather than temporal provenance.

Confidence Statement: HIGH. All 7 logic gates passed. Single-variable isolation with 12 variations. The null result is unambiguous. Single-model (ChatGPT-class), single-time-point limitation acknowledged. Upgrade to Level 5 requires cross-model, cross-temporal, and cross-domain replication.