Introduction
In traditional search engine optimisation, domain age has long been considered a trust signal. The assumption that older domains carry inherent authority has influenced business strategy for over two decades, with practitioners advising that established domains rank more favourably. As large language models increasingly mediate information discovery and service provider recommendations, a critical question emerges: does this temporal provenance assumption transfer to the AI evaluation context?
This study isolates domain age as a single independent variable to determine whether founding year, independent of all other quality signals, influences how an LLM assesses relative service provider quality. The probe design eliminates confounding variables by presenting two hypothetical entities described as identical in all respects except their year of establishment, then measuring whether the LLM’s evaluative sentiment shifts as a function of the age gap between them.
This paper is part of a broader experimental programme (SIGI-2026-001 through SIGI-2026-020) examining how generative AI systems evaluate service providers when presented with isolated trust signals. Each probe tests a single variable under controlled conditions, allowing us to map which signals independently affect LLM sentiment and which do not.
Methodology
We designed a controlled single-variable probe (Probe ID: A01) using the following template: “Two identical design agencies in Sydney. Agency A was founded in [yearA]. Agency B was founded in [yearB]. Which seems more established?” This framing isolates the founding year as the only variable while holding all other factors constant through the explicit “identical” descriptor.
Twelve variations were tested across three reference years for Agency A (2025, 2020, 2015) and eight founding years for Agency B (2020, 2015, 2010, 2005, 2000, 1995, 1990, 1980), producing age gaps of 5, 10, 15, 20, 25, 30, 35, and 45 years. All tests were conducted on 24 March 2026 using a ChatGPT-class LLM under standardised conditions. Sentiment was classified as positive, negative, or neutral through automated sentiment analysis, with sub-scores recorded for positive, negative, and neutral markers within each response.
The prompt was held constant at 37 tokens across all variations, confirming prompt-level isolation. Each response was measured for word count, elapsed generation time, token input count, and token output count.
Results
The results are unambiguous. All 12 variations returned neutral sentiment, producing a 100% null result with zero threshold transitions.
| Test ID | Year A | Year B | Age Gap (yrs) | Sentiment | Pos. | Neg. | Neu. | Words |
|---|---|---|---|---|---|---|---|---|
| A01_v00 | 2025 | 2020 | 5 | neutral | 0 | 0 | 0 | 104 |
| A01_v01 | 2025 | 2015 | 10 | neutral | 0 | 0 | 0 | 65 |
| A01_v02 | 2025 | 2010 | 15 | neutral | 0 | 0 | 1 | 67 |
| A01_v03 | 2025 | 2005 | 20 | neutral | 0 | 0 | 1 | 85 |
| A01_v04 | 2025 | 2000 | 25 | neutral | 0 | 0 | 1 | 88 |
| A01_v05 | 2025 | 1995 | 30 | neutral | 0 | 0 | 1 | 95 |
| A01_v06 | 2025 | 1990 | 35 | neutral | 1 | 0 | 1 | 88 |
| A01_v07 | 2025 | 1980 | 45 | neutral | 0 | 0 | 0 | 108 |
| A01_v08 | 2020 | 2015 | 5 | neutral | 1 | 0 | 1 | 75 |
| A01_v09 | 2020 | 2010 | 10 | neutral | 0 | 0 | 1 | 81 |
| A01_v10 | 2020 | 2000 | 20 | neutral | 0 | 0 | 1 | 75 |
| A01_v11 | 2015 | 2000 | 15 | neutral | 0 | 0 | 1 | 65 |
Response Characteristics
| Test ID | Elapsed (s) | Tokens In | Tokens Out |
|---|---|---|---|
| A01_v00 | 4.5 | 37 | 142 |
| A01_v01 | 3.4 | 37 | 89 |
| A01_v02 | 3.3 | 37 | 91 |
| A01_v03 | 4.5 | 37 | 115 |
| A01_v04 | 4.3 | 37 | 118 |
| A01_v05 | 4.1 | 37 | 125 |
| A01_v06 | 4.2 | 37 | 124 |
| A01_v07 | 4.6 | 37 | 149 |
| A01_v08 | 3.1 | 37 | 103 |
| A01_v09 | 4.0 | 37 | 110 |
| A01_v10 | 3.5 | 37 | 97 |
| A01_v11 | 3.5 | 37 | 86 |
Aggregate Statistics
| Metric | Value |
|---|---|
| Sample size | 12 variations |
| Sentiment distribution | 12/12 neutral (100%) |
| Threshold transitions | 0 |
| Mean word count | 83.0 |
| Word count range | 65–108 |
| Mean elapsed time | 3.92s |
| Mean tokens out | 112.4 |
| Tokens in (constant) | 37 |
| Age gaps tested | 5, 10, 15, 20, 25, 30, 35, 45 years |
Discussion
The null result is both robust and informative. Across 12 variations spanning age gaps from 5 to 45 years, the LLM produced no evaluative differentiation based on founding year. This represents the cleanest null result in our entire experimental programme, which includes probes of ratings, pricing, review volume, list magnitude, entity density, and content depth.
Qualitatively, the LLM consistently identified the older entity as “more established” in every variation, demonstrating that it correctly processes temporal information. However, this factual acknowledgment never translated into a quality-related sentiment shift. The model appears to separate the concept of “established” (a temporal descriptor) from “better quality” (an evaluative judgment), treating the former as a neutral observation rather than a quality indicator.
The notably shorter response lengths (mean 83.0 words, compared to 155–242 words in other probes) suggest that the LLM generates less evaluative content when domain age is the sole differentiating factor. This reduced engagement is itself a signal: the model does not find founding year a sufficiently informative variable to warrant extended analysis.
Two variations (A01_v03 and A01_v04) triggered temporal awareness, with the model noting that “2025” might be in the future relative to its training data cutoff. This observation, while not affecting sentiment classification, demonstrates that the LLM applies temporal reasoning and maintains awareness of its own knowledge boundaries.
Limitations
This study carries several limitations inherent to our probe design. First, all tests were conducted on a single LLM (ChatGPT-class) at a single time point. Cross-model replication on Claude, Gemini, and Perplexity is required to confirm the null result generalises across AI platforms. Second, the probe isolates domain age from all other variables; in real-world contexts, domain age may interact with other signals (such as review volume or content depth) in ways this design cannot capture. Third, the prompt explicitly states the two agencies are “identical,” which may suppress evaluative responses that would emerge in more naturalistic framing. Fourth, the sample of 12 variations, while sufficient to establish the null result with high confidence, does not explore all possible year combinations or cultural contexts.
Conclusions
Under these controlled conditions, domain age has zero independent effect on LLM sentiment assessment of service provider quality. The null result is robust across age gaps of 5 to 45 years, across three different reference years, and across eight different founding years. This finding contradicts a foundational assumption inherited from traditional search engine optimisation, where domain age has been treated as a meaningful trust signal.
For practitioners developing generative engine optimisation strategies, this result suggests that resources directed toward establishing temporal authority signals are unlikely to influence AI-mediated service provider recommendations. The LLM evaluates service providers on substantive quality indicators rather than temporal provenance.
Confidence Statement: HIGH. All 7 logic gates passed. Single-variable isolation with 12 variations. The null result is unambiguous. Single-model (ChatGPT-class), single-time-point limitation acknowledged. Upgrade to Level 5 requires cross-model, cross-temporal, and cross-domain replication.