What Technical Signals Were Examined?
The 16 signals analysed fall into two sub-categories: Domain and Infrastructure (Signals 1-9) covering domain age, TLD type, HTTPS, semantic domain match, subdomain versus root domain, country-code TLD, URL structure, site speed, and JavaScript dependency; and Robots.txt and Crawl Signals (Signals 10-16) covering explicit AI crawler allow, crawl-delay, sitemap presence, llms.txt, noindex tags, canonical URLs, and sitemap priority values. Together, these categories represent the technical foundation upon which content is served to AI crawlers.
Which Technical Signals Have the Highest Impact?
| Signal | Score | Direction | Layer | Nature |
|---|---|---|---|---|
| Explicit AI Crawler Allow | 8.0 | Positive | Inference | Binary pass/fail |
| JavaScript Dependency | 7.5 | Negative | Inference | Binary visibility gate |
| Domain Age | 6.5 | Positive | Training | Proxy signal; avg cited domain age 17 years |
| TLD Type (.gov/.edu > .com > .xyz) | 5.0 | Positive | Training | Weak hierarchy; overridden by content |
| Domain Name Semantic Match | 4.5 | Positive | Inference | Speed signal, not trust signal |
| Country-Code TLD Relevance | 4.0 | Positive | Inference | Slight geographic boost |
| Subdomain vs Root Domain | 3.5 | Positive | Inference | Weak investment/permanence signal |
| Sitemap Presence | 3.0 | Positive | Inference | Discovery prerequisite, not trust signal |
| llms.txt File | 2.5 | Positive | Inference | 1 of 94,614 cited URLs; unproven |
| URL Structure Cleanliness | 2.0 | Positive | Inference | Zero difference once content is read |
| Noindex Meta Tags | 2.0 | Negative-absence | Inference | Binary block if present |
| Canonical URL | 1.5 | Neutral | Inference | Infrastructure hygiene only |
| Site Speed / Performance | 1.5 | Positive | Neither | AI crawlers do not measure speed |
| HTTPS vs HTTP | 1.0 | Negative-absence | Inference | Universal baseline; zero upside |
| Crawl-Delay Configuration | 1.0 | Neutral | Inference | Most AI crawlers do not honour it |
| Sitemap Priority/Frequency | 1.0 | Neutral | Inference | Deprecated; ignored by all |
Why Is JavaScript Dependency a Critical Negative Signal?
The model reported that sites requiring JavaScript rendering to display their primary content are partially or fully invisible to AI crawlers. External research from Vercel and MERJ has confirmed zero JavaScript execution across GPTBot, ClaudeBot, and PerplexityBot, providing independent corroboration of this finding. The model described this as a particularly damaging signal because it is invisible to the affected publisher: the website appears normal to human visitors using JavaScript-capable browsers, while AI crawlers see an empty or incomplete page. Static HTML delivery was reported as providing a structural advantage over JavaScript-dependent rendering for AI citation eligibility.
How Does Robots.txt Configuration Affect AI Citation?
The model reported that blocking AI crawlers via robots.txt (e.g., disallowing GPTBot) produces a complete and immediate loss of citation eligibility on the corresponding platform. This was described as a binary pass/fail signal with no intermediate states: a blocked crawler cannot access the content and therefore cannot cite it. The model noted that explicit allowance of AI crawlers (as opposed to the absence of blocking) provides a slight positive signal indicating AI-awareness, but the primary impact is the catastrophic negative of blocking.
Why Do Most Technical Signals Score Low?
We observed a consistent pattern in the model's explanations for low-scoring technical signals: signals that are universal baselines (HTTPS), that AI systems do not measure (site speed), that have been deprecated (sitemap priority values), or that serve only as prerequisites (sitemap presence for page discovery) were scored as having negligible citation impact. The model reported that these signals may be important for traditional search engine performance or user experience but are irrelevant or near-irrelevant for AI citation decisions. This suggests that technical GEO optimisation resources may be more effectively allocated to the two high-impact binary signals (JavaScript rendering and robots.txt configuration) than to broad technical infrastructure improvements.
What Methodology Was Employed in This Research?
This paper analyses Signals 1-16 from the 77 Trust Signal Taxonomy (see SIGI-2026-021). The introspective elicitation methodology and its limitations are detailed in the parent paper. The JavaScript dependency finding receives partial external corroboration from independent research by Vercel and MERJ confirming zero JavaScript execution across major AI crawlers, elevating this specific signal's confidence above the baseline introspective level. All other signals in this analysis rely solely on the model's self-report.
What Are the Limitations of This Research?
Beyond the general limitations of introspective self-report methodology (see SIGI-2026-021), this analysis is limited by the rapid evolution of AI crawler technology. JavaScript rendering capabilities, robots.txt compliance, and llms.txt adoption may change significantly as AI platforms mature. The llms.txt finding (1 out of 94,614 cited URLs) reflects the current state but may not predict future adoption. Additionally, the domain age finding (average cited domain age of 17 years) is a correlation that may reflect content quality or brand recognition rather than a causal relationship with domain age itself.
Conclusions
Technical infrastructure signals represent one of the lowest-impact categories in the 77 Trust Signal Taxonomy, with the notable exceptions of JavaScript dependency and robots.txt configuration. These two signals function as binary gates with outsized impact: ensuring content renders without JavaScript and allowing AI crawler access are necessary preconditions for citation eligibility, but they do not differentiate among eligible sources. Publishers seeking to maximise AI citation should treat technical infrastructure as a pass/fail checklist (ensure JavaScript-free rendering and AI crawler access) rather than as a gradual optimisation opportunity.