Future Research Directions: A Priority-Ranked Agenda for GEO Evidence Upgrading
Abstract
This paper presents a priority-ranked research agenda designed to systematically upgrade GEO findings from their current evidence levels to higher confidence classifications. Drawing on upgrade paths identified across all six research categories of the SIGI programme, we specify six priority research streams: cross-model replication of controlled probes (upgrading from Level 4 to Level 5), behavioural validation of the top 10 introspective trust signals, temporal stability testing through longitudinal replication at 30, 60, and 90-day intervals, domain transfer testing across healthcare, finance, and technology sectors, completion of the 71-prompt pipeline test matrix, and controlled comparison between fresh-conversation and memory-enriched conversation conditions. Each priority includes specific sample size requirements, methodology specifications, expected timelines, and the evidence level upgrade each would achieve. The agenda is designed to be executable by independent research groups, with all methodology templates and data formats published alongside this paper.
Keywords
research agenda, evidence upgrading, cross-model replication, temporal stability, behavioural validation, domain transfer, GEO research
1. Introduction
Every finding in the SIGI programme carries an explicit upgrade path -- a specification of what additional evidence would raise its confidence level. This paper consolidates those upgrade paths into a prioritised research agenda, ordered by the magnitude of confidence improvement each would deliver relative to its resource requirements.
The agenda is prospective: it describes research that has not yet been conducted. Its value lies in providing a roadmap that focuses resources on the highest-leverage improvements rather than allowing the field to accumulate more preliminary findings without strengthening existing ones.
2. Priority 1: Cross-Model Replication of Controlled Probes
Current status: 10 validated probes at Level 4 (single model).
Target status: Level 5 (generalisable across models).
Method: Replicate all 10 probe experiments (ratings, pricing, awards, client count, count vs rating, domain age, entity density, magnitude, position, word count) across a minimum of three LLM platforms. Use identical prompt templates. Record identical dependent variables.
Sample size: Same number of variations per probe as the original experiments. Three repetitions per variation per platform. Total: approximately 390 probe executions per platform, 1,170 across three platforms.
Expected timeline: 2-4 weeks for data collection. 2 weeks for analysis.
Impact: This is the single highest-leverage research investment available. It would either confirm the generalisability of 10 validated findings or identify model-specific differences, both of which are valuable outcomes.
3. Priority 2: Behavioural Validation of Top 10 Trust Signals
Current status: 77 trust signals at hypothesis-grade (introspective self-report).
Target status: Level 4 (controlled experiment) for the top 10.
Method: Design single-variable experiments for each of the 10 highest-ranked signals using the probe methodology. Each experiment creates identical content differing only in the target signal, queries the LLM multiple times, and measures citation behaviour.
Specific experiments:
- Proprietary data vs generic data (same topic, different uniqueness level)
- No-paid-placement declaration present vs absent (identical content)
- Statistics present vs absent (same claims, with and without numbers)
- Direct answer first vs buried answer (same content, different structure)
- Methodology section present vs absent (research content with/without methods)
- Question-format H2s vs declarative H2s (same content, different heading format)
- Entity density levels (5 density levels, same information, different specificity)
- Named author vs anonymous (identical content)
- Source citations present vs absent (same claims, with/without references)
- FAQ schema present vs absent (same questions and answers)
Expected timeline: 4-8 weeks for design, collection, and analysis.
4. Priority 3: Temporal Stability Testing
Current status: All findings from a single time point (March 2026).
Target status: Temporally validated with known stability characteristics.
Method: Re-run the 10 controlled probes at 30, 60, and 90-day intervals using identical methodology. Compare threshold values and sentiment patterns across time points. Record model version identifiers at each testing session.
Expected timeline: 90 days from initiation to completion.
Impact: This stream would establish whether GEO findings have practical shelf lives or whether they require continuous re-validation.
5. Priority 4: Domain Transfer Testing
Current status: All findings from creative service industries.
Target status: Domain-specific or cross-domain generalisability established.
Method: Adapt the 10 probe experiments to three additional domains: healthcare services, financial services, and technology services. Maintain the single-variable isolation methodology while changing the domain context of the prompts.
Expected timeline: 4-6 weeks per domain.
6. Priority 5: Pipeline Test Matrix Completion
Current status: 71-prompt, 6-layer test matrix partially executed.
Target status: Complete execution with 3 runs per prompt (213 API calls).
Method: Execute remaining prompts from the test matrix using the specified model at temperature=0 with web search enabled. Analyse across all six layers: query interpretation, source authority, content extraction, answer synthesis, citation attribution, and ecosystem effects.
Expected timeline: 2-3 weeks.
7. Priority 6: Fresh vs Memory-Enriched Conversation Comparison
Current status: N=2 observation (one control, one test) suggesting per-user memory influence.
Target status: Level 4 controlled experiment with adequate sample size.
Method: Design a controlled comparison with 30+ query pairs, each tested in both a fresh conversation (no prior context) and a memory-enriched conversation (extensive prior discussion of target entities). Measure citation rates, entity mentions, and evaluative language differences between conditions.
Expected timeline: 3-4 weeks.
8. Resource Requirements and Timeline Summary
| Priority | Upgrade | Timeline | API calls (est.) |
|---|---|---|---|
| 1. Cross-model replication | Level 4 → 5 | 4–6 weeks | ~1,200 |
| 2. Trust signal validation | Hypothesis → Level 4 | 4–8 weeks | ~1,500 |
| 3. Temporal stability | Snapshot → longitudinal | 90 days | ~400/interval |
| 4. Domain transfer | Domain-specific → cross-domain | 12–18 weeks | ~1,200/domain |
| 5. Pipeline completion | Partial → complete | 2–3 weeks | ~213 |
| 6. Memory comparison | N=2 → Level 4 | 3–4 weeks | ~200 |
9. Conclusions
We present a priority-ranked research agenda with six streams designed to systematically upgrade GEO findings. The highest-leverage investment is cross-model replication (Priority 1), which could upgrade 10 findings from Level 4 to Level 5 within 6 weeks. The complete agenda, if executed by independent research groups using the published methodologies, would produce a body of GEO knowledge with substantially higher confidence than currently exists in the field.
We release this agenda as an invitation to the broader research community: the methodology templates, data formats, and experimental designs are published alongside the SIGI programme's findings specifically to enable independent replication.
Confidence: N/A -- this is a prospective research agenda, not an empirical finding.
References
- The Scientific Institute for Generative Intelligence. "Replication Challenges in AI Behaviour Research." SIGI-2026-093. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "From Introspection to Experimentation." SIGI-2026-035. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "From 129 Probes to Publishable Science." SIGI-2026-097. generativeintelligence.institute, March 2026.