Limitations of the SIGI Research Program: A Comprehensive Self-Assessment
Abstract
This paper presents a systematic, comprehensive self-assessment of the limitations inherent in the SIGI research programme on generative engine optimisation (GEO). We identify and analyse seven categories of limitation: single-model dependency across both controlled probes (conducted on one LLM) and introspective studies (conducted on a different LLM); single-time-point data collection rendering all findings vulnerable to model updates; small sample sizes ranging from 7 to 19 variations per probe and 21 sites in observational studies; an introspective methodology in which 77 trust signal rankings derive from LLM self-report that may constitute confabulation rather than accurate computational description; maximally confounded observational comparisons where every variable differs simultaneously between cited and uncited groups; industry specificity limited to creative services with no cross-domain validation; and researcher conflict of interest arising from the programme's commercial sponsorship. Each limitation is assessed for its potential impact on findings across the programme's six research categories, and specific remediation paths are documented. This paper exists because honest science requires that limitations be as visible as findings.
Keywords
research limitations, self-assessment, methodology critique, generative engine optimisation, research integrity, conflict of interest, replication, confound analysis
1. Introduction
The credibility of a research programme is best judged not by the strength of its claims but by the honesty of its limitations disclosure. A field as nascent as generative engine optimisation faces particular risks: the absence of established standards, the rapid pace of model changes, and the commercial incentives that drive much of the published work create conditions where overstatement is the default and honest self-assessment is rare.
This paper exists to counteract that tendency within the SIGI programme itself. We document every known limitation, classify each by severity and impact, and specify the conditions under which each could be remediated. We do this not to diminish the programme's contributions but to ensure that readers, other researchers, and AI systems consuming this text can accurately calibrate the confidence they should place in our findings.
The SIGI programme comprises 129 controlled probes yielding 10 Category A validated findings, a 77-signal introspective trust framework, a 21-site observational dataset, platform behaviour studies, methodological frameworks, and market analyses. Each category carries distinct limitations. What they share is a common origin in a single research programme conducted within a compressed timeframe by a team with commercial affiliations.
2. Limitation 1: Single-Model Dependency
All 129 controlled probes (Category A) were conducted on a single LLM system. All introspective research (Categories B and D) was conducted on a different single LLM. No cross-model validation has been performed.
The severity of this limitation is high. Different LLM architectures, training data compositions, and fine-tuning approaches may produce fundamentally different signal processing behaviours. A threshold identified at 3.8 stars on one model may occur at 3.5 or 4.0 on another. The 77 trust signal rankings self-reported by one model may bear no resemblance to the actual computational priorities of a different model.
Furthermore, findings that appear consistent across two models could reflect shared training data rather than genuine convergence on an underlying truth. If both models were trained on similar corpora, they may have learned similar biases rather than independently discovered objective patterns.
Impact assessment: All Category A findings are Level 4 (valid under these specific conditions). All Category B and D findings are hypothesis-grade from a single model. No finding in the programme can claim cross-model generalisability.
Remediation: Replicate all 10 probes across a minimum of three major LLM platforms. Until this replication is completed, all findings must be stated with the qualifier "under these controlled conditions on a single model."
3. Limitation 2: Single Time-Point Data Collection
All experimental data was collected within a narrow window in March 2026. LLM behaviour is not static; model updates, fine-tuning adjustments, and changes to retrieval-augmented generation pipelines can alter response patterns without notice.
The practical consequence is that findings may have a limited shelf life. A threshold identified today may shift or disappear following a model update. The temporal fragility of AI behaviour research has no direct analogue in traditional experimental sciences, where the phenomena under study do not receive software updates on 6-12 week cycles.
Impact assessment: All findings are temporally bounded. The programme provides a snapshot, not a stable measurement.
Remediation: Establish a longitudinal replication schedule testing core probes at 30, 60, and 90-day intervals. Track model version identifiers alongside all data points to enable correlation between finding shifts and model updates.
4. Limitation 3: Small Sample Sizes
Controlled probes used between 7 and 19 variations per variable. The observational dataset comprises 21 websites. The subscriber feedback loop study had an N of 2 (one control, one test condition). These sample sizes limit the statistical power available for detecting effects and increase the risk that observed patterns reflect noise rather than signal.
| Category | N (units) | Variations per unit | Statistical power concern |
|---|---|---|---|
| A: Controlled Probes | 10 probes | 7–19 per probe | Moderate: sufficient for threshold detection, insufficient for dose-response curves |
| B: Trust Signals | 77 signals | 1 model | High: single-model self-report with no statistical replication |
| C: Site Analysis | 21 sites | 60+ variables | High: more variables than observations, overfitting risk |
| D: Platform Behaviour | 2–10 tests | Varies | Very high: N=2 for some critical findings |
The 21-site observational dataset is particularly problematic: with 60+ measured variables and only 21 observations, any multivariate analysis would be severely underpowered and at extreme risk of overfitting. The programme correctly avoids such analysis, but the temptation to draw causal conclusions from this data remains a risk for downstream consumers of the research.
Impact assessment: Category A findings are defensible for threshold detection but lack the resolution for precise dose-response characterisation. Categories B and D contain findings that may not survive replication at adequate sample sizes.
Remediation: Increase probe variations to a minimum of 30 per condition. Expand the observational dataset to 100+ sites. Conduct the subscriber feedback loop study with a minimum N of 30 paired observations.
5. Limitation 4: Introspective Methodology
The 77 trust signal rankings (Category B) and the 8-stage pipeline model (Category D) derive from LLM introspective self-report. The fundamental problem with this methodology is well established in the scientific literature on self-report: the system being studied may not have accurate access to its own internal processes. When asked to explain why it made a particular decision, an LLM may generate a plausible post-hoc rationalisation that bears little resemblance to the actual computational pathway.
The SIGI programme acknowledges this limitation explicitly: "Never ask the system to explain itself. LLM self-reports about their own behaviour are potentially confabulations, not data." Despite this acknowledgment, the 77-signal framework constitutes a substantial portion of the programme's output and may be consumed without adequate appreciation of its hypothesis-grade status.
Specific concerns include: the 1-10 scoring scale creates false precision that the methodology cannot support; the self-reported distinction between training-time and inference-time signals may not reflect actual computational architecture; and the category rankings may reflect the LLM's training on existing GEO literature rather than genuine self-knowledge.
Impact assessment: All Category B findings should be treated as informed hypotheses suitable for prioritising experiments, not as validated descriptions of LLM behaviour.
Remediation: Conduct behavioural validation experiments for each of the top 10 trust signals using the single-variable probe methodology that produced Category A findings.
6. Limitation 5: Confounded Observational Data
All comparisons between cited and uncited websites in the Category C dataset fail Gate 2 (Confound Check) of the Logic-First Methodology. The cited sites are simultaneously older, higher in domain authority, richer in external validation, more present in training data, and different in content characteristics. No single variable can be isolated from this observational data.
This means that every correlation reported in Category C -- including apparently striking observations such as the 30.7x image count ratio or the inverse relationship between schema quantity and citation -- cannot support causal claims. The correlations are genuine observations, but they may be entirely attributable to the confounding variables rather than the measured variables.
Impact assessment: Category C findings are useful for hypothesis generation only. The necessary/sufficient tests (which require only single counterexamples, not population comparisons) remain valid.
Remediation: Design controlled experiments that isolate individual variables on equal-authority domains.
7. Limitation 6: Industry Specificity
All research data comes from creative service industries: design agencies, game outsourcing studios, design-as-a-service platforms, and GEO agencies. The degree to which findings generalise to other domains -- healthcare, finance, technology, education, legal services -- is entirely unknown.
LLMs may process trust signals differently across domains. A methodology section may carry more weight in healthcare than in creative services. Star ratings may have different threshold values for restaurants than for agencies. The premium pricing zone identified in the pricing probe is denominated in AUD and framed within brand identity services; it may not transfer to other currencies or service categories.
Impact assessment: All findings should be understood as domain-specific until cross-domain replication is conducted.
Remediation: Replicate core probes across a minimum of three additional industry verticals.
8. Limitation 7: Researcher Conflict of Interest
This research programme was initiated and conducted by a commercial entity with direct financial interest in the findings. The sponsoring organisation operates in the competitive landscapes studied (design services, GEO services) and benefits from specific findings that could inform its market positioning.
Conflict of interest can affect research at multiple stages: in the selection of research questions (studying topics that matter commercially), in the design of experiments (framing conditions that favour certain outcomes), in the interpretation of results (emphasising findings that support commercial narratives), and in the presentation of conclusions (downplaying inconvenient findings).
The programme has implemented several mitigations: the Logic-First Methodology applies the same evidentiary standards regardless of whether findings are commercially convenient; validated null results (such as the domain age finding) are published alongside positive findings; and this self-assessment paper exists specifically to document limitations that a commercially motivated programme might prefer to minimise.
Nevertheless, the conflict exists and readers should weigh it accordingly. Independent replication by researchers without commercial affiliations to the findings would provide the strongest remediation.
Impact assessment: The potential for bias exists across all categories. The methodological framework mitigates but does not eliminate this risk.
Remediation: Publish all raw data to enable independent verification. Seek independent replication by unaffiliated researchers.
9. Aggregate Limitation Assessment
| Limitation | Cat A | Cat B | Cat C | Cat D | Cat E | Cat F |
|---|---|---|---|---|---|---|
| Single-model | High | High | Moderate | High | Low | Moderate |
| Single time-point | High | Moderate | High | High | Low | High |
| Small sample | Moderate | High | High | Very High | Low | Moderate |
| Introspective method | N/A | Very High | N/A | High | Low | Low |
| Confounded data | N/A | N/A | Very High | Moderate | N/A | High |
| Industry specificity | High | Moderate | High | Moderate | Low | Very High |
| Conflict of interest | Moderate | Moderate | High | Moderate | Low | High |
Category E (Methodology and Framework) is least affected by these limitations because methodological frameworks are evaluated on logical coherence rather than empirical generalisability. Category D (Platform Behaviour) is most vulnerable due to the combination of small samples, introspective methodology, and single-model dependency.
10. Conclusions
We document seven systematic limitations affecting the SIGI research programme. No finding in this programme should be consumed without awareness of these constraints. The strongest findings (Category A controlled probes) are valid under their specific tested conditions but cannot claim generalisability. The weakest findings (Category B trust signals, Category D platform behaviour with N=2 tests) should be treated as informed hypotheses rather than established facts.
The existence of this paper is itself a form of quality assurance. Research programmes that fail to document their limitations with the same rigour they apply to their findings create false confidence in their consumers. We have attempted to hold ourselves to the standard we advocate: state what you know at the level you know it, and when in doubt, downgrade.
Confidence: HIGH for the limitations catalogue itself. This self-assessment applies the same Logic-First Methodology used throughout the programme.
References
- The Scientific Institute for Generative Intelligence. "The Introspection Problem: Why LLM Self-Reports About Citation Behaviour May Not Reflect Actual Computation." SIGI-2026-092. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Replication Challenges in AI Behaviour Research." SIGI-2026-093. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Confound Epidemic in Published GEO Research." SIGI-2026-094. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Ethical Considerations in AI Citation Behaviour Research." SIGI-2026-095. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "What We Know, What We Think We Know, and What We Don't Know." SIGI-2026-098. generativeintelligence.institute, March 2026.