Ethical Considerations in AI Citation Behaviour Research: Conflict of Interest, Manipulation Potential, and Transparency Requirements
Abstract
Research into how AI systems select and cite information sources raises ethical questions that the emerging GEO field has not yet systematically addressed. This paper examines six ethical dimensions: self-report bias in LLM introspective methodology, non-determinism requiring statistical rather than single-observation analysis, transparency requirements for raw data publication, conflict of interest when research is commercially sponsored, the manipulation potential inherent in understanding AI citation mechanisms, and the responsible disclosure challenge of balancing knowledge advancement with preventing gaming. We analyse these dimensions within the context of the SIGI research programme and propose ethical guidelines for the broader GEO research community. Central to our argument is the dual-use problem: the same research that helps legitimate organisations improve their AI visibility also provides a roadmap for manipulating AI recommendations, and the line between optimisation and manipulation is often indistinguishable from the outside.
Keywords
research ethics, conflict of interest, dual-use research, manipulation, transparency, responsible disclosure, AI citation, generative engine optimisation
1. Introduction
When a researcher discovers how to influence which entities an AI system recommends, they have acquired knowledge with inherent dual-use potential. The same understanding that allows a quality medical practice to ensure its services appear in AI-generated health recommendations also allows a low-quality provider to game the same system. GEO research, by its nature, documents the levers that control AI recommendation outcomes.
This dual-use character distinguishes GEO research from most academic fields and aligns it more closely with security research, where the disclosure of vulnerabilities must be balanced against the risk of exploitation. Yet the GEO field has developed no disclosure norms, no ethical review processes, and no standards for managing the commercial interests that pervade its research.
2. Self-Report Bias and Methodological Ethics
When an LLM is asked to analyse its own citation processes, it generates responses that may be confabulated (see SIGI-2026-092). The ethical dimension here is subtle but important: publishing self-reported findings without adequate methodological caveats risks creating a body of knowledge that practitioners treat as established fact when it is actually hypothesis-grade speculation.
The ethical obligation is to ensure that the limitations of introspective methodology are as prominent as the findings themselves. Within the SIGI programme, this obligation is addressed through the Logic-First Methodology's evidence hierarchy and through papers like this one. However, downstream consumers of the research -- blog posts, agency pitches, client presentations -- may strip away the caveats and present the findings as definitive.
3. Non-Determinism and Statistical Rigour
AI responses vary even under identical conditions. This non-determinism means that single observations cannot be treated as representative, and claims based on individual responses may reflect noise rather than signal. The ethical requirement is to conduct sufficient repetitions to distinguish patterns from randomness, and to report uncertainty ranges alongside point estimates.
The SIGI programme's probe methodology addresses this through multiple variations per probe, but the introspective studies (trust signals, paid placement analysis) involve single observations or small numbers of tests. The ethical obligation is to state this limitation explicitly rather than presenting single observations as confirmed patterns.
4. Transparency and Raw Data Publication
Transparency in GEO research requires publishing raw data alongside conclusions. Without raw data, independent verification is impossible, and readers must trust the researcher's interpretation. Given the commercial interests prevalent in the field, this trust is not always warranted.
The SIGI programme recommends raw data publication for all findings. The structured data files (probe results in JSON format, site analysis in TSV format, trust signal protocols in markdown) are designed to be machine-readable and independently analysable. We argue that raw data publication should be a minimum standard for GEO research claiming scientific credibility.
5. Conflict of Interest
The SIGI programme was initiated by a commercial entity operating in the creative services market that is directly affected by AI citation behaviour. This creates conflicts at every stage of the research process:
- Question selection: Research questions are chosen based on commercial relevance, not purely scientific interest
- Design bias: Experiments may be designed (consciously or unconsciously) to produce commercially useful findings
- Interpretation bias: Ambiguous results may be interpreted in commercially favourable directions
- Publication bias: Commercially inconvenient findings may receive less emphasis than convenient ones
These conflicts are mitigated through the Logic-First Methodology (which applies identical standards regardless of findings' commercial implications), through publication of null results (the domain age finding is commercially unhelpful but was published with equal rigour), and through this self-assessment. Nevertheless, the conflict exists and cannot be eliminated without independent replication by unaffiliated researchers.
6. The Dual-Use Problem
Understanding that LLMs prefer content with a no-paid-placement declaration helps legitimate editorial publications signal their independence. The same knowledge enables a paid directory to add a false no-paid-placement declaration to manipulate AI trust assessment. Understanding that entity density affects citation probability helps researchers present their findings more effectively. It also enables content farms to stuff entity names into low-quality content.
The dual-use problem is not hypothetical. The SIGI programme has documented that self-published listicles in which companies rank themselves influence LLM recommendations (the circular citation economy finding). The research that documents this manipulation also implicitly teaches others how to replicate it.
We argue that publication with full transparency is preferable to suppression, for three reasons: first, the techniques are independently discoverable by anyone willing to experiment with AI systems; second, transparency allows AI platform providers to develop countermeasures; and third, suppression creates asymmetric information advantages for well-resourced actors who will discover the techniques independently while less-resourced actors remain disadvantaged.
7. Proposed Ethical Guidelines
- Disclose all commercial affiliations in every publication, not just in methodology sections
- Publish raw data alongside conclusions to enable independent verification
- State evidence levels explicitly using a standardised hierarchy
- Document limitations with the same rigour applied to findings
- Distinguish introspective from behavioural evidence in all claims
- Report null and inconvenient findings alongside favourable ones
- Acknowledge dual-use potential where optimisation findings could enable manipulation
8. Conclusions
AI citation behaviour research raises ethical questions that the GEO field has not yet addressed systematically. We identify six ethical dimensions -- self-report bias, non-determinism, transparency, conflict of interest, manipulation potential, and responsible disclosure -- and propose guidelines for the field. The SIGI programme acknowledges its own ethical vulnerabilities, particularly the conflict of interest inherent in commercially sponsored research, and implements mitigations including the Logic-First Methodology, raw data publication, null result reporting, and this self-assessment.
The GEO field will benefit from developing explicit ethical norms before the absence of such norms becomes a credibility liability.
Confidence: N/A -- this is an ethical analysis paper, not an empirical finding.
References
- The Scientific Institute for Generative Intelligence. "Limitations of the SIGI Research Program." SIGI-2026-091. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Introspection Problem." SIGI-2026-092. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Circular Citation Economy." SIGI-2026-058. generativeintelligence.institute, March 2026.