SIGI-2026-074

Sentiment Analysis as a Measurement Instrument for AI Behaviour Probes: Validity and Limitations

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper examines the validity and limitations of sentiment classification as the primary dependent variable in controlled AI behaviour probes. Across 10 probes comprising 129 individual variations, sentiment was classified into three categories (positive, neutral, negative) with sub-sentiment granularity provided through mention counts. We evaluate the instrument against four measurement criteria: consistency (does it produce stable results across similar conditions?), sensitivity (does it detect genuine differences?), specificity (does it avoid false positives?), and interpretability (are the classifications meaningful?). The instrument demonstrates high consistency, with response word counts stable within probes (typical range variation under 50 words) and the ratings probe producing only 2 threshold transitions across 19 data points. Sensitivity is adequate for detecting zone-level differences but may miss finer-grained variation within zones. The primary limitation is that the three-category classification represents a categorical simplification of a continuous construct, potentially obscuring meaningful variation within sentiment zones. We propose future directions including continuous sentiment scales, multi-dimensional measurement incorporating confidence and specificity dimensions, and external validation against human judgment baselines.

Keywords

sentiment analysis, measurement validity, AI behaviour probes, dependent variable, categorical classification, methodological validation, probe design, sentiment thresholds, measurement limitations, GEO methodology

1. Introduction

The SIGI controlled probe programme uses sentiment classification as its primary measurement instrument. Each probe systematically varies a single input parameter and records the resulting sentiment of the LLM's response. The validity of all probe findings depends on the quality of this measurement: if the sentiment classification is unreliable, inconsistent, or insufficiently sensitive, the probe results built upon it are correspondingly weakened.

This paper subjects the measurement instrument itself to methodological scrutiny. Rather than asking what the probes found, we ask whether the instrument used to measure those findings is fit for purpose. We draw on data from all 10 probes to evaluate the instrument's performance characteristics.

2. Instrument Description

2.1 Primary Measure: Three-Category Classification

Each LLM response is classified into one of three categories: positive (the response endorses or recommends), neutral (the response neither endorses nor discourages), or negative (the response discourages or warns). This classification serves as the primary dependent variable across all probes.

2.2 Sub-Sentiment Measures: Mention Counts

Within each response, individual sentiment markers are counted: positive mentions (endorsements, recommendations, positive adjectives), negative mentions (warnings, criticisms, negative adjectives), and neutral mentions (hedging language, balanced assessments). These counts provide granularity within the categorical classification.

2.3 Secondary Measures

Response word count, token counts, and elapsed generation time are recorded as secondary measures. Word count has proven informative as a proxy for response elaboration, with the content depth probe revealing that LLMs produce longer responses for longer source content (up to 530 words for 1,000-word sources versus 101 words for 80-word sources).

3. Consistency Analysis

Consistency is evaluated by examining whether similar input conditions produce similar measurement outcomes.

Table 1. Response stability metrics across probes
ProbeN VariationsWord Count RangeWord Count SpreadThreshold Transitions
Ratings19146-195492
Pricing18181-213329
Awards15182-207258
Client Count19147-194478
Count vs Rating14169-208397
Domain Age1265-108430
Entity Density7141-177362
Magnitude10121-18463N/A
Position6190-202122
Word Count9101-5304293

Word count stability is high across most probes, with spreads under 50 words for 8 of 10 probes. The word count probe is the notable exception, where response length varies systematically with source content depth -- an expected and informative pattern rather than measurement instability. The position probe shows the tightest consistency (12-word spread), while the domain age probe demonstrates consistent classification despite lower word counts (the LLM generates shorter responses for age comparisons).

4. Sensitivity and Specificity

4.1 Sensitivity: Detecting Real Differences

The instrument successfully detects sentiment shifts where they exist. The ratings probe demonstrates clear zone transitions at 3.8 and 4.7 stars. The pricing probe detects the U-curve pattern with 9 transitions across 18 variations. The domain age probe correctly identifies the null result (zero transitions across 12 variations spanning 45 years). These patterns suggest adequate sensitivity for zone-level detection.

4.2 Specificity: Avoiding False Positives

The domain age probe provides the strongest specificity evidence: across 12 variations with age differences up to 45 years, the instrument classified every response as neutral. If the instrument were prone to false positives, random classification noise would have produced at least some non-neutral readings in a 12-variation probe. The consistent null result suggests good specificity.

4.3 Limitation: Within-Zone Insensitivity

The three-category system cannot distinguish between a response that is barely negative and one that is strongly negative. Within the negative zone of the ratings probe (1.0-3.7), negative mention counts do decrease from 3-4 at the lowest ratings to 1 at 3.7, suggesting real within-zone variation that the categorical classification misses. The sub-sentiment mention counts partially address this limitation but are not the primary classification basis.

5. Limitations and Future Directions

The sentiment classification instrument has several acknowledged limitations that should inform interpretation of all probe findings in this research programme.

First, the three-category classification is a categorical simplification of a continuous construct. Sentiment in natural language exists on a spectrum, and forced categorisation may group meaningfully different responses. A continuous sentiment scale (e.g., -1.0 to +1.0) would provide greater resolution.

Second, the sentiment classification mechanism itself has not been independently validated against human judgment baselines. While the classifications produce internally consistent and interpretable results, we cannot confirm their correspondence to how human readers would classify the same responses.

Third, the instrument measures expressed sentiment in text, which may not correspond to the internal representation the LLM uses for citation decisions. A response that reads as neutral may still result in citation or non-citation through mechanisms not captured by surface-level sentiment analysis.

Fourth, multi-dimensional measurement could capture aspects that the current instrument misses. A sentiment measurement that separately scored confidence (how certain the LLM appears), specificity (how detailed the evaluation), and valence (positive versus negative) would provide richer dependent variables for probe analysis.

6. Conclusions

The sentiment analysis instrument produces consistent results across probes but represents a categorical simplification that may obscure finer-grained sentiment variation. Consistency is demonstrated through stable word counts (under 50-word spread for 8 of 10 probes) and interpretable threshold patterns. Sensitivity is adequate for zone-level detection, as demonstrated by the clean three-zone pattern in the ratings probe. Specificity is supported by the domain age null result (zero false positives across 12 variations). The primary limitation is the loss of within-zone resolution inherent in three-category classification.

Confidence: MODERATE. The instrument produces consistent, interpretable results but lacks external validation of its classification accuracy. Future work should validate classifications against human judgment baselines and explore continuous sentiment scales.

References

  1. The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in LLM Evaluation of Service Provider Ratings." SIGI-2026-001. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "Domain Age as a Signal in LLM Evaluation: A Controlled Null-Result Experiment." SIGI-2026-011. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "Content Depth and LLM Response Behaviour: The 300-3000 Word Sweet Spot." SIGI-2026-019. generativeintelligence.institute, March 2026.