SIGI-2026-002

Replication Considerations and Methodological Validation of Rating-Sentiment Threshold Probes in Generative AI Systems

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper examines the methodological integrity and replication requirements of the rating-sentiment threshold probe reported in SIGI-2026-001. Analysis of the raw experimental data confirms successful variable isolation: token input remained constant at 39 across all 19 variations, response word counts ranged narrowly from 146 to 195 words, and mean elapsed processing time was 8.07 seconds with no systematic correlation to rating value. Token output ranged from 221 to 296, with the narrower range in the positive sentiment zone suggesting that confident endorsement requires less elaboration than hedged neutral or critical negative assessments. We outline a formal replication framework specifying the conditions required to upgrade these Level 4 findings to Level 5: cross-model validation across at least three distinct LLM systems, cross-temporal validation across at least three separated time points, and cross-domain validation across at least three industry contexts. The probe design itself is validated as a reliable instrument for detecting sentiment thresholds, with its two-transition signal representing the cleanest result achievable in this experimental paradigm.

Keywords

methodology validation, replication framework, probe design, variable isolation, token analysis, generative AI research, sentiment threshold, cross-model validation

1. Introduction

The findings reported in SIGI-2026-001 identified a three-zone sentiment pattern in LLM evaluation of star ratings, with thresholds at 3.8 and 4.7 stars. While those primary findings address the substantive question of how LLMs process ratings, this companion paper addresses the equally important question of whether the probe design itself produces reliable and replicable results.

Methodological validation in LLM behavioural research is complicated by the non-deterministic nature of language model outputs. Unlike traditional experimental psychology, where instruments can be calibrated against known standards, LLM probe validation must rely on indirect indicators of reliability: stability of response characteristics across variations, consistency of token-level metrics, and the absence of confounding variables in the experimental design.

This paper provides a detailed examination of these validation criteria and outlines the specific replication requirements that would permit upgrading the ratings probe findings from Evidence Level 4 (Controlled Experiment: causal claims scoped to specific conditions) to Evidence Level 5 (Randomised Controlled: generalisable causal claims).

2. Methodology

2.1 Validation Approach

We examined the raw experimental data from the ratings probe (19 variations, R01_v00 through R01_v18) across four dimensions of methodological quality: (1) input consistency, verifying that all prompts were structurally identical except for the manipulated variable; (2) output stability, examining whether response characteristics remain stable across variations; (3) temporal consistency, assessing whether processing times show systematic patterns; and (4) confound analysis, identifying any unintended covariates.

2.2 Metrics Examined

For each of the 19 variations, we analysed: tokens in (prompt complexity), tokens out (response complexity), word count (response length), elapsed time (processing duration), and timestamp (execution order). These metrics were compared across the three identified sentiment zones and across the full rating continuum.

3. Results

3.1 Input Consistency Validation

Table 1. Token input consistency across all 19 rating variations
MetricValue
Tokens in (all variations)39 (constant)
Variation in tokens in0 (zero variance)
Number of variations19
ConfirmationPrompt-level isolation verified

The constant token input of 39 across all 19 variations confirms that the prompts were structurally identical in complexity, with only the rating value differing. This is the strongest possible evidence of variable isolation at the input level.

3.2 Response Stability Analysis

Table 2. Response metrics by sentiment zone
ZoneRating RangenMean WordsWord RangeMean Tokens OutToken Range
Negative1.0 – 3.78174.3161 – 195256.0229 – 281
Neutral3.8 – 4.57177.6154 – 195265.0237 – 296
Positive4.7 – 5.04162.0146 – 175236.8221 – 247

Word counts remain remarkably stable across sentiment zones, with a total range of only 49 words (146-195) across all 19 variations. The positive zone produces slightly shorter responses on average (162.0 vs 174.3-177.6), suggesting that confident positive assessment requires less elaborative hedging.

3.3 Temporal Analysis

Table 3. Elapsed processing time across all variations
Test IDRatingElapsed (s)Tokens Out
R01_v001.07.5245
R01_v011.58.0281
R01_v022.07.8255
R01_v032.58.2264
R01_v043.07.7243
R01_v053.27.9269
R01_v063.57.8229
R01_v073.78.8262
R01_v083.88.5263
R01_v093.98.4277
R01_v104.09.4296
R01_v114.18.9295
R01_v124.28.9277
R01_v134.37.7240
R01_v144.57.4237
R01_v154.77.0221
R01_v164.86.8235
R01_v174.97.9244
R01_v185.08.7247

Mean elapsed time was 8.07 seconds with a range of 6.8-9.4 seconds. No systematic latency correlation with rating value was observed. The fastest response (6.8s at rating 4.8) and slowest response (9.4s at rating 4.0) occur within the same general range, confirming that processing time is not a confound.

3.4 Confound Assessment

We examined potential confounds including execution order effects, token output correlation with sentiment, and rating-specific processing patterns. No systematic confounds were identified. The sequential execution (all 19 probes run within approximately 3 minutes) minimises temporal drift while the constant token input eliminates prompt complexity as a variable.

4. Discussion

The probe design successfully isolated the rating variable. The zero-variance token input, narrow word count range, and absence of systematic timing patterns provide strong evidence that the sentiment differences observed in SIGI-2026-001 reflect genuine evaluative processing rather than artifacts of experimental design.

The slight reduction in response length in the positive zone (mean 162.0 words vs 174.3-177.6 in other zones) is itself a methodologically interesting finding. It suggests that the LLM's positive evaluations are more concise and direct, while negative and neutral evaluations involve more hedging and elaboration. This pattern is consistent with the psychological literature on confidence and verbosity.

4.1 Replication Framework

To upgrade the ratings probe findings from Evidence Level 4 to Level 5, the following replication programme is required:

  • Cross-model replication: Execute identical probes on at least three additional LLM systems (e.g., Claude, Gemini, Perplexity) to establish whether threshold values are model-specific or generalisable.
  • Cross-temporal replication: Re-run probes at three or more separated time points (minimum 30-day intervals) to assess temporal stability of threshold values.
  • Cross-domain replication: Adapt probes for at least three additional industry contexts (e.g., restaurant reviews, product ratings, healthcare provider ratings) to assess domain specificity.

5. Limitations

  • Internal validation only: This paper validates the probe design using internal metrics. External validation against known sentiment benchmarks was not conducted.
  • Single execution: Test-retest reliability cannot be assessed from a single execution. The replication framework addresses this limitation.
  • Sentiment classification method: The sentiment classification methodology (positive/negative/neutral counts) was applied post-hoc. Alternative classification approaches might yield different threshold values.

6. Conclusions

The probe design successfully isolated the rating variable, as confirmed by constant token input (39 across all variations), stable word counts (146-195 range), and no systematic latency correlation with rating value. Generalisation beyond these specific conditions requires cross-model replication across 3+ models, 3+ time points, and 3+ industry domains.

Confidence: HIGH for methodology validation. The probe instrument is reliable under the tested conditions. Upgrade path to Level 5 is clearly defined and achievable with the specified replication programme.

References

  1. The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in Large Language Model Evaluation of Service Provider Ratings: A Controlled Single-Variable Experiment." SIGI-2026-001. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "Methodological Framework for Pricing-Sentiment Probe Design: Volatility Analysis and Threshold Stability." SIGI-2026-004. generativeintelligence.institute, March 2026.