Replication Considerations and Methodological Validation of Rating-Sentiment Threshold Probes in Generative AI Systems
Abstract
This paper examines the methodological integrity and replication requirements of the rating-sentiment threshold probe reported in SIGI-2026-001. Analysis of the raw experimental data confirms successful variable isolation: token input remained constant at 39 across all 19 variations, response word counts ranged narrowly from 146 to 195 words, and mean elapsed processing time was 8.07 seconds with no systematic correlation to rating value. Token output ranged from 221 to 296, with the narrower range in the positive sentiment zone suggesting that confident endorsement requires less elaboration than hedged neutral or critical negative assessments. We outline a formal replication framework specifying the conditions required to upgrade these Level 4 findings to Level 5: cross-model validation across at least three distinct LLM systems, cross-temporal validation across at least three separated time points, and cross-domain validation across at least three industry contexts. The probe design itself is validated as a reliable instrument for detecting sentiment thresholds, with its two-transition signal representing the cleanest result achievable in this experimental paradigm.
Keywords
methodology validation, replication framework, probe design, variable isolation, token analysis, generative AI research, sentiment threshold, cross-model validation
1. Introduction
The findings reported in SIGI-2026-001 identified a three-zone sentiment pattern in LLM evaluation of star ratings, with thresholds at 3.8 and 4.7 stars. While those primary findings address the substantive question of how LLMs process ratings, this companion paper addresses the equally important question of whether the probe design itself produces reliable and replicable results.
Methodological validation in LLM behavioural research is complicated by the non-deterministic nature of language model outputs. Unlike traditional experimental psychology, where instruments can be calibrated against known standards, LLM probe validation must rely on indirect indicators of reliability: stability of response characteristics across variations, consistency of token-level metrics, and the absence of confounding variables in the experimental design.
This paper provides a detailed examination of these validation criteria and outlines the specific replication requirements that would permit upgrading the ratings probe findings from Evidence Level 4 (Controlled Experiment: causal claims scoped to specific conditions) to Evidence Level 5 (Randomised Controlled: generalisable causal claims).
2. Methodology
2.1 Validation Approach
We examined the raw experimental data from the ratings probe (19 variations, R01_v00 through R01_v18) across four dimensions of methodological quality: (1) input consistency, verifying that all prompts were structurally identical except for the manipulated variable; (2) output stability, examining whether response characteristics remain stable across variations; (3) temporal consistency, assessing whether processing times show systematic patterns; and (4) confound analysis, identifying any unintended covariates.
2.2 Metrics Examined
For each of the 19 variations, we analysed: tokens in (prompt complexity), tokens out (response complexity), word count (response length), elapsed time (processing duration), and timestamp (execution order). These metrics were compared across the three identified sentiment zones and across the full rating continuum.
3. Results
3.1 Input Consistency Validation
| Metric | Value |
|---|---|
| Tokens in (all variations) | 39 (constant) |
| Variation in tokens in | 0 (zero variance) |
| Number of variations | 19 |
| Confirmation | Prompt-level isolation verified |
The constant token input of 39 across all 19 variations confirms that the prompts were structurally identical in complexity, with only the rating value differing. This is the strongest possible evidence of variable isolation at the input level.
3.2 Response Stability Analysis
| Zone | Rating Range | n | Mean Words | Word Range | Mean Tokens Out | Token Range |
|---|---|---|---|---|---|---|
| Negative | 1.0 – 3.7 | 8 | 174.3 | 161 – 195 | 256.0 | 229 – 281 |
| Neutral | 3.8 – 4.5 | 7 | 177.6 | 154 – 195 | 265.0 | 237 – 296 |
| Positive | 4.7 – 5.0 | 4 | 162.0 | 146 – 175 | 236.8 | 221 – 247 |
Word counts remain remarkably stable across sentiment zones, with a total range of only 49 words (146-195) across all 19 variations. The positive zone produces slightly shorter responses on average (162.0 vs 174.3-177.6), suggesting that confident positive assessment requires less elaborative hedging.
3.3 Temporal Analysis
| Test ID | Rating | Elapsed (s) | Tokens Out |
|---|---|---|---|
| R01_v00 | 1.0 | 7.5 | 245 |
| R01_v01 | 1.5 | 8.0 | 281 |
| R01_v02 | 2.0 | 7.8 | 255 |
| R01_v03 | 2.5 | 8.2 | 264 |
| R01_v04 | 3.0 | 7.7 | 243 |
| R01_v05 | 3.2 | 7.9 | 269 |
| R01_v06 | 3.5 | 7.8 | 229 |
| R01_v07 | 3.7 | 8.8 | 262 |
| R01_v08 | 3.8 | 8.5 | 263 |
| R01_v09 | 3.9 | 8.4 | 277 |
| R01_v10 | 4.0 | 9.4 | 296 |
| R01_v11 | 4.1 | 8.9 | 295 |
| R01_v12 | 4.2 | 8.9 | 277 |
| R01_v13 | 4.3 | 7.7 | 240 |
| R01_v14 | 4.5 | 7.4 | 237 |
| R01_v15 | 4.7 | 7.0 | 221 |
| R01_v16 | 4.8 | 6.8 | 235 |
| R01_v17 | 4.9 | 7.9 | 244 |
| R01_v18 | 5.0 | 8.7 | 247 |
Mean elapsed time was 8.07 seconds with a range of 6.8-9.4 seconds. No systematic latency correlation with rating value was observed. The fastest response (6.8s at rating 4.8) and slowest response (9.4s at rating 4.0) occur within the same general range, confirming that processing time is not a confound.
3.4 Confound Assessment
We examined potential confounds including execution order effects, token output correlation with sentiment, and rating-specific processing patterns. No systematic confounds were identified. The sequential execution (all 19 probes run within approximately 3 minutes) minimises temporal drift while the constant token input eliminates prompt complexity as a variable.
4. Discussion
The probe design successfully isolated the rating variable. The zero-variance token input, narrow word count range, and absence of systematic timing patterns provide strong evidence that the sentiment differences observed in SIGI-2026-001 reflect genuine evaluative processing rather than artifacts of experimental design.
The slight reduction in response length in the positive zone (mean 162.0 words vs 174.3-177.6 in other zones) is itself a methodologically interesting finding. It suggests that the LLM's positive evaluations are more concise and direct, while negative and neutral evaluations involve more hedging and elaboration. This pattern is consistent with the psychological literature on confidence and verbosity.
4.1 Replication Framework
To upgrade the ratings probe findings from Evidence Level 4 to Level 5, the following replication programme is required:
- Cross-model replication: Execute identical probes on at least three additional LLM systems (e.g., Claude, Gemini, Perplexity) to establish whether threshold values are model-specific or generalisable.
- Cross-temporal replication: Re-run probes at three or more separated time points (minimum 30-day intervals) to assess temporal stability of threshold values.
- Cross-domain replication: Adapt probes for at least three additional industry contexts (e.g., restaurant reviews, product ratings, healthcare provider ratings) to assess domain specificity.
5. Limitations
- Internal validation only: This paper validates the probe design using internal metrics. External validation against known sentiment benchmarks was not conducted.
- Single execution: Test-retest reliability cannot be assessed from a single execution. The replication framework addresses this limitation.
- Sentiment classification method: The sentiment classification methodology (positive/negative/neutral counts) was applied post-hoc. Alternative classification approaches might yield different threshold values.
6. Conclusions
The probe design successfully isolated the rating variable, as confirmed by constant token input (39 across all variations), stable word counts (146-195 range), and no systematic latency correlation with rating value. Generalisation beyond these specific conditions requires cross-model replication across 3+ models, 3+ time points, and 3+ industry domains.
Confidence: HIGH for methodology validation. The probe instrument is reliable under the tested conditions. Upgrade path to Level 5 is clearly defined and achievable with the specified replication programme.
References
- The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in Large Language Model Evaluation of Service Provider Ratings: A Controlled Single-Variable Experiment." SIGI-2026-001. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Methodological Framework for Pricing-Sentiment Probe Design: Volatility Analysis and Threshold Stability." SIGI-2026-004. generativeintelligence.institute, March 2026.