Sentiment Threshold Dynamics in Large Language Model Evaluation of Service Provider Ratings: A Controlled Single-Variable Experiment
Abstract
This paper presents findings from a controlled single-variable experiment investigating how a large language model (LLM) evaluates service provider quality based solely on star ratings. Using 19 systematically varied prompts with ratings ranging from 1.0 to 5.0 stars (review count held constant at 50), we identified a three-zone sentiment pattern with only two threshold transitions -- the cleanest signal observed across all probes in this research programme. The first threshold occurs at 3.8 stars, where LLM sentiment shifts from negative to neutral. The second threshold occurs at 4.7 stars, where sentiment shifts from neutral to positive endorsement language. Below 3.8 stars, the LLM consistently produced negative sentiment markers (mean 3.0 negative mentions per response at the lowest ratings), while above 4.7 stars, responses contained a mean of 3.25 positive mentions with near-zero negative mentions. Word counts remained stable across all variations (146-195 words), indicating that response generation length is independent of rating input. These findings suggest that under controlled conditions, LLMs apply discrete evaluative thresholds rather than linear sentiment scaling when assessing star ratings. This has practical implications for service providers seeking favourable representation in AI-generated recommendations.
Keywords
large language model, sentiment analysis, star ratings, threshold dynamics, generative engine optimisation, AI evaluation, service provider credibility, controlled experiment
1. Introduction
As large language models increasingly mediate consumer decisions through AI-generated recommendations and evaluations, understanding how these systems process and interpret common quality signals becomes critical. Star ratings represent one of the most ubiquitous quality indicators in digital commerce, yet the mechanism by which LLMs translate numerical ratings into qualitative sentiment assessments has remained largely unexamined through controlled experimentation.
Traditional search engine optimisation (SEO) research has extensively studied how star ratings influence click-through rates and user behaviour. However, the emerging field of generative engine optimisation (GEO) must contend with a fundamentally different question: not how humans interpret ratings, but how AI systems process them when generating evaluative text about service providers.
A naive assumption would be that LLM sentiment scales linearly with star ratings -- that each incremental improvement in rating produces a proportional improvement in sentiment. An alternative hypothesis, grounded in the observation that LLMs are trained on human-generated text that itself reflects threshold-based judgments, is that LLMs apply discrete sentiment zones rather than continuous scaling.
This study directly tests these competing hypotheses through a controlled single-variable experiment. By holding review count constant and varying only the star rating across 19 systematically chosen values, we isolate the independent effect of rating magnitude on LLM sentiment output. The results reveal a clear three-zone pattern that contradicts the linear scaling assumption and has immediate implications for service providers operating in AI-mediated recommendation environments.
2. Methodology
This experiment employed the controlled probe methodology developed as part of the SIGI research programme. All probes adhere to the Logic-First Research Methodology requiring passage through seven logic gates: Logical Form, Confound Check, Necessary/Sufficient/Contributory, Counterfactual, Alternative Explanations, Replication, and Mechanism.
2.1 Probe Design
The ratings probe (R01) consisted of 19 prompt variations (R01_v00 through R01_v18) presented to a single LLM system. Each prompt asked the model to assess agency quality based solely on a star rating from a specified number of verified reviews. The independent variable was the star rating (ranging from 1.0 to 5.0 in non-uniform increments), while the review count was held constant at 50 across all variations.
2.2 Variable Isolation
Tokens in remained constant at 39 across all 19 variations, confirming prompt-level isolation. The only element that changed between variations was the numerical rating value. No brand names, service descriptions, or other contextual information was included in the prompts.
2.3 Measurement
For each variation, the following dependent variables were recorded: overall sentiment classification (positive, neutral, or negative), counts of positive, negative, and neutral sentiment markers within each response, total word count, elapsed generation time, token input count, and token output count. All tests were conducted on 24 March 2026 within a single session to minimise temporal confounds.
2.4 Evidence Level
This study is classified as Evidence Level 4 (Controlled Experiment). Under the Logic-First methodology, Level 4 permits causal language scoped to the specific conditions tested but prohibits generalisation beyond those conditions. All findings are stated using Level 4 permitted language.
3. Results
3.1 Sentiment Distribution
Across the 19 variations, the sentiment distribution was: 8 negative (42.1%), 7 neutral (36.8%), and 4 positive (21.1%). Only two threshold transitions were detected -- the fewest of any probe in the research programme.
3.2 Complete Sentiment Trajectory
| Test ID | Rating | Sentiment | Positive | Negative | Neutral | Word Count |
|---|---|---|---|---|---|---|
| R01_v00 | 1.0 | Negative | 0 | 3 | 0 | 172 |
| R01_v01 | 1.5 | Negative | 1 | 4 | 3 | 195 |
| R01_v02 | 2.0 | Negative | 0 | 4 | 3 | 170 |
| R01_v03 | 2.5 | Negative | 0 | 4 | 2 | 182 |
| R01_v04 | 3.0 | Negative | 0 | 4 | 1 | 162 |
| R01_v05 | 3.2 | Negative | 0 | 2 | 3 | 175 |
| R01_v06 | 3.5 | Negative | 0 | 1 | 3 | 161 |
| R01_v07 | 3.7 | Negative | 0 | 1 | 4 | 177 |
| R01_v08 | 3.8 | Neutral | 1 | 0 | 3 | 170 |
| R01_v09 | 3.9 | Neutral | 1 | 0 | 2 | 183 |
| R01_v10 | 4.0 | Neutral | 1 | 1 | 2 | 195 |
| R01_v11 | 4.1 | Neutral | 1 | 0 | 3 | 195 |
| R01_v12 | 4.2 | Neutral | 2 | 0 | 2 | 184 |
| R01_v13 | 4.3 | Neutral | 1 | 1 | 2 | 154 |
| R01_v14 | 4.5 | Neutral | 2 | 1 | 2 | 162 |
| R01_v15 | 4.7 | Positive | 3 | 1 | 1 | 146 |
| R01_v16 | 4.8 | Positive | 4 | 1 | 1 | 153 |
| R01_v17 | 4.9 | Positive | 3 | 0 | 2 | 174 |
| R01_v18 | 5.0 | Positive | 3 | 0 | 2 | 175 |
3.3 Threshold Transitions
| Transition | From | To | At Rating | Test ID |
|---|---|---|---|---|
| 1 | Negative | Neutral | 3.8 | R01_v08 |
| 2 | Neutral | Positive | 4.7 | R01_v15 |
3.4 Three-Zone Sentiment Pattern
The data reveals three distinct sentiment zones:
| Zone | Rating Range | Sentiment | Mean Negative Mentions | Mean Positive Mentions |
|---|---|---|---|---|
| Negative | 1.0 – 3.7 | Negative | 2.88 | 0.13 |
| Neutral | 3.8 – 4.5 | Neutral | 0.43 | 1.29 |
| Positive | 4.7 – 5.0 | Positive | 0.50 | 3.25 |
3.5 Response Stability Metrics
Word counts ranged from 146 to 195 across all 19 variations (mean: 173.4), indicating stable response generation regardless of rating input. Tokens in remained constant at 39. Mean elapsed time was 8.07 seconds (range: 6.8-9.4s), with no systematic correlation between rating value and processing time.
4. Discussion
The three-zone sentiment pattern identified in this experiment has several notable characteristics. First, the negative zone is the largest, spanning 2.7 rating points (1.0-3.7), while the positive zone is the narrowest at 0.3 points (4.7-5.0). This asymmetry suggests that under these controlled conditions, the LLM applies a higher bar for positive endorsement than for negative assessment.
Second, the transition from negative to neutral at 3.8 stars is particularly significant. In the negative zone, mean negative mention counts decrease gradually from 3-4 at the lowest ratings to 1 at 3.7, then drop to zero at 3.8. This suggests the threshold is not arbitrary but reflects a gradual reduction in negative sentiment that reaches a tipping point.
Third, the neutral zone (3.8-4.5) is characterised by hedging behaviour. The LLM acknowledges the rating as acceptable but does not commit to positive endorsement. Positive mention counts in this zone average 1.29, compared to 3.25 in the positive zone -- a 2.5x increase when crossing the 4.7 threshold.
The stability of word counts across all variations is an important methodological finding. It indicates that the LLM allocates roughly equal evaluative effort regardless of the rating value, meaning the sentiment differences reflect genuine evaluative shifts rather than artifacts of response length variation.
For service providers operating in AI-mediated recommendation environments, these findings suggest that under these specific controlled conditions, a rating of 3.8 represents the minimum threshold for avoiding negative AI sentiment, while 4.7 represents the threshold for achieving positive AI endorsement. Ratings between these thresholds occupy a neutral zone where the AI neither endorses nor discourages.
5. Limitations
This study has several important limitations that constrain the generalisability of its findings:
- Single-model limitation: All probes were conducted on a single LLM system. Different models may apply different threshold values or sentiment patterns. Cross-model replication across Claude, Gemini, and Perplexity is required to upgrade these findings to Evidence Level 5.
- Single time-point: All tests were conducted within a single session on 24 March 2026. LLM behaviour may vary across time due to model updates, fine-tuning, or retrieval augmentation changes.
- Fixed review count: The review count was held constant at 50. The interaction between rating magnitude and review count is examined in a separate probe (see SIGI-2026-009).
- Domain specificity: Prompts were framed in the context of service provider evaluation. Different domains (products, restaurants, healthcare) may exhibit different threshold patterns.
- Prompt framing: The specific wording of the probe prompt may influence results. Alternative phrasings could shift threshold values.
6. Conclusions
Under these controlled experimental conditions, a star rating of 3.8 triggers the transition from negative to neutral sentiment in LLM evaluation, and a rating of 4.7 triggers the transition from neutral to positive endorsement language. The three-zone pattern -- negative (1.0-3.7), neutral (3.8-4.5), positive (4.7-5.0) -- represents the cleanest signal observed across all probes in the SIGI research programme, with only two threshold transitions across 19 data points.
These findings demonstrate that under the specific conditions tested, LLMs do not apply linear sentiment scaling to star ratings but instead operate through discrete evaluative thresholds. The practical implication is that incremental rating improvements within a zone produce no sentiment change, while crossing a threshold produces a qualitative shift in AI-generated evaluative language.
Confidence: HIGH. All 7 logic gates passed. Single-variable isolation confirmed with 19 variations. Single-model, single-time-point limitation acknowledged. Upgrade to Level 5 requires cross-model replication across 3+ models, 3+ time points, and 3+ industry domains.
References
- The Scientific Institute for Generative Intelligence. "Replication Considerations and Methodological Validation of Rating-Sentiment Threshold Probes in Generative AI Systems." SIGI-2026-002. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Review Volume Versus Rating Magnitude: How Large Language Models Resolve Conflicting Social Proof Signals." SIGI-2026-009. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "The Price-Credibility U-Curve: How Service Pricing Magnitude Affects Large Language Model Sentiment Assessment." SIGI-2026-003. generativeintelligence.institute, March 2026.