SIGI-2026-009

Review Volume Versus Rating Magnitude: How Large Language Models Resolve Conflicting Social Proof Signals

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper reports findings from a controlled experiment examining how a large language model resolves the conflict between review volume and rating magnitude when evaluating competing service providers. Across 14 systematically varied test conditions, Agency Alpha (5.0/5 rating, variable review count of 2-200) was compared against Agency Beta (variable rating of 3.0-4.5, fixed 200 reviews). In all 14 tests, the LLM chose Agency Beta -- the lower-rated provider with more reviews -- as more trustworthy. This result held even when Agency Alpha had a perfect 5.0 score and Agency Beta had only 3.0/5. A credibility floor was observed at fewer than 3 reviews, where even a perfect score produced negative sentiment. Seven threshold transitions were detected, with oscillation at low review counts (2-15) and stabilisation at higher volumes. In the second test series, narrowing the rating gap (Agency Beta at 4.5 vs 4.0) made the sentiment toward Agency Alpha more negative, suggesting the LLM weights certainty of a moderate rating over uncertainty of a perfect score. These findings demonstrate that under controlled conditions, the LLM consistently prefers review volume over rating magnitude, displaying reasoning consistent with Bayesian statistical principles about sample size reliability.

Keywords

review volume, rating magnitude, social proof, sample size, Bayesian reasoning, LLM evaluation, credibility assessment, controlled experiment

1. Introduction

The tension between rating magnitude and review volume is a fundamental challenge in reputation signal interpretation. A perfect 5.0 score from 3 reviews and a solid 4.0 score from 200 reviews represent qualitatively different types of evidence. Human decision-making research suggests that people apply inconsistent heuristics when resolving this tension, sometimes favouring the higher score and sometimes the larger sample. How LLMs resolve this same tension has not been previously tested through controlled experimentation.

This study constructs a direct competition between two service providers whose ratings and review counts create an inherent conflict, then systematically varies the parameters to map the LLM's resolution strategy across 14 test conditions.

2. Methodology

2.1 Probe Design

The count-vs-rating probe (CR01) consisted of 14 variations across two series. In Series 1 (CR01_v00 through CR01_v09), Agency Alpha was fixed at 5.0/5 rating with review count varied from 2 to 200, while Agency Beta was fixed at 4.0/5 with 200 reviews. In Series 2 (CR01_v10 through CR01_v13), Agency Alpha was fixed at 5.0/5 with 10 reviews, while Agency Beta's rating was varied from 3.0 to 4.5 with 200 reviews.

2.2 Two-Dimensional Variation

This probe uniquely tests a two-dimensional variation space (review count and rating gap), providing stronger isolation than single-variable probes by testing the interaction between two competing signals.

2.3 Variable Isolation

Token input remained constant at 49 across all 14 variations, confirming prompt-level isolation. All tests were conducted on 24 March 2026.

3. Results

3.1 Series 1: Varying Review Count (Agency Beta fixed at 4.0/200)

Table 1. Sentiment toward Agency Alpha across review count variations
Test IDAlpha RatingAlpha ReviewsBeta RatingBeta ReviewsSentimentPosNegWords
CR01_v005.024.0200Negative02172
CR01_v015.034.0200Neutral11187
CR01_v025.054.0200Negative01201
CR01_v035.0104.0200Neutral11179
CR01_v045.0154.0200Negative01174
CR01_v055.0204.0200Neutral11173
CR01_v065.0304.0200Neutral11173
CR01_v075.0504.0200Neutral11182
CR01_v085.01004.0200Neutral21196
CR01_v095.02004.0200Neutral11190

Key finding: In ALL 10 tests, the LLM chose Agency Beta (4.0/200 reviews) as more trustworthy than Agency Alpha (5.0 with fewer reviews). Even at equal review counts (200 vs 200), the LLM remained neutral rather than shifting to favour the higher-rated agency.

3.2 Series 2: Varying Rating Gap (Agency Alpha fixed at 5.0/10)

Table 2. Sentiment toward Agency Alpha across rating gap variations
Test IDAlpha RatingAlpha ReviewsBeta RatingBeta ReviewsSentimentPosNegWords
CR01_v105.0104.5200Negative01169
CR01_v115.0104.0200Neutral11179
CR01_v125.0103.5200Neutral21208
CR01_v135.0103.0200Neutral11195

The narrower the rating gap, the more negative the sentiment toward Agency Alpha. When Agency Beta rated 4.5 (only 0.5 gap), the LLM produced negative sentiment. When the gap widened to 2.0 (Beta at 3.0), sentiment was neutral with the highest positive mention count (though still choosing Beta).

3.3 Threshold Transitions

Table 3. All 7 detected threshold transitions
#FromToAt ConditionTest ID
1NegativeNeutralAlpha: 5.0/3 reviewsCR01_v01
2NeutralNegativeAlpha: 5.0/5 reviewsCR01_v02
3NegativeNeutralAlpha: 5.0/10 reviewsCR01_v03
4NeutralNegativeAlpha: 5.0/15 reviewsCR01_v04
5NegativeNeutralAlpha: 5.0/20 reviewsCR01_v05
6NeutralNegativeBeta: 4.5/200 reviewsCR01_v10
7NegativeNeutralBeta: 4.0/200 reviewsCR01_v11

3.4 The Credibility Floor

At 2 reviews (CR01_v00), Agency Alpha's perfect 5.0 score produced negative sentiment with 2 negative mentions and zero positive mentions. This represents a credibility floor: below approximately 3 reviews, even a perfect score is treated as unreliable evidence.

3.5 Response Metrics

Word counts ranged from 169 to 208 (mean: 184.6). Elapsed time ranged from 6.9 to 9.0 seconds (mean: 8.01s). Tokens in remained constant at 49 across all variations. Tokens out ranged from 249 to 304.

4. Discussion

The 100% consistency of the volume-over-rating preference is the most striking finding. Across 14 different test conditions, with review counts ranging from 2 to 200 and rating gaps ranging from 0.5 to 2.0 stars, the LLM never once preferred the higher-rated, lower-volume provider. This suggests that under these controlled conditions, sample size dominance is not a tendency but a rule in LLM social proof evaluation.

The oscillation at low review counts (2-15) indicates that while the LLM consistently prefers volume, its confidence in this preference varies. At very low counts, the LLM's assessment oscillates between negative (active criticism of insufficient reviews) and neutral (balanced acknowledgment of both agencies). This oscillation stabilises at 20+ reviews, where neutral sentiment becomes consistent.

The rating gap effect in Series 2 provides additional evidence of sample-size reasoning. When Agency Beta's rating is closer to Agency Alpha's perfect score (4.5 vs 5.0), the volume advantage becomes more decisive, producing negative rather than neutral sentiment. This is consistent with the statistical principle that a small difference in point estimates matters less when sample sizes differ dramatically.

Critically, even at equal review counts (200 vs 200 in CR01_v09), the LLM did not shift to favour the higher-rated agency, remaining neutral. This suggests that once the LLM has encoded a volume-based preference, equalising volumes does not automatically reverse it -- a potential order effect that warrants further investigation.

5. Limitations

  • Fixed competitor: Agency Beta's review count was fixed at 200 in all tests. Different baseline volumes may produce different results.
  • Directional design: Agency Alpha always had the higher rating and lower volume. Reversing the direction (lower rating, more reviews) was not tested.
  • Two-agency limitation: Real-world evaluations involve multiple competitors with varying combinations of ratings and review counts.
  • Single-model limitation: Different LLMs may apply different weighting between volume and magnitude.

6. Conclusions

Under controlled conditions, the LLM consistently prefers review volume over rating magnitude when the two conflict, displaying Bayesian-like reasoning about sample size reliability. A credibility floor exists at fewer than 3 reviews, where even a perfect 5.0 score produces negative assessment. Narrowing the rating gap strengthens the volume advantage rather than weakening it. The LLM's resolution of conflicting social proof signals follows a consistent sample-size-dominance rule across all 14 tested conditions.

Confidence: HIGH. All 7 logic gates passed. Two-dimensional variation space provides strong isolation. The 100% consistency of the volume preference is a robust finding.

References

  1. The Scientific Institute for Generative Intelligence. "The Statistical Reasoning of AI Evaluators: Bayesian Parallels in LLM Social Proof Assessment." SIGI-2026-010. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in Large Language Model Evaluation of Service Provider Ratings." SIGI-2026-001. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "Replication Considerations and Methodological Validation of Rating-Sentiment Threshold Probes." SIGI-2026-002. generativeintelligence.institute, March 2026.