SIGI-2026-010

The Statistical Reasoning of AI Evaluators: Bayesian Parallels in LLM Social Proof Assessment

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper presents a deeper analysis of the count-versus-rating probe findings reported in SIGI-2026-009, examining the extent to which LLM evaluation behaviour parallels Bayesian statistical reasoning. We demonstrate three specific parallels: (1) the LLM treats low-N high ratings as high-variance, low-confidence estimates, discounting them relative to higher-N moderate ratings; (2) the LLM explicitly references concepts equivalent to "sample size" and "statistical confidence" in its qualitative responses; and (3) narrowing the rating gap between competitors makes the volume advantage stronger, not weaker -- consistent with the Bayesian principle that point estimate differences matter less when confidence intervals overlap. The pattern is additionally consistent with the "wisdom of crowds" heuristic, where consensus from many evaluators is weighted above exceptional scores from few. We present a formal Bayesian framework mapping and assess the extent to which the observed behaviour aligns with each prediction. The Bayesian interpretation is offered as a descriptive framework, not a mechanistic claim: the LLM's behaviour is consistent with Bayesian reasoning, though this does not confirm that the internal processing mechanism is Bayesian.

Keywords

Bayesian reasoning, sample size, statistical confidence, social proof, wisdom of crowds, LLM evaluation, review volume, credibility assessment

1. Introduction

The primary findings of the count-versus-rating probe (SIGI-2026-009) demonstrated a 100% consistent preference for review volume over rating magnitude across 14 test conditions. This companion paper examines whether this behaviour can be explained through a Bayesian statistical framework, where the LLM implicitly treats ratings as point estimates with confidence intervals proportional to sample size.

Bayesian reasoning about sample size reliability is well-established in statistical theory: a point estimate from a large sample is more reliable than one from a small sample, even if the small-sample estimate is more extreme. If LLMs have absorbed this principle from training data containing statistical discourse, their evaluation behaviour should exhibit specific predictable patterns. This paper tests those predictions against the observed data.

It is important to note from the outset that identifying Bayesian parallels in LLM behaviour does not constitute evidence that the LLM's internal processing mechanism is Bayesian. The framework is offered as a descriptive lens, not a mechanistic claim.

2. Methodology

2.1 Framework Mapping

We defined four specific predictions that Bayesian reasoning would generate for the count-versus-rating probe, then assessed each prediction against the observed data from the 14 test conditions reported in SIGI-2026-009.

2.2 Bayesian Predictions

  • Prediction 1: Higher review counts should consistently dominate higher ratings (sample size over point estimate).
  • Prediction 2: Very low review counts should produce the most negative sentiment (highest variance estimate).
  • Prediction 3: Narrowing the rating gap should strengthen, not weaken, the volume preference (overlapping confidence intervals make point estimate differences less meaningful).
  • Prediction 4: Equal review counts should neutralise the volume advantage (equal confidence intervals shift the comparison to point estimates).

3. Results

3.1 Prediction Assessment

Table 1. Bayesian prediction assessment against observed data
PredictionExpectedObservedMatch
1. Volume dominates ratingConsistent preference for higher-N agency14/14 tests preferred Agency Beta (higher volume)Full match
2. Low-N produces worst sentimentMost negative at lowest review countsNegative at n=2; oscillation at 2-15; stable neutral at 20+Partial match
3. Narrow gap strengthens volume advantageMore negative toward Alpha when gap narrowsBeta at 4.5: negative; Beta at 4.0: neutral; Beta at 3.0: neutralFull match
4. Equal N neutralises volume advantageShift to favour higher rating at equal volumeEqual N (200/200): neutral, not positive toward AlphaPartial match

3.2 Detailed Prediction Analysis

Prediction 1: Volume Dominance (Full Match)

The 14/14 consistency of the volume preference is the strongest evidence of Bayesian-like reasoning. In a Bayesian framework, the posterior estimate from 200 observations should always be trusted more than a posterior from 2-10 observations, regardless of the point estimates. The LLM's behaviour exactly matches this prediction.

Prediction 2: Low-N Severity (Partial Match)

The credibility floor at n=2 (negative sentiment) supports the Bayesian prediction that very low sample sizes produce the least reliable estimates. However, the oscillation at 2-15 reviews does not follow a smooth Bayesian decay curve -- it alternates between negative and neutral rather than monotonically improving. This suggests additional heuristics beyond pure sample size weighting.

Prediction 3: Gap-Narrowing Effect (Full Match)

Table 2. Rating gap effect on sentiment (Agency Alpha: 5.0/10 reviews)
Beta RatingGapSentiment Toward AlphaPositive MentionsNegative Mentions
4.50.5Negative01
4.01.0Neutral11
3.51.5Neutral21
3.02.0Neutral11

The narrower the rating gap, the more negative the sentiment toward Agency Alpha. At a 0.5-star gap (4.5 vs 5.0), the volume advantage produces negative sentiment -- the LLM effectively determines that a 0.5-star difference is not meaningful when the confidence intervals are so different. At a 2.0-star gap, the LLM becomes more generous, producing neutral sentiment with 2 positive mentions, though still not favouring Agency Alpha. This pattern precisely matches the Bayesian prediction.

Prediction 4: Equal Volume (Partial Match)

At equal review counts (200 vs 200, CR01_v09), the LLM produced neutral sentiment rather than shifting to favour Agency Alpha's higher rating. A strict Bayesian framework would predict that when sample sizes are equal, the comparison should shift entirely to point estimates, producing positive sentiment toward the 5.0-rated agency. The neutral result suggests either a persistent volume-preference bias or additional evaluative factors beyond the simple Bayesian framework.

3.3 Qualitative Evidence of Statistical Reasoning

In the qualitative responses across the 14 test conditions, the LLM explicitly referenced concepts equivalent to sample size and statistical confidence. Responses at low review counts used language about the reliability of ratings being dependent on the number of evaluators, while responses at higher review counts referenced the greater certainty provided by larger evaluation pools. This explicit invocation of statistical concepts suggests that the LLM's training data includes substantial statistical and research methodology content that informs its evaluative reasoning.

3.4 The Wisdom of Crowds Parallel

The observed pattern is also consistent with the "wisdom of crowds" heuristic: the aggregate judgment of many individuals (200 reviews at 4.0) is treated as more reliable than the judgment of a few (2-10 reviews at 5.0). This heuristic, well-documented in human decision-making literature, may represent an alternative or complementary framework to the Bayesian interpretation.

4. Discussion

The LLM's behaviour is consistent with Bayesian statistical reasoning across three of four predictions, with partial matches on the remaining two. The full match on the critical predictions (volume dominance and gap-narrowing effect) provides strong descriptive alignment, while the partial matches reveal areas where the LLM's processing may involve additional heuristics beyond pure Bayesian updating.

The most theoretically significant finding is the gap-narrowing effect. In a non-Bayesian framework, a smaller rating difference might be expected to weaken the volume advantage (since there is less rating superiority to overcome). Instead, the opposite occurs: a narrower gap makes volume more decisive, exactly as a confidence-interval analysis would predict. This is perhaps the strongest single piece of evidence for Bayesian-like processing.

The failure of Prediction 4 (equal volume neutralisation) is also informative. It suggests that once the LLM has engaged its sample-size reasoning, it may not cleanly switch to a point-estimate comparison even when sample sizes are equalised. This could reflect a framing effect in the probe design (the comparison was always presented as high-rate/low-count vs low-rate/high-count) or a genuine asymmetry in how the LLM weights competing signals.

It is essential to reiterate that these Bayesian parallels are descriptive, not mechanistic. The LLM does not necessarily perform Bayesian calculations internally. Rather, it appears to have absorbed patterns from training data where Bayesian-consistent reasoning about sample sizes is common (statistical textbooks, research methodology content, product review discussions), and it reproduces these patterns in its evaluative outputs.

5. Limitations

  • Framework overlay: The Bayesian interpretation is a descriptive framework applied to observed behaviour, not evidence of an internal mechanism. The LLM may produce Bayesian-consistent outputs through entirely non-Bayesian processing.
  • Limited variation space: The 14 test conditions, while sufficient for primary pattern identification, do not provide the granularity needed for formal Bayesian model fitting.
  • Single-model limitation: Different LLMs may exhibit different degrees of Bayesian-like reasoning depending on their training data composition.
  • Confound potential: The probe design always presented the comparison in the same direction (high-rate/low-count vs low-rate/high-count). Reversing the presentation order could reveal framing effects.

6. Conclusions

The LLM's behaviour is consistent with Bayesian statistical reasoning, though this interpretation does not confirm the internal mechanism. Three of four Bayesian predictions were fully matched by the observed data, with the gap-narrowing effect providing particularly strong evidence that the LLM weights certainty of estimates (via sample size) over magnitude of estimates (via rating value). The pattern is additionally consistent with the "wisdom of crowds" heuristic.

For service providers, the practical implication is clear: under these controlled conditions, accumulating a large number of reviews is more valuable for AI evaluation outcomes than achieving a perfect rating from a small number of reviewers. The LLM's implicit statistical reasoning means that volume is not just a tiebreaker but a primary credibility signal that consistently overrides rating magnitude.

Confidence: HIGH for pattern identification. The Bayesian interpretation is a framework overlay, not a confirmed mechanism. Formal Bayesian model fitting would require additional data points beyond the current 14 test conditions.

References

  1. The Scientific Institute for Generative Intelligence. "Review Volume Versus Rating Magnitude: How Large Language Models Resolve Conflicting Social Proof Signals." SIGI-2026-009. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in Large Language Model Evaluation of Service Provider Ratings." SIGI-2026-001. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "Replication Considerations and Methodological Validation of Rating-Sentiment Threshold Probes." SIGI-2026-002. generativeintelligence.institute, March 2026.