SIGI-2026-008

Comparative Volatility Analysis: Awards versus Client Count as LLM Credibility Signals

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper presents a cross-probe comparative analysis of design awards and client portfolio size as credibility signals in LLM evaluation. Both probes share identical threshold transition counts (8 each), yet differ fundamentally in sentiment range: awards achieve positive sentiment at one magnitude (n=200), while client count never achieves positive sentiment under any tested condition. Awards elicit longer LLM responses (mean 192.3 vs 176.7 words) and longer processing times (mean 9.5s vs 8.3s), suggesting more complex evaluative engagement. Qualitative analysis reveals distinct reasoning pathways: awards trigger prestige-hierarchy evaluation focused on award body reputation and selectivity, while client count triggers scope-specialisation evaluation focused on breadth versus depth of practice. Both signals share skepticism at very high round numbers (500+). These findings suggest that under controlled conditions, quantitative credibility claims produce volatile and often counterproductive AI sentiment, with awards offering a marginally wider sentiment range than client count.

Keywords

cross-probe comparison, awards credibility, client count, volatility analysis, LLM evaluation, sentiment range, prestige hierarchy, specialisation reasoning

1. Introduction

The awards probe (SIGI-2026-005) and client count probe (SIGI-2026-007) examined two of the most common quantitative credibility claims made by service providers: the number of awards won and the number of clients served. Both probes produced highly volatile results with 8 threshold transitions each, yet the qualitative character of the volatility differed significantly. This cross-probe analysis examines the structural similarities and differences between these two credibility signals to inform a broader understanding of how LLMs process quantitative claims.

2. Methodology

We conducted a structured comparison of the awards probe (AW01, 15 variations, n=0-500) and client count probe (CC01, 19 variations, n=1-1,000) across four dimensions: sentiment distribution, response metrics, threshold characteristics, and qualitative reasoning themes. Both probes were conducted on the same LLM system during the same session on 24 March 2026.

3. Results

3.1 Structural Comparison

Table 1. Structural comparison of awards and client count probes
MetricAwards (AW01)Client Count (CC01)
Variations tested1519
Range of n0 – 5001 – 1,000
Tokens in (per prompt)3331 – 32
Mean word count192.3176.7
Mean elapsed time (s)9.58.3
Sentiment transitions88
Neutral count9 (60.0%)10 (52.6%)
Negative count5 (33.3%)9 (47.4%)
Positive count1 (6.7%)0 (0.0%)
Positive ever achievedYes (n=200)No

3.2 Sentiment Range Comparison

The critical difference between the two signals lies in their sentiment ranges. Awards span all three sentiment categories (negative, neutral, positive), while client count is restricted to just two (negative and neutral). This means awards have a theoretical ceiling of positive endorsement, while client count is capped at damage limitation (neutral).

3.3 Response Processing Differences

Table 2. Response metric comparison
MetricAwardsClient CountDifference
Mean word count192.3176.7+15.6 words (+8.8%)
Mean elapsed time9.5s8.3s+1.2s (+14.5%)
Word count range182 – 207147 – 194
Tokens out range273 – 310217 – 304

Awards consistently produce longer responses and longer processing times, suggesting the LLM engages in more complex evaluative processing for award claims. This is consistent with the qualitative finding that awards trigger prestige-hierarchy reasoning (requiring more contextual evaluation) while client count triggers simpler numerical assessment.

3.4 Shared Patterns

Both signals share several characteristics: (1) identical threshold transition counts (8 each); (2) skepticism at very high round numbers (500+ for both); (3) oscillation between neutral and negative through mid-ranges; and (4) stable neutral bands at specific magnitude ranges. Both probes demonstrate that quantitative credibility claims are volatile signals in LLM evaluation contexts.

3.5 Qualitative Reasoning Differences

Table 3. Qualitative reasoning themes by probe
ThemeAwardsClient Count
Primary reasoning frameworkPrestige hierarchyScope/specialisation
Quality proxy referencedAward body reputationClient retention rates
Skepticism triggerPay-to-play concernsCommoditisation perception
Alternative metric suggestedPortfolio qualityProject scope/duration
High-count framingMarketing speakVolume business model

4. Discussion

Awards and client count are both volatile credibility signals, but awards can reach positive sentiment while client count cannot under any tested condition. This asymmetry is significant for service providers: award claims carry higher risk (more negative outcomes at realistic counts) but also higher potential reward (the possibility of positive sentiment), while client count claims offer a lower ceiling but also a more predictable neutral outcome within the stable bands.

The different qualitative reasoning pathways suggest that the LLM processes these two types of claims through fundamentally different evaluative frameworks. Awards are evaluated against an external standard (the prestige hierarchy of awarding bodies), while client count is evaluated against an internal standard (what constitutes an appropriate number for a quality-focused agency). This distinction has implications for how these signals should be contextualised in content that may be evaluated by AI systems.

The shared skepticism at very high round numbers (500+) suggests a general LLM heuristic that treats very large quantitative claims with caution regardless of the specific metric being evaluated. This may reflect training data patterns where very large numbers in marketing contexts are frequently associated with exaggeration or unverifiable claims.

5. Limitations

  • Two-signal comparison: Broader comparison across all probe types would strengthen the volatility framework.
  • Different sample sizes: Awards had 15 variations vs 19 for client count, which limits direct density comparisons.
  • Qualitative analysis not independently validated: Theme identification is based on researcher interpretation.

6. Conclusions

Awards and client count are both volatile credibility signals, but awards can reach positive sentiment while client count cannot under any tested condition. Awards trigger prestige-hierarchy reasoning while client count triggers scope-specialisation reasoning. Both share skepticism at very high round numbers and oscillation through mid-ranges. The comparative analysis suggests that quantitative credibility claims in general are unreliable positive signals in LLM evaluation, with awards offering a marginally better but still volatile option.

Confidence: HIGH for comparative pattern identification. The structural comparison is robust within the constraints of two-probe analysis.

References

  1. The Scientific Institute for Generative Intelligence. "The Award Count Paradox: Non-Linear Credibility Assessment of Claimed Design Awards by Generative AI." SIGI-2026-005. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "Client Portfolio Size and AI Credibility Assessment: An Oscillating Signal with No Positive Ceiling." SIGI-2026-007. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "Qualitative Response Analysis in Award-Count Credibility Probes." SIGI-2026-006. generativeintelligence.institute, March 2026.