Comparative Volatility Analysis: Awards versus Client Count as LLM Credibility Signals
Abstract
This paper presents a cross-probe comparative analysis of design awards and client portfolio size as credibility signals in LLM evaluation. Both probes share identical threshold transition counts (8 each), yet differ fundamentally in sentiment range: awards achieve positive sentiment at one magnitude (n=200), while client count never achieves positive sentiment under any tested condition. Awards elicit longer LLM responses (mean 192.3 vs 176.7 words) and longer processing times (mean 9.5s vs 8.3s), suggesting more complex evaluative engagement. Qualitative analysis reveals distinct reasoning pathways: awards trigger prestige-hierarchy evaluation focused on award body reputation and selectivity, while client count triggers scope-specialisation evaluation focused on breadth versus depth of practice. Both signals share skepticism at very high round numbers (500+). These findings suggest that under controlled conditions, quantitative credibility claims produce volatile and often counterproductive AI sentiment, with awards offering a marginally wider sentiment range than client count.
Keywords
cross-probe comparison, awards credibility, client count, volatility analysis, LLM evaluation, sentiment range, prestige hierarchy, specialisation reasoning
1. Introduction
The awards probe (SIGI-2026-005) and client count probe (SIGI-2026-007) examined two of the most common quantitative credibility claims made by service providers: the number of awards won and the number of clients served. Both probes produced highly volatile results with 8 threshold transitions each, yet the qualitative character of the volatility differed significantly. This cross-probe analysis examines the structural similarities and differences between these two credibility signals to inform a broader understanding of how LLMs process quantitative claims.
2. Methodology
We conducted a structured comparison of the awards probe (AW01, 15 variations, n=0-500) and client count probe (CC01, 19 variations, n=1-1,000) across four dimensions: sentiment distribution, response metrics, threshold characteristics, and qualitative reasoning themes. Both probes were conducted on the same LLM system during the same session on 24 March 2026.
3. Results
3.1 Structural Comparison
| Metric | Awards (AW01) | Client Count (CC01) |
|---|---|---|
| Variations tested | 15 | 19 |
| Range of n | 0 – 500 | 1 – 1,000 |
| Tokens in (per prompt) | 33 | 31 – 32 |
| Mean word count | 192.3 | 176.7 |
| Mean elapsed time (s) | 9.5 | 8.3 |
| Sentiment transitions | 8 | 8 |
| Neutral count | 9 (60.0%) | 10 (52.6%) |
| Negative count | 5 (33.3%) | 9 (47.4%) |
| Positive count | 1 (6.7%) | 0 (0.0%) |
| Positive ever achieved | Yes (n=200) | No |
3.2 Sentiment Range Comparison
The critical difference between the two signals lies in their sentiment ranges. Awards span all three sentiment categories (negative, neutral, positive), while client count is restricted to just two (negative and neutral). This means awards have a theoretical ceiling of positive endorsement, while client count is capped at damage limitation (neutral).
3.3 Response Processing Differences
| Metric | Awards | Client Count | Difference |
|---|---|---|---|
| Mean word count | 192.3 | 176.7 | +15.6 words (+8.8%) |
| Mean elapsed time | 9.5s | 8.3s | +1.2s (+14.5%) |
| Word count range | 182 – 207 | 147 – 194 | |
| Tokens out range | 273 – 310 | 217 – 304 |
Awards consistently produce longer responses and longer processing times, suggesting the LLM engages in more complex evaluative processing for award claims. This is consistent with the qualitative finding that awards trigger prestige-hierarchy reasoning (requiring more contextual evaluation) while client count triggers simpler numerical assessment.
3.4 Shared Patterns
Both signals share several characteristics: (1) identical threshold transition counts (8 each); (2) skepticism at very high round numbers (500+ for both); (3) oscillation between neutral and negative through mid-ranges; and (4) stable neutral bands at specific magnitude ranges. Both probes demonstrate that quantitative credibility claims are volatile signals in LLM evaluation contexts.
3.5 Qualitative Reasoning Differences
| Theme | Awards | Client Count |
|---|---|---|
| Primary reasoning framework | Prestige hierarchy | Scope/specialisation |
| Quality proxy referenced | Award body reputation | Client retention rates |
| Skepticism trigger | Pay-to-play concerns | Commoditisation perception |
| Alternative metric suggested | Portfolio quality | Project scope/duration |
| High-count framing | Marketing speak | Volume business model |
4. Discussion
Awards and client count are both volatile credibility signals, but awards can reach positive sentiment while client count cannot under any tested condition. This asymmetry is significant for service providers: award claims carry higher risk (more negative outcomes at realistic counts) but also higher potential reward (the possibility of positive sentiment), while client count claims offer a lower ceiling but also a more predictable neutral outcome within the stable bands.
The different qualitative reasoning pathways suggest that the LLM processes these two types of claims through fundamentally different evaluative frameworks. Awards are evaluated against an external standard (the prestige hierarchy of awarding bodies), while client count is evaluated against an internal standard (what constitutes an appropriate number for a quality-focused agency). This distinction has implications for how these signals should be contextualised in content that may be evaluated by AI systems.
The shared skepticism at very high round numbers (500+) suggests a general LLM heuristic that treats very large quantitative claims with caution regardless of the specific metric being evaluated. This may reflect training data patterns where very large numbers in marketing contexts are frequently associated with exaggeration or unverifiable claims.
5. Limitations
- Two-signal comparison: Broader comparison across all probe types would strengthen the volatility framework.
- Different sample sizes: Awards had 15 variations vs 19 for client count, which limits direct density comparisons.
- Qualitative analysis not independently validated: Theme identification is based on researcher interpretation.
6. Conclusions
Awards and client count are both volatile credibility signals, but awards can reach positive sentiment while client count cannot under any tested condition. Awards trigger prestige-hierarchy reasoning while client count triggers scope-specialisation reasoning. Both share skepticism at very high round numbers and oscillation through mid-ranges. The comparative analysis suggests that quantitative credibility claims in general are unreliable positive signals in LLM evaluation, with awards offering a marginally better but still volatile option.
Confidence: HIGH for comparative pattern identification. The structural comparison is robust within the constraints of two-probe analysis.
References
- The Scientific Institute for Generative Intelligence. "The Award Count Paradox: Non-Linear Credibility Assessment of Claimed Design Awards by Generative AI." SIGI-2026-005. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Client Portfolio Size and AI Credibility Assessment: An Oscillating Signal with No Positive Ceiling." SIGI-2026-007. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Qualitative Response Analysis in Award-Count Credibility Probes." SIGI-2026-006. generativeintelligence.institute, March 2026.