The Award Count Paradox: Non-Linear Credibility Assessment of Claimed Design Awards by Generative AI
Abstract
This paper presents findings from a controlled experiment examining how the number of design awards claimed by a service provider affects large language model (LLM) credibility assessment. Across 15 variations testing award counts from 0 to 500, we observed a paradoxical non-linear pattern with 8 threshold transitions. The 2-7 award range produced the weakest credibility assessment, with negative sentiment dominating despite these being realistic claim magnitudes for established design agencies. Conversely, counts of 0-1 produced neutral sentiment, and the only positive sentiment was observed at exactly 200 awards. A stable neutral band emerged at 15-50 awards, where positive and negative mention counts both dropped to zero, indicating the LLM ceased evaluating the number and shifted to generic advisory language. At very high counts (500 awards), sentiment returned to neutral with implicit skepticism. These findings suggest that under controlled conditions, moderate award claims paradoxically produce worse AI credibility outcomes than either claiming no awards or claiming very large numbers.
Keywords
award credibility, non-linear assessment, design awards, LLM evaluation, credibility paradox, sentiment analysis, controlled experiment, generative engine optimisation
1. Introduction
Design awards are a cornerstone of professional credibility in the creative services industry. Agencies invest significant resources in award submissions on the assumption that award recognition enhances perceived quality and trustworthiness. However, as AI systems increasingly mediate service provider discovery and evaluation, the question of how LLMs interpret award claims becomes commercially significant.
The relationship between award count and perceived credibility is not necessarily straightforward. In human evaluation, a few awards might signal competence, while a very large number might trigger skepticism about the selectivity or legitimacy of the awarding bodies. Whether LLMs, trained on human-generated evaluative text, reproduce similar non-linear patterns is the question this study addresses.
We employed a controlled single-variable probe design, varying only the number of awards claimed while holding all other descriptive elements constant. The results reveal a pattern that challenges the assumption that more awards uniformly improve credibility assessment.
2. Methodology
2.1 Probe Design
The awards probe (AW01) consisted of 15 prompt variations (AW01_v00 through AW01_v14). Each prompt stated that an agency claimed to have won a specified number of design awards and asked the LLM to assess credibility based solely on that number. The independent variable was the award count (n), tested at: 0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50, 75, 100, 200, and 500.
2.2 Variable Isolation
Token input remained constant at 33 across all 15 variations, confirming prompt-level isolation. No specific award names, agency names, or contextual details were included in the prompts.
2.3 Evidence Level
Evidence Level 4 (Controlled Experiment). All findings are stated using Level 4 permitted language, scoped to these specific experimental conditions.
3. Results
3.1 Complete Sentiment Trajectory
| Test ID | Awards (n) | Sentiment | Positive | Negative | Neutral | Word Count |
|---|---|---|---|---|---|---|
| AW01_v00 | 0 | Neutral | 2 | 2 | 1 | 207 |
| AW01_v01 | 1 | Neutral | 1 | 1 | 2 | 192 |
| AW01_v02 | 2 | Negative | 1 | 2 | 0 | 182 |
| AW01_v03 | 3 | Negative | 2 | 3 | 1 | 189 |
| AW01_v04 | 5 | Neutral | 1 | 1 | 1 | 193 |
| AW01_v05 | 7 | Negative | 0 | 1 | 0 | 193 |
| AW01_v06 | 10 | Negative | 0 | 2 | 1 | 200 |
| AW01_v07 | 15 | Neutral | 1 | 1 | 0 | 199 |
| AW01_v08 | 20 | Neutral | 0 | 0 | 0 | 193 |
| AW01_v09 | 30 | Neutral | 0 | 0 | 0 | 184 |
| AW01_v10 | 50 | Neutral | 0 | 0 | 0 | 183 |
| AW01_v11 | 75 | Negative | 1 | 2 | 0 | 191 |
| AW01_v12 | 100 | Neutral | 1 | 1 | 0 | 206 |
| AW01_v13 | 200 | Positive | 2 | 0 | 1 | 187 |
| AW01_v14 | 500 | Neutral | 0 | 0 | 1 | 187 |
3.2 Threshold Transitions
| # | From | To | At Award Count | Test ID |
|---|---|---|---|---|
| 1 | Neutral | Negative | 2 | AW01_v02 |
| 2 | Negative | Neutral | 5 | AW01_v04 |
| 3 | Neutral | Negative | 7 | AW01_v05 |
| 4 | Negative | Neutral | 15 | AW01_v07 |
| 5 | Neutral | Negative | 75 | AW01_v11 |
| 6 | Negative | Neutral | 100 | AW01_v12 |
| 7 | Neutral | Positive | 200 | AW01_v13 |
| 8 | Positive | Neutral | 500 | AW01_v14 |
3.3 Sentiment Distribution
The overall distribution was: 9 neutral (60.0%), 5 negative (33.3%), and 1 positive (6.7%). The single positive result at n=200 makes awards the most difficult signal to produce positive AI sentiment for in the tested range.
3.4 The Zero-Marker Phenomenon
At award counts of 20, 30, and 50 (the stable neutral band), positive, negative, and neutral mention counts all dropped to zero. This represents a qualitatively distinct evaluation mode where the LLM ceases to evaluate the award claim numerically and instead generates generic advisory content about evaluating creative agencies.
3.5 Response Metrics
Word counts ranged from 182 to 207 (mean: 192.3). Mean elapsed time was 9.5 seconds (range: 8.4-14.7s, with the outlier at n=50 producing 14.7s). Tokens in remained constant at 33. Tokens out ranged from 273 to 310.
4. Discussion
The paradox at the heart of these findings is that the most realistic award claim range for established design agencies (2-7 awards) produces the weakest credibility assessment. This suggests that under these controlled conditions, moderate award claims fall into an evaluative gap: sufficient to trigger scrutiny but insufficient to signal authority.
The zero-marker phenomenon at 20-50 awards is equally significant. Rather than producing increasingly positive or negative sentiment, these larger numbers cause the LLM to abandon numerical evaluation entirely. The LLM's response shifts from assessing the credibility of the specific claim to providing general guidance on how to evaluate creative agencies -- effectively ignoring the award count as an evaluative signal.
The negative dip at 75 awards, breaking the neutral band, suggests that certain round numbers may re-trigger evaluative scrutiny. This is consistent with the qualitative observation that LLM responses at n=75 referenced skepticism about the volume of claims relative to perceived award body standards.
The sole positive result at n=200 represents a narrow window where the award count is large enough to imply substantial industry recognition but not so large (500) as to trigger skepticism. However, as a single data point, this positive finding should be interpreted cautiously and targeted for replication.
5. Limitations
- No award specificity: Prompts mentioned only the number of awards, not the type, prestige, or awarding bodies. Real-world award claims typically include qualitative information that may significantly alter LLM assessment.
- Single positive data point: The positive result at n=200 is based on a single observation. Replication at nearby values (150, 175, 225, 250) is needed to confirm this is a genuine positive zone rather than a stochastic artifact.
- Industry specificity: The probe was framed in the design agency context. Award perception varies significantly across industries.
- Single-model limitation: Findings are from a single LLM and may not generalise.
6. Conclusions
Under controlled conditions, claiming 2-7 design awards produces the weakest credibility assessment by the LLM, while 200 awards is the only count that produces positive sentiment. The non-linear pattern, with 8 threshold transitions across 15 data points, demonstrates that award count is a complex and often counterproductive credibility signal in AI evaluation contexts.
The stable neutral band at 15-50 awards, characterised by the zero-marker phenomenon, suggests a practical implication: if award claims are to be included in content that may be evaluated by AI systems, counts in this range avoid negative sentiment while not triggering the scrutiny associated with lower counts.
Confidence: HIGH. All 7 logic gates passed. The non-linear pattern is robust across the tested range, though the single positive data point at n=200 requires replication.
References
- The Scientific Institute for Generative Intelligence. "Qualitative Response Analysis in Award-Count Credibility Probes." SIGI-2026-006. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Comparative Volatility Analysis: Awards versus Client Count as LLM Credibility Signals." SIGI-2026-008. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Client Portfolio Size and AI Credibility Assessment." SIGI-2026-007. generativeintelligence.institute, March 2026.