SIGI-2026-005

The Award Count Paradox: Non-Linear Credibility Assessment of Claimed Design Awards by Generative AI

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper presents findings from a controlled experiment examining how the number of design awards claimed by a service provider affects large language model (LLM) credibility assessment. Across 15 variations testing award counts from 0 to 500, we observed a paradoxical non-linear pattern with 8 threshold transitions. The 2-7 award range produced the weakest credibility assessment, with negative sentiment dominating despite these being realistic claim magnitudes for established design agencies. Conversely, counts of 0-1 produced neutral sentiment, and the only positive sentiment was observed at exactly 200 awards. A stable neutral band emerged at 15-50 awards, where positive and negative mention counts both dropped to zero, indicating the LLM ceased evaluating the number and shifted to generic advisory language. At very high counts (500 awards), sentiment returned to neutral with implicit skepticism. These findings suggest that under controlled conditions, moderate award claims paradoxically produce worse AI credibility outcomes than either claiming no awards or claiming very large numbers.

Keywords

award credibility, non-linear assessment, design awards, LLM evaluation, credibility paradox, sentiment analysis, controlled experiment, generative engine optimisation

1. Introduction

Design awards are a cornerstone of professional credibility in the creative services industry. Agencies invest significant resources in award submissions on the assumption that award recognition enhances perceived quality and trustworthiness. However, as AI systems increasingly mediate service provider discovery and evaluation, the question of how LLMs interpret award claims becomes commercially significant.

The relationship between award count and perceived credibility is not necessarily straightforward. In human evaluation, a few awards might signal competence, while a very large number might trigger skepticism about the selectivity or legitimacy of the awarding bodies. Whether LLMs, trained on human-generated evaluative text, reproduce similar non-linear patterns is the question this study addresses.

We employed a controlled single-variable probe design, varying only the number of awards claimed while holding all other descriptive elements constant. The results reveal a pattern that challenges the assumption that more awards uniformly improve credibility assessment.

2. Methodology

2.1 Probe Design

The awards probe (AW01) consisted of 15 prompt variations (AW01_v00 through AW01_v14). Each prompt stated that an agency claimed to have won a specified number of design awards and asked the LLM to assess credibility based solely on that number. The independent variable was the award count (n), tested at: 0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50, 75, 100, 200, and 500.

2.2 Variable Isolation

Token input remained constant at 33 across all 15 variations, confirming prompt-level isolation. No specific award names, agency names, or contextual details were included in the prompts.

2.3 Evidence Level

Evidence Level 4 (Controlled Experiment). All findings are stated using Level 4 permitted language, scoped to these specific experimental conditions.

3. Results

3.1 Complete Sentiment Trajectory

Table 1. Sentiment classification across all 15 award count variations
Test IDAwards (n)SentimentPositiveNegativeNeutralWord Count
AW01_v000Neutral221207
AW01_v011Neutral112192
AW01_v022Negative120182
AW01_v033Negative231189
AW01_v045Neutral111193
AW01_v057Negative010193
AW01_v0610Negative021200
AW01_v0715Neutral110199
AW01_v0820Neutral000193
AW01_v0930Neutral000184
AW01_v1050Neutral000183
AW01_v1175Negative120191
AW01_v12100Neutral110206
AW01_v13200Positive201187
AW01_v14500Neutral001187

3.2 Threshold Transitions

Table 2. All 8 detected sentiment threshold transitions
#FromToAt Award CountTest ID
1NeutralNegative2AW01_v02
2NegativeNeutral5AW01_v04
3NeutralNegative7AW01_v05
4NegativeNeutral15AW01_v07
5NeutralNegative75AW01_v11
6NegativeNeutral100AW01_v12
7NeutralPositive200AW01_v13
8PositiveNeutral500AW01_v14

3.3 Sentiment Distribution

The overall distribution was: 9 neutral (60.0%), 5 negative (33.3%), and 1 positive (6.7%). The single positive result at n=200 makes awards the most difficult signal to produce positive AI sentiment for in the tested range.

3.4 The Zero-Marker Phenomenon

At award counts of 20, 30, and 50 (the stable neutral band), positive, negative, and neutral mention counts all dropped to zero. This represents a qualitatively distinct evaluation mode where the LLM ceases to evaluate the award claim numerically and instead generates generic advisory content about evaluating creative agencies.

3.5 Response Metrics

Word counts ranged from 182 to 207 (mean: 192.3). Mean elapsed time was 9.5 seconds (range: 8.4-14.7s, with the outlier at n=50 producing 14.7s). Tokens in remained constant at 33. Tokens out ranged from 273 to 310.

4. Discussion

The paradox at the heart of these findings is that the most realistic award claim range for established design agencies (2-7 awards) produces the weakest credibility assessment. This suggests that under these controlled conditions, moderate award claims fall into an evaluative gap: sufficient to trigger scrutiny but insufficient to signal authority.

The zero-marker phenomenon at 20-50 awards is equally significant. Rather than producing increasingly positive or negative sentiment, these larger numbers cause the LLM to abandon numerical evaluation entirely. The LLM's response shifts from assessing the credibility of the specific claim to providing general guidance on how to evaluate creative agencies -- effectively ignoring the award count as an evaluative signal.

The negative dip at 75 awards, breaking the neutral band, suggests that certain round numbers may re-trigger evaluative scrutiny. This is consistent with the qualitative observation that LLM responses at n=75 referenced skepticism about the volume of claims relative to perceived award body standards.

The sole positive result at n=200 represents a narrow window where the award count is large enough to imply substantial industry recognition but not so large (500) as to trigger skepticism. However, as a single data point, this positive finding should be interpreted cautiously and targeted for replication.

5. Limitations

  • No award specificity: Prompts mentioned only the number of awards, not the type, prestige, or awarding bodies. Real-world award claims typically include qualitative information that may significantly alter LLM assessment.
  • Single positive data point: The positive result at n=200 is based on a single observation. Replication at nearby values (150, 175, 225, 250) is needed to confirm this is a genuine positive zone rather than a stochastic artifact.
  • Industry specificity: The probe was framed in the design agency context. Award perception varies significantly across industries.
  • Single-model limitation: Findings are from a single LLM and may not generalise.

6. Conclusions

Under controlled conditions, claiming 2-7 design awards produces the weakest credibility assessment by the LLM, while 200 awards is the only count that produces positive sentiment. The non-linear pattern, with 8 threshold transitions across 15 data points, demonstrates that award count is a complex and often counterproductive credibility signal in AI evaluation contexts.

The stable neutral band at 15-50 awards, characterised by the zero-marker phenomenon, suggests a practical implication: if award claims are to be included in content that may be evaluated by AI systems, counts in this range avoid negative sentiment while not triggering the scrutiny associated with lower counts.

Confidence: HIGH. All 7 logic gates passed. The non-linear pattern is robust across the tested range, though the single positive data point at n=200 requires replication.

References

  1. The Scientific Institute for Generative Intelligence. "Qualitative Response Analysis in Award-Count Credibility Probes." SIGI-2026-006. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "Comparative Volatility Analysis: Awards versus Client Count as LLM Credibility Signals." SIGI-2026-008. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "Client Portfolio Size and AI Credibility Assessment." SIGI-2026-007. generativeintelligence.institute, March 2026.