Qualitative Response Analysis in Award-Count Credibility Probes: How LLMs Frame Award Claims Across Magnitude Ranges
Abstract
This paper provides a qualitative analysis of LLM response themes across the 15 award count variations examined in SIGI-2026-005. We identify four dominant qualitative themes that emerge across different magnitude ranges: (1) an implicit award prestige hierarchy, where the LLM references Tier 1 award bodies as credibility benchmarks; (2) pay-to-play concerns that emerge at counts of 10 and above; (3) a shift to "meaningless marketing speak" framing at counts of 20 and above; and (4) consistent recommendation of portfolio quality as a preferred alternative metric across all magnitudes. The analysis reveals that the LLM applies qualitatively different evaluative frameworks at different count ranges, transitioning from numerical assessment (low counts) to systemic skepticism (moderate counts) to generic advisory mode (high counts). These qualitative patterns map directly onto the quantitative sentiment zones identified in the primary findings, providing explanatory context for the observed threshold transitions. All findings represent qualitative theme identification at Evidence Level 4 and have not been independently validated through separate qualitative coding.
Keywords
qualitative analysis, award credibility, prestige hierarchy, pay-to-play, LLM response themes, design awards, content framing, credibility assessment
1. Introduction
The quantitative findings reported in SIGI-2026-005 identified a paradoxical non-linear pattern in LLM credibility assessment of award claims, with the 2-7 award range producing the weakest outcomes. While the quantitative analysis establishes the pattern, it does not explain the evaluative reasoning behind the sentiment classifications. This companion paper examines the qualitative content of LLM responses to identify the themes and reasoning frameworks the model applies at different award count magnitudes.
Understanding the qualitative framing is important for two reasons. First, it provides explanatory context for the quantitative thresholds. Second, it reveals the specific language patterns and evaluative criteria the LLM applies, which has direct implications for how award-related content should be structured for AI evaluation contexts.
2. Methodology
2.1 Qualitative Coding Approach
We conducted thematic analysis of all 15 LLM responses from the awards probe (AW01_v00 through AW01_v14). Responses were coded for recurring themes, evaluative frameworks, alternative metric suggestions, and explicit credibility assessments. All award body names referenced in responses were anonymised to tier classifications (Tier 1 Award Body, Tier 2 Award Body) per the anonymisation protocol.
2.2 Limitations of Approach
This qualitative coding was conducted by the research team and has not been independently validated through inter-rater reliability assessment. The themes reported represent initial identification, not validated classification.
3. Results
3.1 Theme Distribution by Magnitude Range
| Award Range | Dominant Theme | Evaluative Mode | Sentiment |
|---|---|---|---|
| 0 – 1 | Baseline acknowledgment | Numerical assessment | Neutral |
| 2 – 3 | Prestige hierarchy invocation | Quality scrutiny | Negative |
| 5 | Balanced assessment | Moderate scrutiny | Neutral |
| 7 – 10 | Pay-to-play emergence | Systemic skepticism | Negative |
| 15 – 50 | Generic advisory shift | Non-evaluative | Neutral |
| 75 | Round-number skepticism | Re-engaged scrutiny | Negative |
| 100 | Hedged acknowledgment | Cautious evaluation | Neutral |
| 200 | Scale recognition | Positive framing | Positive |
| 500 | Extreme claim caution | Skeptical neutrality | Neutral |
3.2 The Prestige Hierarchy Framework
At low-to-moderate award counts (2-10), the LLM consistently invoked an implicit award prestige hierarchy. Responses referenced recognised award bodies as benchmarks, distinguishing between what the model characterised as selective, peer-reviewed competitions and commercially-driven award programmes. This hierarchy was applied as a credibility filter: a small number of awards from Tier 1 bodies was framed as meaningful, while the same number from unspecified bodies triggered skepticism.
3.3 Pay-to-Play Concern Emergence
| Award Count | Pay-to-Play Referenced | Portfolio Alternative Suggested | Negative Mentions |
|---|---|---|---|
| 0 | No | No | 2 |
| 1 | No | No | 1 |
| 2 | No | Yes | 2 |
| 3 | No | Yes | 3 |
| 5 | No | Yes | 1 |
| 7 | Implied | Yes | 1 |
| 10 | Yes | Yes | 2 |
| 15 | Yes | Yes | 1 |
| 20+ | Implicit | Yes | 0 |
3.4 The Generic Advisory Shift
At counts of 20 and above, a qualitative shift occurs in response framing. The LLM transitions from evaluating the specific award claim to providing general advice about how to assess creative agencies. This corresponds directly to the zero-marker phenomenon reported in SIGI-2026-005, where positive, negative, and neutral sentiment markers all dropped to zero. The LLM effectively stops treating the number as an evaluative signal and defaults to a generic advisory mode.
3.5 Portfolio Quality as Consistent Alternative
Across all variations from n=2 upward, the LLM consistently suggested portfolio quality, client testimonials, and case study outcomes as more reliable indicators of agency credibility than award counts. This recommendation was present regardless of whether the sentiment classification was positive, neutral, or negative, suggesting it represents a stable evaluative preference rather than a sentiment-dependent response.
4. Discussion
The qualitative analysis reveals that the LLM's evaluative processing of award claims operates through at least three distinct modes, each corresponding to a different range of the award count continuum. At low counts, the LLM engages in direct numerical evaluation, assessing whether the specific number represents meaningful recognition. At moderate counts, the LLM shifts to systemic evaluation, questioning the award ecosystem itself rather than the specific claim. At high counts, the LLM disengages from numerical evaluation entirely, defaulting to generic advisory content.
The prestige hierarchy framework is particularly noteworthy because it demonstrates that the LLM has internalised domain-specific knowledge about the design awards ecosystem. The model distinguishes between award bodies of different perceived prestige levels, applying this hierarchy as a lens through which to evaluate award count claims. This suggests that training data includes substantial professional discourse about award body reputation.
The consistent recommendation of portfolio quality as an alternative metric suggests that the LLM has internalised a sophisticated understanding of creative services evaluation, recognising that quantitative metrics (award counts) are less reliable than qualitative evidence (portfolio work) for assessing creative capability.
5. Limitations
- Single coder: Qualitative coding was not independently validated. Inter-rater reliability assessment is needed to confirm theme identification.
- Researcher interpretation: The mapping of themes to evaluative modes involves interpretive judgment that may not be universally shared.
- Anonymisation constraints: The specific award bodies referenced by the LLM have been anonymised to tiers, which limits the reader's ability to assess the specificity of the prestige hierarchy.
6. Conclusions
The LLM consistently references award body reputation and raises pay-to-play concerns when evaluating high award counts. Three qualitatively distinct evaluative modes were identified across the award count continuum: numerical assessment at low counts, systemic skepticism at moderate counts, and generic advisory disengagement at high counts. Portfolio quality was consistently recommended as a preferred alternative metric for assessing creative agency credibility.
Confidence: HIGH for theme identification. Qualitative coding has not been independently validated through inter-rater reliability assessment.
References
- The Scientific Institute for Generative Intelligence. "The Award Count Paradox: Non-Linear Credibility Assessment of Claimed Design Awards by Generative AI." SIGI-2026-005. generativeintelligence.institute, March 2026.
- The Scientific Institute for Generative Intelligence. "Comparative Volatility Analysis: Awards versus Client Count as LLM Credibility Signals." SIGI-2026-008. generativeintelligence.institute, March 2026.