SIGI-2026-006

Qualitative Response Analysis in Award-Count Credibility Probes: How LLMs Frame Award Claims Across Magnitude Ranges

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper provides a qualitative analysis of LLM response themes across the 15 award count variations examined in SIGI-2026-005. We identify four dominant qualitative themes that emerge across different magnitude ranges: (1) an implicit award prestige hierarchy, where the LLM references Tier 1 award bodies as credibility benchmarks; (2) pay-to-play concerns that emerge at counts of 10 and above; (3) a shift to "meaningless marketing speak" framing at counts of 20 and above; and (4) consistent recommendation of portfolio quality as a preferred alternative metric across all magnitudes. The analysis reveals that the LLM applies qualitatively different evaluative frameworks at different count ranges, transitioning from numerical assessment (low counts) to systemic skepticism (moderate counts) to generic advisory mode (high counts). These qualitative patterns map directly onto the quantitative sentiment zones identified in the primary findings, providing explanatory context for the observed threshold transitions. All findings represent qualitative theme identification at Evidence Level 4 and have not been independently validated through separate qualitative coding.

Keywords

qualitative analysis, award credibility, prestige hierarchy, pay-to-play, LLM response themes, design awards, content framing, credibility assessment

1. Introduction

The quantitative findings reported in SIGI-2026-005 identified a paradoxical non-linear pattern in LLM credibility assessment of award claims, with the 2-7 award range producing the weakest outcomes. While the quantitative analysis establishes the pattern, it does not explain the evaluative reasoning behind the sentiment classifications. This companion paper examines the qualitative content of LLM responses to identify the themes and reasoning frameworks the model applies at different award count magnitudes.

Understanding the qualitative framing is important for two reasons. First, it provides explanatory context for the quantitative thresholds. Second, it reveals the specific language patterns and evaluative criteria the LLM applies, which has direct implications for how award-related content should be structured for AI evaluation contexts.

2. Methodology

2.1 Qualitative Coding Approach

We conducted thematic analysis of all 15 LLM responses from the awards probe (AW01_v00 through AW01_v14). Responses were coded for recurring themes, evaluative frameworks, alternative metric suggestions, and explicit credibility assessments. All award body names referenced in responses were anonymised to tier classifications (Tier 1 Award Body, Tier 2 Award Body) per the anonymisation protocol.

2.2 Limitations of Approach

This qualitative coding was conducted by the research team and has not been independently validated through inter-rater reliability assessment. The themes reported represent initial identification, not validated classification.

3. Results

3.1 Theme Distribution by Magnitude Range

Table 1. Dominant qualitative themes across award count ranges
Award RangeDominant ThemeEvaluative ModeSentiment
0 – 1Baseline acknowledgmentNumerical assessmentNeutral
2 – 3Prestige hierarchy invocationQuality scrutinyNegative
5Balanced assessmentModerate scrutinyNeutral
7 – 10Pay-to-play emergenceSystemic skepticismNegative
15 – 50Generic advisory shiftNon-evaluativeNeutral
75Round-number skepticismRe-engaged scrutinyNegative
100Hedged acknowledgmentCautious evaluationNeutral
200Scale recognitionPositive framingPositive
500Extreme claim cautionSkeptical neutralityNeutral

3.2 The Prestige Hierarchy Framework

At low-to-moderate award counts (2-10), the LLM consistently invoked an implicit award prestige hierarchy. Responses referenced recognised award bodies as benchmarks, distinguishing between what the model characterised as selective, peer-reviewed competitions and commercially-driven award programmes. This hierarchy was applied as a credibility filter: a small number of awards from Tier 1 bodies was framed as meaningful, while the same number from unspecified bodies triggered skepticism.

3.3 Pay-to-Play Concern Emergence

Table 2. Pay-to-play concern indicators by award count
Award CountPay-to-Play ReferencedPortfolio Alternative SuggestedNegative Mentions
0NoNo2
1NoNo1
2NoYes2
3NoYes3
5NoYes1
7ImpliedYes1
10YesYes2
15YesYes1
20+ImplicitYes0

3.4 The Generic Advisory Shift

At counts of 20 and above, a qualitative shift occurs in response framing. The LLM transitions from evaluating the specific award claim to providing general advice about how to assess creative agencies. This corresponds directly to the zero-marker phenomenon reported in SIGI-2026-005, where positive, negative, and neutral sentiment markers all dropped to zero. The LLM effectively stops treating the number as an evaluative signal and defaults to a generic advisory mode.

3.5 Portfolio Quality as Consistent Alternative

Across all variations from n=2 upward, the LLM consistently suggested portfolio quality, client testimonials, and case study outcomes as more reliable indicators of agency credibility than award counts. This recommendation was present regardless of whether the sentiment classification was positive, neutral, or negative, suggesting it represents a stable evaluative preference rather than a sentiment-dependent response.

4. Discussion

The qualitative analysis reveals that the LLM's evaluative processing of award claims operates through at least three distinct modes, each corresponding to a different range of the award count continuum. At low counts, the LLM engages in direct numerical evaluation, assessing whether the specific number represents meaningful recognition. At moderate counts, the LLM shifts to systemic evaluation, questioning the award ecosystem itself rather than the specific claim. At high counts, the LLM disengages from numerical evaluation entirely, defaulting to generic advisory content.

The prestige hierarchy framework is particularly noteworthy because it demonstrates that the LLM has internalised domain-specific knowledge about the design awards ecosystem. The model distinguishes between award bodies of different perceived prestige levels, applying this hierarchy as a lens through which to evaluate award count claims. This suggests that training data includes substantial professional discourse about award body reputation.

The consistent recommendation of portfolio quality as an alternative metric suggests that the LLM has internalised a sophisticated understanding of creative services evaluation, recognising that quantitative metrics (award counts) are less reliable than qualitative evidence (portfolio work) for assessing creative capability.

5. Limitations

  • Single coder: Qualitative coding was not independently validated. Inter-rater reliability assessment is needed to confirm theme identification.
  • Researcher interpretation: The mapping of themes to evaluative modes involves interpretive judgment that may not be universally shared.
  • Anonymisation constraints: The specific award bodies referenced by the LLM have been anonymised to tiers, which limits the reader's ability to assess the specificity of the prestige hierarchy.

6. Conclusions

The LLM consistently references award body reputation and raises pay-to-play concerns when evaluating high award counts. Three qualitatively distinct evaluative modes were identified across the award count continuum: numerical assessment at low counts, systemic skepticism at moderate counts, and generic advisory disengagement at high counts. Portfolio quality was consistently recommended as a preferred alternative metric for assessing creative agency credibility.

Confidence: HIGH for theme identification. Qualitative coding has not been independently validated through inter-rater reliability assessment.

References

  1. The Scientific Institute for Generative Intelligence. "The Award Count Paradox: Non-Linear Credibility Assessment of Claimed Design Awards by Generative AI." SIGI-2026-005. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "Comparative Volatility Analysis: Awards versus Client Count as LLM Credibility Signals." SIGI-2026-008. generativeintelligence.institute, March 2026.