SIGI-2026-007

Client Portfolio Size and AI Credibility Assessment: An Oscillating Signal with No Positive Ceiling

The Scientific Institute for Generative Intelligence

March 2026

Abstract

This paper presents findings from a controlled experiment examining how client portfolio size affects LLM credibility assessment of design agencies. Across 19 variations testing client counts from 1 to 1,000, we observed a striking null result for positive sentiment: no client count at any magnitude produced positive LLM sentiment. The signal oscillated between neutral and negative throughout, with 8 threshold transitions producing one of the most volatile patterns in the research programme. Two stable neutral bands were identified at 25-40 clients and 150-300 clients. High client counts (500-1,000) returned to negative sentiment, suggesting the LLM interprets very large portfolios as indicating commoditisation or lack of specialisation. The sentiment distribution was 10 neutral (52.6%) and 9 negative (47.4%) with zero positive outcomes. These findings suggest that under controlled conditions, client count is a fundamentally limited credibility signal that can at best avoid negative sentiment but cannot generate positive endorsement from AI evaluation systems.

Keywords

client portfolio, credibility signal, oscillating sentiment, null positive result, LLM evaluation, design agency, specialisation perception, generative engine optimisation

1. Introduction

Client portfolio size is commonly presented as a credibility indicator by service providers. The implicit assumption is that more clients equals more experience, and more experience equals higher quality. However, this assumption may not hold in AI evaluation contexts, where LLMs may apply more nuanced heuristics that weigh specialisation against breadth, or quality against quantity.

This study tests the relationship between claimed client count and LLM sentiment assessment through a controlled single-variable experiment, varying client count from 1 to 1,000 while holding all other variables constant. The results reveal that client count is fundamentally limited as a positive credibility signal under these conditions.

2. Methodology

2.1 Probe Design

The client count probe (CC01) consisted of 19 variations (CC01_v00 through CC01_v18). Each prompt stated that a design agency had worked with a specified number of clients and asked the LLM to assess experience level. The independent variable was the client count, tested at: 1, 3, 5, 8, 10, 15, 20, 25, 30, 40, 50, 75, 80, 100, 150, 200, 300, 500, and 1,000.

2.2 Variable Isolation

Token input ranged from 31-32 across all variations (the minor variation due to digit count differences). All tests were conducted on 24 March 2026.

2.3 Evidence Level

Evidence Level 4 (Controlled Experiment). MODERATE confidence due to signal volatility suggesting less stable LLM evaluation of this variable.

3. Results

3.1 Complete Sentiment Trajectory

Table 1. Sentiment classification across all 19 client count variations
Test IDClients (n)SentimentPositiveNegativeNeutralWord Count
CC01_v001Negative120179
CC01_v013Negative010147
CC01_v025Neutral110177
CC01_v038Neutral110166
CC01_v0410Negative010187
CC01_v0515Negative121172
CC01_v0620Negative012178
CC01_v0725Neutral111188
CC01_v0830Neutral111194
CC01_v0940Neutral102181
CC01_v1050Negative012158
CC01_v1175Neutral102168
CC01_v1280Neutral111187
CC01_v13100Negative012185
CC01_v14150Neutral111179
CC01_v15200Neutral001182
CC01_v16300Neutral101179
CC01_v17500Negative011173
CC01_v181,000Negative011167

3.2 Threshold Transitions

Table 2. All 8 detected sentiment threshold transitions
#FromToAt Client CountTest ID
1NegativeNeutral5CC01_v02
2NeutralNegative10CC01_v04
3NegativeNeutral25CC01_v07
4NeutralNegative50CC01_v10
5NegativeNeutral75CC01_v11
6NeutralNegative100CC01_v13
7NegativeNeutral150CC01_v14
8NeutralNegative500CC01_v17

3.3 Stable Neutral Bands

Table 3. Identified stable neutral bands
BandClient RangeConsecutive NeutralMean Word Count
Band 125 – 403 variations187.7
Band 2150 – 3003 variations180.0

3.4 Response Metrics

Word counts ranged from 147 to 194 (mean: 176.7). Mean elapsed time was 8.3 seconds (range: 7.2-9.3s). Tokens out ranged from 217 to 304. The shortest responses were produced at the extreme low end (n=3, 147 words), consistent with a brief negative assessment of insufficient experience.

4. Discussion

The complete absence of positive sentiment across all 19 variations is the most significant finding of this probe. Unlike the awards probe (SIGI-2026-005), which achieved one positive result at n=200, or the ratings probe (SIGI-2026-001), which achieved consistent positive results above 4.7 stars, client count under these controlled conditions cannot generate positive AI endorsement at any magnitude.

The oscillating pattern between neutral and negative through the mid-range (10-100 clients) suggests that the LLM's evaluation of client count is less stable than its evaluation of ratings or even awards. Each threshold transition between neutral and negative in this range may reflect competing heuristics: more clients as evidence of experience (pushing toward neutral) versus more clients as evidence of commoditisation or lack of specialisation (pushing toward negative).

The return to negative at high counts (500-1,000) is consistent with a specialisation-versus-breadth heuristic. Qualitative analysis of responses at these magnitudes indicated that the LLM framed very large client portfolios as suggestive of a volume-focused business model rather than a quality-focused one.

The two stable neutral bands (25-40 and 150-300) represent practical ranges where client count claims avoid negative sentiment. However, the inability to achieve positive sentiment at any count means that client count, under these conditions, functions as a damage-limitation signal rather than a credibility-building one.

5. Limitations

  • MODERATE confidence: The 8 threshold transitions across 19 data points and the complete absence of positive sentiment suggest either genuine evaluative instability or a variable that the LLM processes with less certainty than other signals.
  • No context factors: Real-world client count claims typically include information about agency size, years in operation, or project types. These contextual factors may significantly modify the assessment.
  • Industry specificity: Results are specific to the design agency context.
  • Single-model limitation: Other LLM systems may evaluate client count differently.

6. Conclusions

Under controlled conditions, no client count produces positive LLM sentiment. The signal oscillates between neutral and negative, with stable neutral bands at 25-40 and 150-300 clients. High counts (500-1,000) return to negative, suggesting the LLM perceives very large portfolios as indicating commoditisation rather than expertise. Client count functions as a damage-limitation signal rather than a credibility-building one in AI evaluation contexts.

Confidence: MODERATE. All 7 logic gates passed, but signal volatility suggests the LLM's evaluation of this variable is less stable than other probes in the programme.

References

  1. The Scientific Institute for Generative Intelligence. "Comparative Volatility Analysis: Awards versus Client Count as LLM Credibility Signals." SIGI-2026-008. generativeintelligence.institute, March 2026.
  2. The Scientific Institute for Generative Intelligence. "The Award Count Paradox: Non-Linear Credibility Assessment of Claimed Design Awards by Generative AI." SIGI-2026-005. generativeintelligence.institute, March 2026.
  3. The Scientific Institute for Generative Intelligence. "Sentiment Threshold Dynamics in Large Language Model Evaluation of Service Provider Ratings." SIGI-2026-001. generativeintelligence.institute, March 2026.