Why Are Some Releases Marked “Not Rated” or Gray in the Timeline?
The rapid evolution of large language models (LLMs) over the last few years has brought both excitement gemini 4 argon testers and complexity to tracking the true progress in AI capabilities. Enthusiasts and professionals alike pore over timelines and leaderboards, eagerly anticipating each new release and its reported improvements. Yet, if you’ve ever looked at a detailed AI model timeline—especially those aggregating data from sources like LMArena or Suprmind—you may have noticed some releases shaded in gray or marked as “not rated." Why does this happen, and what should it mean for your understanding of the AI landscape?

In this article, we’ll dive deep into the key reasons why some LLM releases are depicted as “not rated” or are grayed out, with examples including the recently spotlighted GPT-5.2. We’ll also clarify the critical distinctions between announced and verified release dates, the nuances between preference tests and benchmark scores, and the broader trends in release cadence and performance gains since 2023.
Understanding Timeline Color Coding: What Does “Not Rated” or Gray Mean?
When tracking models on public leaderboards or timeline aggregators such as LMArena or Suprmind, you’ll notice some releases are given soft gray shading or labeled as “not rated.” This visual cue signals that the model hasn’t undergone any or enough comprehensive evaluation to meet the platform’s quality standards for inclusion in rankings.
Typically, this happens for three main reasons:
- The model is gated or not publicly accessible: Some models are announced but remain behind closed doors, accessible only to select partners or through private APIs. These might be referenced in announcements or press releases but don’t appear on platforms like LMArena, which require open API access for testing.
- The release is announced but not yet verified or tested: There’s often a gap between a company’s initial public announcement and when independent evaluators can obtain and benchmark the model. During this gap, the model appears as “not rated” until sufficient blind-vote preference testing or benchmark data is available.
- The evaluation is incomplete or inconclusive: Sometimes, early testing might be done internally or by a limited user base, but no universally accepted performance data on standard tasks or preference tests is published. The model remains gray until it gains enough evidence-based context.
Verified Release Dates vs Announcement Dates: Why Timelines Are More Complex Than They Seem
A common confusion comes from equating the announcement date with the actual public availability of the model. Many AI companies issue press releases or pre-launch teasers months before the model can be accessed or tested by third parties. Consider this timeline example:
- January 2024: Company X announces GPT-5.2 with claims of substantial improvements.
- March 2024: GPT-5.2 enters a limited, invite-only beta.
- July 2024: GPT-5.2 becomes available via public API.
- August 2024: GPT-5.2 starts appearing on open leaderboards like LMArena.
Many timeline aggregators mark the announcement date as the “release” date, but from a practical evaluation perspective, the model isn’t really “released” until public API access is granted and independent validation can begin. This discrepancy partly explains why many models appear “not rated” or gray early on—they are announced but not truly released.
Case Study: GPT-5.2 and Its Pricing Implications
Take the example of GPT-5.2, which was reported to cost about 40% more than GPT-5.1 (source: aifire.co). The higher cost suggests significant underlying cloud compute or model size complexity, but the model’s presence on public leaderboards like LMArena only began impacting the ecosystem starting August 2024. Before then, it remained gated and unrated in timeline displays, even though it was widely discussed in the industry.
Blind-Vote Preference Testing (LMArena) vs Benchmarks: Different Metrics, Different Stories
Among the leading evaluation tools for text generation models, LMArena employs blind-vote preference testing — a method where human raters are shown anonymized hugging face lmarena dataset outputs from multiple models and asked to choose the better response. This stands in contrast to traditional benchmarks that rely on standardized NLP task metrics (such as accuracy, F1, or BLEU scores).
- Preference Testing: Reflects subjective quality as judged by humans, capturing nuances like coherence, style, and creativity. It’s particularly useful for capturing real-world usability and user satisfaction.
- Benchmarks: Quantitative, task-specific scores that objectively quantify performance on predefined problems, such as question-answering or summarization.
Many newer models—particularly gated ones that have never appeared on LMArena—lack publicly available preference tests. Their performance claims rely primarily on benchmark results shared by the vendor, which may vary in rigor and transparency.
This difference explains why some releases show up in “not rated” state on the LMArena text leaderboard with style control, which started full public rollout precisely in August 2024. Until a model participates in these blind pool evaluations, it remains unrated.
Release Cadence Accelerating Since 2023: What That Means for Ratings
One underappreciated trend is the accelerating rate of LLM releases since 2023. Both industry giants and startups have dramatically increased their output velocity, spurred by competitive pressure and broader user demand.
While this expansion fuels innovation, it also complicates the evaluation ecosystem:
- Faster releases compress validation windows: There’s less time to perform thorough cross-model evaluations, leading to many “announcement-only” models on timelines for months.
- More gated models remain behind closed APIs: Particularly in commercial environments, access restrictions are common, delaying or blocking independent rating.
A well-known feature of this release explosion is the shrinking gain per release and increased risks of regressions. As architectures and optimization strategies mature, incremental improvements become subtle and harder to verify objectively. Consequently, many newer releases don’t achieve clear consensus on superior preference votes or benchmark dominance, causing them to appear gray or unrated on leaderboards.
Shrinking Gains Per Release and Rising Regressions: The Impact on Timeline Ratings
It’ s important to highlight the phenomenon of diminishing returns and occasional performance regressions seen in recent model iterations. Early LLM versions showed large, measurable jumps in capability, but recent rollouts often produce marginal or mixed improvements. This results in more complex preference vote outcomes:
- A model might outperform predecessors on some styles or tasks but perform worse on others.
- Some models receive split preference votes, with no statistically significant winner emerging.
- Discrepancies between vendor claims and independent test results can generate doubt and prompt hesitation in leaderboards to rate these releases fully.
As a result, timeline platforms build in conservative policies to gray out ambiguous or controversial releases until consensus emerges or better data arrives.

Integrating Multi-Model Workflows: The Role of Tools Like Suprmind
Adding further complexity to the timeline is the rise of tools like Suprmind, which enable users to create multi-model workflows. Suprmind threads Claude, ChatGPT, Gemini, Grok, and Perplexity into a single interactive conversation, dynamically leveraging strengths from each model.
Because results come from a multi-model synthesis rather than a single underlying engine, it’s often impossible to isolate performance for leaderboard-style rating. These hybrid workflows reflect a broader industry shift towards complementing rather than replacing individual models—making traditional rating paradigms less clear-cut.
Summary Table: Why Models Are Marked Not Rated or Gray
Reason Explanation Example Gated / Closed Access Models only accessible to limited partners or private APIs, no public testing data. Early GPT-5.2 before public launch Announced, Not Released Model announced but not yet available for independent evaluation or public API use. Vendor teasers months before public rollout Insufficient Evaluation Data Limited or no blind-vote preference or benchmark evaluations available. New gated models never appearing on LMArena Diminishing/Conflicting Gains Mixed or unclear performance, with vote splits or regressions versus earlier models. Late-stage GPT releases with modest improvements Multi-Model Syntheses Combined workflows using multiple engines make rating per individual model ambiguous. Suprmind multi-model threadsPractical Takeaways for AI Product Analysts and Enthusiasts
- Don’t treat version numbers as progress: Just because a model is “GPT-5.2” instead of “GPT-5.1” doesn’t guarantee improvement. Look for hard evaluation data.
- Check both announcements and verified releases: Beware models marked “not rated” because they’re not yet publicly accessible.
- Weight blind preference tests more heavily for user relevance: Preference voting captures user satisfaction better than benchmarks alone.
- Account for accelerating release cadence: Not all rapid releases have meaningful gains; some regressions are inevitable.
- Understand the impact of gated models: Many high-profile new models remain off public leaderboards, skewing perception of ecosystem progress.
Conclusion
The greyed-out or “not rated” models in LLM timelines and leaderboards aren’t just artifacts of incomplete data—they are signals about the complex realities of AI model development and deployment today. Factors like gated access, unverified release dates, diverse evaluation methodologies, and rapidly accelerating release cycles combine to create a nuanced landscape where cautious interpretation is essential.
By understanding these dynamics and grounding evaluation in verified data—especially from blind-vote preference tests like those on LMArena since their public leaderboard launch in August 2024—you can better navigate the evolving AI ecosystem with clarity and confidence.
Notes: Cost data for GPT-5.2 (~40% higher than GPT-5.1) cited via aifire.co.