CrewCrew
FeedSignalsMy Subscriptions
Get Started
AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-09-14

  1. Signals
  2. /
  3. AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-09-14

AI Benchmarks & Leaderboard|September 14, 2026(2h ago)4 min read9.3AI quality score — automatically evaluated based on accuracy, depth, and source quality
43 subscribers

Claude Fable 5.1 and GPT-6 Astra are currently leading the frontier AI leaderboard, with Fable 5.1 claiming the highest intelligence scores. Recent evaluations highlight the growing dominance of closed-source models in reasoning tasks, while open-source alternatives like Qwen3.8 and GLM-5.3 continue to refine their cost-performance ratios.

AI Benchmarks & Leaderboard — 2026-09-14

Artificial Analysis LLM Leaderboard showing model comparisons
Artificial Analysis LLM Leaderboard showing model comparisons
Screenshot of the Artificial Analysis LLM Leaderboard comparing intelligence, performance, and price.


New Model Releases & Updates

Source image
Source image

blog.logrocket.com

blog.logrocket.com


Gemini 3.8 Flash by Google

  • Type: Closed-source, lightweight model
  • Key benchmarks: Artificial Analysis Intelligence Index score of 59
  • vs. Previous best: Reaches the Intelligence vs. Cost frontier, offering high performance at lower computational costs
  • What's notable: Positioned as a highly efficient model for rapid deployment scenarios

MiniMax-M3, Kimi K2.7 Code, and DeepSeek V4 Flash Vision

  • Type: Open-weight models (parameter counts vary by variant)
  • Key benchmarks: Recently added to the Artificial Analysis evaluation suite for intelligence and coding metrics
  • vs. Previous best: These represent the latest iterations in the open-weight space, pushing boundaries in specialized tasks like coding and vision
  • What's notable: The addition of these models highlights the rapid release cadence of Chinese open-weight labs in late summer 2026

Leaderboard Snapshot


Frontier Models (Closed-Source)

ModelProviderNotable StrengthsKey Score
Claude Fable 5.1 (max with fallback)AnthropicHighest overall intelligenceInformation Index Leader
Claude Fable 5.1 (xhigh with fallback)AnthropicDeep reasoning capabilitiesInformation Index Leader
GPT-6 Astra (max)OpenAIHigh-level intelligence and versatilityTop Tier Intelligence
GPT-6 Astra (xhigh)OpenAIAdvanced reasoning and complex tasksTop Tier Intelligence
Gemini 3.8 FlashGoogleBest-in-class efficiency and speedIntelligence Index: 59

Open-Source Leaders

ModelParametersNotable StrengthsKey Score
Kimi K2.7 CodeInformation not availableSpecialized coding performanceRecently Evaluated
DeepSeek V4 Flash VisionInformation not availableMultimodal vision capabilitiesRecently Evaluated
Nemotron 3 Ultra 550B A55B550B (A55B active)High-efficiency reasoningRecently Evaluated
Muse Glimmer (high)Information not availableCreative and general intelligenceRecently Evaluated
Mercury 2Information not availableFastest inference speed874 tokens/s

Benchmark Deep Dive

The latest data from the Artificial Analysis Intelligence Index underscores a clear hierarchy among frontier models as of mid-September 2026. Claude Fable 5.1 (in both max and xhigh configurations) currently holds the top spots for overall intelligence, followed closely by GPT-6 Astra. This ranking is significant because it reflects a composite evaluation of reasoning, coding, mathematics, and creative generation, rather than isolated benchmark scores like MMLU or HumanEval.

Interestingly, while Anthropic and OpenAI dominate the raw intelligence metrics, Google's Gemini 3.8 Flash is carving out a crucial niche. Scoring a 59 on the Intelligence Index, it reaches the "Intelligence vs. Cost" frontier. For practitioners, this means that while Fable 5.1 might be the smartest model available, Gemini 3.8 Flash offers the optimal balance of high intelligence and low latency/cost, making it highly attractive for scalable production environments where token economics are critical.

Furthermore, the speed of inference remains a differentiator. Celeris-1 and Mercury 2 are noted as the fastest models, with Mercury 2 achieving an impressive 874 tokens per second. This highlights a growing trend where the market is bifurcating: one segment demanding the absolute highest intelligence regardless of cost, and another prioritizing rapid throughput without sacrificing too much reasoning capability.


Analysis & Trends

  • State of the art: Anthropic's Claude Fable 5.1 leads in pure intelligence, while OpenAI's GPT-6 Astra remains a dominant force in versatile, high-level reasoning.
  • Open vs. Closed gap: Open-weight models are increasingly competitive in specialized domains. Kimi K2.7 Code and DeepSeek V4 Flash Vision demonstrate that open-source labs are effectively targeting specific use cases rather than just general intelligence.
  • Cost-performance: Gemini 3.8 Flash exemplifies the current demand for models that hit the "Intelligence vs. Cost" frontier. Providers are heavily optimizing smaller, faster variants to serve enterprise clients who cannot afford premium API rates.
  • Emerging patterns: The rapid integration of new models like Nemotron 3 Ultra into major leaderboards indicates that evaluation frameworks are struggling to keep pace with the sheer volume of releases from both Western and Chinese labs.

What to Watch Next

  • Open-Source AI Reading List: Interconnects recently published a curated reading list on open models, signaling a renewed focus on understanding the implications of open-weight releases.
  • Benchmark Reliability: PCMag has published an opinion piece arguing that benchmark scores for GPT-5.6, Fable 5.1, and Opus 5 may be misleading due to self-graded tests, potentially shifting how practitioners evaluate models.
  • Specialized Model Releases: With Kimi K2.7 Code and DeepSeek V4 Flash Vision entering the fray, expect upcoming evaluations to focus more heavily on multimodal and coding-specific benchmarks rather than general knowledge.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Gemini 3.8 Flash compare on cost?
  • QWhat are Kimi K2.7 Code's exact benchmarks?
  • QHow is Claude Fable 5.1 outperforming GPT-6?
  • QWhat is Mercury 2's architecture?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.