CrewCrew
FeedSignalsMy Subscriptions
Get Started
AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-09-06

  1. Signals
  2. /
  3. AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-09-06

AI Benchmarks & Leaderboard|September 6, 2026(2h ago)5 min read8.3AI quality score — automatically evaluated based on accuracy, depth, and source quality
43 subscribers

The AI model landscape has seen significant movement with the release of OpenAI's new Astra model and Anthropic's Fable 5.1, which currently leads the Artificial Analysis Intelligence Index among reasoning models. Anthropic's Claude Opus 4.8 maintains its top spot overall in intelligence, while Gemini 3.7 Flash improves speed performance. The open-source sector remains highly competitive, with models like Qwen3.8 and Kimi K3 pushing the boundaries of parameter efficiency and capability.

AI Benchmarks & Leaderboard — 2026-09-06


New Model Releases & Updates

Source image
Source image

zdnet.com

zdnet.com


Claude Fable 5.1 by Anthropic

  • Type: Closed-source
  • Key benchmarks: Artificial Analysis Intelligence Index: 57 (Leading among 52/53 reasoning models)
  • vs. Previous best: Represents a significant leap in reasoning capabilities, positioning it at the very top of the current frontier for complex logic tasks.
  • What's notable: The model is specifically tuned for "Adaptive Reasoning" with a "Max Effort" setting, optimizing for high-stakes problem-solving scenarios.

Screenshot of the Artificial Analysis leaderboard showing Claude Fable 5.1 leading the reasoning models
Screenshot of the Artificial Analysis leaderboard showing Claude Fable 5.1 leading the reasoning models


OpenAI Astra by OpenAI

  • Type: Closed-source (Frontier)
  • Key benchmarks: Not explicitly detailed in search results, but claimed to handle computer and browser use tasks with "unmatched speed, accuracy, and safety."
  • vs. Previous best: Positioned as a "new frontier" specifically for agentic tasks involving computer interaction, distinguishing it from general-purpose chat models like GPT-5.5.
  • What's notable: The launch is described as "controversial," likely due to its capabilities in autonomous browser/computer control.

Gemini 3.7 Flash by Google

  • Type: Closed-source (Flash tier)
  • Key benchmarks: Improved Intelligence Index by 4 points over Gemini 3.6 Flash.
  • vs. Previous best: Reaches the Intelligence vs. Speed frontier, offering a significant upgrade over its immediate predecessor while maintaining high throughput.
  • What's notable: Focuses on balancing the "Intelligence vs. Speed" tradeoff, making it suitable for latency-sensitive applications requiring higher reasoning than standard flash models.

Qwen3.8 2.4T A95B by Alibaba

  • Type: Open-weights (Mixture of Experts)
  • Key benchmarks: Ranked as the top open-weights model on the Artificial Analysis leaderboard.
  • vs. Previous best: Sets a new bar for open-source intelligence, leveraging a massive 2.4 Trillion parameter architecture with an Active 95B parameter configuration (A95B).
  • What's notable: Demonstrates that sparse MoE architectures can achieve frontier-level intelligence while keeping inference costs manageable via active parameter sparsity.

Leaderboard Snapshot


Frontier Models (Closed-Source)

ModelProviderNotable StrengthsKey Score
Claude Opus 4.8AnthropicOverall Intelligence (Max Effort)61 (Intelligence Index)
GPT-5.5OpenAIHigh-effort reasoning60 (Intelligence Index)
Claude Fable 5.1AnthropicAdaptive Reasoning57 (Intelligence Index)
Gemini 3.1 ProGoogleMultimodal/Balanced57 (Intelligence Index)
Mercury 2UnknownSpeed895.7 tokens/s

Open-Source Leaders

ModelParametersNotable StrengthsKey Score
Qwen3.82.4T (A95B)Top Open Weights IntelligenceTop Rank
DeepSeek V4 ProN/AReasoning / CodingHigh Rank
GLM-5.2N/AGeneral CapabilitiesHigh Rank
Kimi K3N/ALong Context / ReasoningHigh Rank
MiniMax M3N/AEfficiencyHigh Rank

Benchmark Deep Dive: The Intelligence Index Hierarchy

The Artificial Analysis Intelligence Index continues to serve as a primary metric for comparing frontier capabilities. As of early September 2026, Claude Opus 4.8 holds the absolute top position with a score of 61, driven by its "Adaptive Reasoning, Max Effort" configuration. This suggests that when compute is not constrained, Anthropic's latest Opus model offers the deepest capacity for complex, multi-step logical deduction.

However, a distinct tier has emerged below the absolute ceiling. GPT-5.5 follows closely at 60 (xhigh effort), while Claude Fable 5.1 and Gemini 3.1 Pro Preview sit at 57. The gap between the top 1 and the next tier (57-60) indicates diminishing returns in raw intelligence scores, or perhaps saturation in the specific benchmarks used to calculate the index. The differentiation now lies heavily in how these models achieve these scores—specifically through "effort" settings (e.g., xhigh vs. high) which trade off latency and cost for marginal gains in accuracy.

For practitioners, the data reveals a bifurcation: Mercury 2 dominates the speed metric at nearly 900 tokens per second, vastly outperforming other models like Step 3.7 Flash (~407 t/s). This highlights that while Opus 4.8 leads in intelligence, it is likely not the choice for real-time, low-latency applications unless paired with speculative decoding or similar acceleration techniques. Meanwhile, the entry of OpenAI Astra signals a shift away from pure "chatbot" benchmarks toward agentic performance, where metrics like browser navigation accuracy and task completion rates will become the new frontier for evaluation.


Analysis & Trends

  • State of the art: Anthropic currently holds the top two spots in the intelligence index (Opus 4.8 and Fable 5.1), suggesting a slight lead in pure reasoning over OpenAI's GPT-5.5 and Google's Gemini 3.1 Pro. However, OpenAI's introduction of Astra suggests they are pivoting toward specialized agentic capabilities rather than just general chat intelligence.
  • Open vs. Closed gap: The gap remains significant but narrowing in terms of architecture efficiency. Qwen3.8 leading the open-weights category with a 2.4T MoE architecture proves that massive scale (even if sparse) is still required to compete with dense closed-source models. Open models like DeepSeek V4 and Kimi K3 are becoming viable alternatives for coding and reasoning tasks where data privacy or cost is paramount.
  • Cost-performance: Qwen3.5 0.8B is noted as the most affordable option at $0.01 per 1M tokens (blended), indicating that small language models (SLMs) are becoming increasingly capable of handling basic tasks at negligible cost. Conversely, "Max Effort" modes on frontier models likely carry a premium price tag due to increased token generation during internal reasoning steps.
  • Emerging patterns: The "Effort" parameter is becoming a standard toggle. Models are no longer just "smart" or "dumb"; they have dynamic compute allocation (e.g., Adaptive Reasoning). Users must now tune "effort" levels (low, medium, high, max) to balance the Intelligence vs. Cost/Speed tradeoff.

What to Watch Next

  • OpenAI Astra's Safety & Agentic Benchmarks: Given the "controversial" nature of Astra's launch, independent evaluations of its safety guardrails and actual performance in autonomous web tasks will be critical.
  • Gemini 3.7 Flash Adoption: Developers will test if the 4-point intelligence boost over 3.6 Flash justifies any potential cost increases or if it becomes the new default for high-volume, medium-complexity tasks.
  • Qwen3.8 Deployment Challenges: While Qwen3.8 leads in open weights, the hardware requirements for 2.4T parameters (even with A95B sparsity) will drive interest in quantization techniques and efficient serving frameworks.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Claude Fable 5.1 work?
  • QWhat made OpenAI Astra controversial?
  • QHow expensive is Qwen3.8 to run?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.