CrewCrew
FeedSignalsMy Subscriptions
Get Started
AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-09-03

  1. Signals
  2. /
  3. AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-09-03

AI Benchmarks & Leaderboard|September 3, 2026(2h ago)4 min read8.7AI quality score — automatically evaluated based on accuracy, depth, and source quality
43 subscribers

Meta released Muse Spark 1.3, claiming it surpasses GPT-5.6 Sol in coding capabilities, marking a significant shift in the open-source vs. closed-source model landscape. Meanwhile, Artificial Analysis updated its leaderboard, placing Claude Fable 5.1 at the top of reasoning models with an Intelligence Index of 66, while Kimi K3 leads the open-weights category.

AI Benchmarks & Leaderboard — 2026-09-03


New Model Releases & Updates


Muse Spark 1.3 by Meta

  • Type: Closed-source (implied by "released" context, though Meta often releases open weights; specific parameter count not disclosed in snippet)
  • Key benchmarks: Chief AI Officer Alexandr Wang claimed it surpasses OpenAI's GPT-5.6 Sol in coding capabilities.
  • vs. Previous best: Directly challenges GPT-5.6 Sol, currently a top-tier competitor in coding tasks.
  • What's notable: This release represents Meta's most powerful AI model to date, intensifying competition in the coding and agent task sectors.

Meta's new AI model release
Meta's new AI model release


Leaderboard Snapshot


Frontier Models (Closed-Source)

ModelProviderNotable StrengthsKey Score
Claude Fable 5.1AnthropicAdaptive Reasoning, Max EffortIntelligence Index: 66
Claude Opus 4.8AnthropicAdaptive Reasoning, Max EffortIntelligence Index: 61
GPT-5.5 (xhigh)OpenAIHigh Reasoning EffortIntelligence Index: 60
GPT-5.5 (high)OpenAIHigh Reasoning EffortIntelligence Index: 59
Claude Opus 4.7AnthropicAdaptive Reasoning, Max EffortIntelligence Index: 57
Gemini 3.1 Pro PreviewGooglePro-level ReasoningIntelligence Index: 57

Open-Source Leaders

ModelParametersNotable StrengthsKey Score
Kimi K3 (max)N/AHighest ranked open weightsIntelligence Index: 60
GLM-5.3 (max)N/ATop open weights contenderIntelligence Index: 60
Qwen3.8 2.4T A95B2.4TLarge scale open weightsIntelligence Index: 58
Granite 4.2 3B3BLowest cost per Intelligence Index task ($0.0047)Cost Efficiency Leader
GPT-5.6 Luna (low)N/ALow cost optionCost per task: $0.01

Artificial Analysis Leaderboard
Artificial Analysis Leaderboard


Benchmark Deep Dive

The latest updates from Artificial Analysis highlight a tightening race between frontier closed-source models and high-performing open-weights models. Claude Fable 5.1 currently leads among 150+ reasoning models with an Intelligence Index score of 66. This score aggregates performance across 10 challenging evaluations, providing a robust measure of general intelligence rather than just specific task proficiency.

A notable trend is the cost-performance ratio. While Claude Fable 5.1 leads in raw intelligence, Granite 4.2 3B has emerged as the most cost-efficient model, with the lowest cost per Intelligence Index task at $0.0047. This suggests that for practitioners where budget is a primary constraint and absolute state-of-the-art reasoning is not required, smaller, optimized models are becoming increasingly viable.

The gap between the top open-weights model, Kimi K3 (max), and the top closed-source model, Claude Fable 5.1, is now just 6 points on the Intelligence Index (60 vs. 66). This narrow margin indicates that open-source models are reaching parity with closed-source counterparts for many enterprise use cases, particularly when combined with specialized fine-tuning or retrieval-augmented generation techniques.


Analysis & Trends

  • State of the art: Claude Fable 5.1 leads in general reasoning and adaptive effort scenarios. For coding, Meta's newly released Muse Spark 1.3 claims to outperform GPT-5.6 Sol, suggesting a potential shift in leadership for developer-focused tasks.
  • Open vs. Closed gap: The gap is narrowing significantly. Kimi K3 and GLM-5.3 are within striking distance of the top closed-source models on the Intelligence Index, offering competitive performance with greater flexibility for deployment.
  • Cost-performance: Granite 4.2 3B and GPT-5.6 Luna (low) are leading in cost efficiency, making them attractive for high-volume, lower-complexity tasks. The cost per Intelligence Index task is a key metric for practitioners balancing budget against capability.
  • Emerging patterns: There is a clear trend towards "adaptive reasoning" models like Claude Fable 5.1 and Opus 4.8, which allow users to balance latency and cost by adjusting reasoning effort levels.

What to Watch Next

  • Independent Verification of Muse Spark 1.3: Third-party benchmarks will be crucial to validate Meta's claim that Muse Spark 1.3 surpasses GPT-5.6 Sol in coding tasks.
  • Open-Weights Momentum: Continued improvements from Kimi K3 and GLM-5.3 could see them overtake current closed-source leaders in specific niche benchmarks or cost-adjusted intelligence metrics.
  • Gemini 3.7 Flash Performance: Recent changelogs indicate Gemini 3.7 Flash improved 4 points over its predecessor, potentially reshuffling the mid-tier leaderboard if it becomes widely available.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Muse Spark 1.3 perform on non-coding tasks?
  • QWhat testing criteria make up the Intelligence Index?
  • QHow much does running Claude Fable 5.1 cost?
  • QWhat hardware is required for Qwen3.8?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.