AI Benchmarks & Leaderboard — 2026-09-03
Meta released Muse Spark 1.3, claiming it surpasses GPT-5.6 Sol in coding capabilities, marking a significant shift in the open-source vs. closed-source model landscape. Meanwhile, Artificial Analysis updated its leaderboard, placing Claude Fable 5.1 at the top of reasoning models with an Intelligence Index of 66, while Kimi K3 leads the open-weights category.
AI Benchmarks & Leaderboard — 2026-09-03
New Model Releases & Updates
Muse Spark 1.3 by Meta
- Type: Closed-source (implied by "released" context, though Meta often releases open weights; specific parameter count not disclosed in snippet)
- Key benchmarks: Chief AI Officer Alexandr Wang claimed it surpasses OpenAI's GPT-5.6 Sol in coding capabilities.
- vs. Previous best: Directly challenges GPT-5.6 Sol, currently a top-tier competitor in coding tasks.
- What's notable: This release represents Meta's most powerful AI model to date, intensifying competition in the coding and agent task sectors.

Leaderboard Snapshot
Frontier Models (Closed-Source)
| Model | Provider | Notable Strengths | Key Score |
|---|---|---|---|
| Claude Fable 5.1 | Anthropic | Adaptive Reasoning, Max Effort | Intelligence Index: 66 |
| Claude Opus 4.8 | Anthropic | Adaptive Reasoning, Max Effort | Intelligence Index: 61 |
| GPT-5.5 (xhigh) | OpenAI | High Reasoning Effort | Intelligence Index: 60 |
| GPT-5.5 (high) | OpenAI | High Reasoning Effort | Intelligence Index: 59 |
| Claude Opus 4.7 | Anthropic | Adaptive Reasoning, Max Effort | Intelligence Index: 57 |
| Gemini 3.1 Pro Preview | Pro-level Reasoning | Intelligence Index: 57 |
Open-Source Leaders
| Model | Parameters | Notable Strengths | Key Score |
|---|---|---|---|
| Kimi K3 (max) | N/A | Highest ranked open weights | Intelligence Index: 60 |
| GLM-5.3 (max) | N/A | Top open weights contender | Intelligence Index: 60 |
| Qwen3.8 2.4T A95B | 2.4T | Large scale open weights | Intelligence Index: 58 |
| Granite 4.2 3B | 3B | Lowest cost per Intelligence Index task ($0.0047) | Cost Efficiency Leader |
| GPT-5.6 Luna (low) | N/A | Low cost option | Cost per task: $0.01 |
Benchmark Deep Dive
The latest updates from Artificial Analysis highlight a tightening race between frontier closed-source models and high-performing open-weights models. Claude Fable 5.1 currently leads among 150+ reasoning models with an Intelligence Index score of 66. This score aggregates performance across 10 challenging evaluations, providing a robust measure of general intelligence rather than just specific task proficiency.
A notable trend is the cost-performance ratio. While Claude Fable 5.1 leads in raw intelligence, Granite 4.2 3B has emerged as the most cost-efficient model, with the lowest cost per Intelligence Index task at $0.0047. This suggests that for practitioners where budget is a primary constraint and absolute state-of-the-art reasoning is not required, smaller, optimized models are becoming increasingly viable.
The gap between the top open-weights model, Kimi K3 (max), and the top closed-source model, Claude Fable 5.1, is now just 6 points on the Intelligence Index (60 vs. 66). This narrow margin indicates that open-source models are reaching parity with closed-source counterparts for many enterprise use cases, particularly when combined with specialized fine-tuning or retrieval-augmented generation techniques.
Analysis & Trends
- State of the art: Claude Fable 5.1 leads in general reasoning and adaptive effort scenarios. For coding, Meta's newly released Muse Spark 1.3 claims to outperform GPT-5.6 Sol, suggesting a potential shift in leadership for developer-focused tasks.
- Open vs. Closed gap: The gap is narrowing significantly. Kimi K3 and GLM-5.3 are within striking distance of the top closed-source models on the Intelligence Index, offering competitive performance with greater flexibility for deployment.
- Cost-performance: Granite 4.2 3B and GPT-5.6 Luna (low) are leading in cost efficiency, making them attractive for high-volume, lower-complexity tasks. The cost per Intelligence Index task is a key metric for practitioners balancing budget against capability.
- Emerging patterns: There is a clear trend towards "adaptive reasoning" models like Claude Fable 5.1 and Opus 4.8, which allow users to balance latency and cost by adjusting reasoning effort levels.
What to Watch Next
- Independent Verification of Muse Spark 1.3: Third-party benchmarks will be crucial to validate Meta's claim that Muse Spark 1.3 surpasses GPT-5.6 Sol in coding tasks.
- Open-Weights Momentum: Continued improvements from Kimi K3 and GLM-5.3 could see them overtake current closed-source leaders in specific niche benchmarks or cost-adjusted intelligence metrics.
- Gemini 3.7 Flash Performance: Recent changelogs indicate Gemini 3.7 Flash improved 4 points over its predecessor, potentially reshuffling the mid-tier leaderboard if it becomes widely available.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.