AI Benchmarks & Leaderboard — 2026-08-28
This week saw the release of Qwen3.8-Flash-Next, a 125B parameter model with a massive 262k context window. Meanwhile, Artificial Analysis's latest leaderboard confirms Claude Opus 5 (Max Effort) as the top reasoning model with an Intelligence Index of 63, while Gemini 3.1 Pro remains a strong contender in the open-weights category. The gap between frontier closed and open models continues to narrow, with Qwen3.8-Max and Kimi K3 showing significant benchmark improvements.
AI Benchmarks & Leaderboard — 2026-08-28
New Model Releases & Updates

Qwen3.8-Flash-Next by Alibaba Cloud
- Type: Open-source, 125B parameters
- Key benchmarks: Context window expanded to 262k tokens. Specific MMLU/GPQA scores pending full technical report release, but positioned as a high-efficiency alternative to larger models.
- vs. Previous best: Offers a significantly larger context window than many competitors at this parameter size, targeting long-document analysis tasks.
- What's notable: Released on August 26, 2026. Designed for efficiency and long-context reasoning without requiring the compute resources of 2T+ parameter models.
Claude Opus 5 by Anthropic
- Type: Closed-source
- Key benchmarks: Intelligence Index score of 63 (Adaptive Reasoning, Max Effort). Leads among 143 reasoning models evaluated by Artificial Analysis.
- vs. Previous best: Maintains its lead over GPT-5.6 and Gemini 3.1 Pro in complex reasoning tasks, though GPT-5.6 Luna is noted for cost-efficiency.
- What's notable: Continues to dominate the "Intelligence Index" leaderboard, particularly in tasks requiring multi-step adaptive reasoning.
Gemini 3.1 Pro by Google
- Type: Closed-source (Preview)
- Key benchmarks: Intelligence Index score of 57. Strong performance across general reasoning and multimodal tasks.
- vs. Previous best: Trails Claude Opus 5 in pure reasoning but offers competitive speed and integration capabilities.
- What's notable: Part of Google's broader Gemini 3 series, which includes Flash variants optimized for speed.
GPT-5.6 by OpenAI
- Type: Closed-source
- Key benchmarks: Intelligence Index scores vary by effort level; GPT-5.6 Luna is highlighted for low cost per task ($0.0047 per Intelligence Index task).
- vs. Previous best: Offers a balance between high intelligence (scores ~59-60) and cost-effectiveness, challenging Anthropic's dominance in value-per-intelligence metrics.
- What's notable: The "Luna" variant is specifically noted for its energy efficiency and low operational cost.
Kimi K3 by Moonshot AI
- Type: Open-weights, 2.8T total parameters (104B active)
- Key benchmarks: Terminal-Bench 2.1 score of 88.3. 1M context window.
- vs. Previous best: Currently ranks as one of the top open-source models for coding and agentic tasks.
- What's notable: Released with open weights on July 27, 2026, but continues to be a benchmark leader in August evaluations.
Leaderboard Snapshot
Frontier Models (Closed-Source)
| Model | Provider | Notable Strengths | Key Score |
|---|---|---|---|
| Claude Opus 5 | Anthropic | Complex Reasoning, Adaptive Effort | 63 (Intelligence Index) |
| GPT-5.6 | OpenAI | Balanced Performance, Cost Efficiency | 60 (Intelligence Index) |
| Gemini 3.1 Pro | Multimodal, Speed | 57 (Intelligence Index) | |
| Grok 4.6 | xAI | Real-time Data Integration | N/A (Top Tier) |
| GPT-5.6 Luna | OpenAI | Low Cost, Energy Efficiency | $0.0047/task |
Open-Source Leaders
| Model | Parameters | Notable Strengths | Key Score |
|---|---|---|---|
| Kimi K3 | 2.8T (104B active) | Coding, Agentic Tasks | 88.3 (Terminal-Bench) |
| Qwen3.8-Max | 2.4T | General Reasoning, Long Context | High (Specifics vary) |
| GLM-5.2 | N/A | Multilingual, Reasoning | High |
| DeepSeek V4-Pro | N/A | Cost-effective Reasoning | High |
| Nemotron 3.5 | N/A | NVIDIA-optimized efficiency | N/A |
Benchmark Deep Dive: The Rise of "Effort-Based" Intelligence
The most significant trend in recent benchmarking is the shift from static model comparisons to "effort-based" evaluation. Artificial Analysis's latest data highlights that models like Claude Opus 5 and GPT-5.6 do not have a single intelligence score; instead, their performance varies dramatically based on the computational "effort" allowed during inference. For instance, Claude Opus 5 achieves an Intelligence Index of 63 only when set to "Max Effort," whereas lower-effort settings yield scores comparable to smaller, faster models.
This development is crucial for practitioners because it decouples model capability from inference cost more granularly than ever before. Previously, choosing a "frontier" model meant accepting a fixed high cost. Now, users can dial in the exact level of reasoning depth required for a task. For simple classification, a low-effort setting on a frontier model might outperform a mid-tier model while costing less. For complex legal or medical reasoning, the "Max Effort" setting justifies the premium.
Furthermore, the emergence of specialized variants like GPT-5.6 Luna suggests a bifurcation in the market. We are seeing distinct models optimized for speed (Celeris-1 at 1,574 t/s), cost (Granite 4.2 3B at $0.0047/task), and peak intelligence (Claude Opus 5). This fragmentation means "best model" is no longer a valid question without specifying the constraint: budget, latency, or accuracy.
Analysis & Trends
- State of the art: Claude Opus 5 leads in pure reasoning complexity. GPT-5.6 Luna leads in cost-performance ratio. Celeris-1 leads in raw token throughput.
- Open vs. Closed gap: The gap is narrowing rapidly. Kimi K3 and Qwen3.8-Max are achieving scores within striking distance of closed frontier models on specific benchmarks like Terminal-Bench, often with larger context windows.
- Cost-performance: Granite 4.2 3B is noted for having the lowest cost per Intelligence Index task, indicating that small language models (SLMs) are becoming viable for many enterprise tasks previously reserved for LLMs.
- Emerging patterns: Long-context capabilities are becoming standard, with new releases like Qwen3.8-Flash-Next offering 262k tokens natively.
What to Watch Next
- Qwen3.8 Technical Report: Full benchmark details for the newly released Qwen3.8 series are expected soon, which may disrupt current open-source rankings.
- Artificial Analysis Grader Updates: The platform recently upgraded its graders for HLE and AA-LCR using GPT-5.6 Luna, which may cause slight shifts in reported scores for all models as grading becomes more accurate.
- NVIDIA Nemotron Ecosystem: Continued releases from NVIDIA's local AI community are focusing on efficient, open-weight models optimized for consumer hardware, potentially democratizing access to high-performance AI.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.