CrewCrew
FeedSignalsMy Subscriptions
Get Started
AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-08-28

  1. Signals
  2. /
  3. AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-08-28

AI Benchmarks & Leaderboard|August 28, 2026(2h ago)5 min read8.3AI quality score — automatically evaluated based on accuracy, depth, and source quality
43 subscribers

This week saw the release of Qwen3.8-Flash-Next, a 125B parameter model with a massive 262k context window. Meanwhile, Artificial Analysis's latest leaderboard confirms Claude Opus 5 (Max Effort) as the top reasoning model with an Intelligence Index of 63, while Gemini 3.1 Pro remains a strong contender in the open-weights category. The gap between frontier closed and open models continues to narrow, with Qwen3.8-Max and Kimi K3 showing significant benchmark improvements.

AI Benchmarks & Leaderboard — 2026-08-28


New Model Releases & Updates

Source image
Source image

local-ai-zone.github.io

local-ai-zone.github.io

local-ai-zone.github.io

local-ai-zone.github.io


Qwen3.8-Flash-Next by Alibaba Cloud

  • Type: Open-source, 125B parameters
  • Key benchmarks: Context window expanded to 262k tokens. Specific MMLU/GPQA scores pending full technical report release, but positioned as a high-efficiency alternative to larger models.
  • vs. Previous best: Offers a significantly larger context window than many competitors at this parameter size, targeting long-document analysis tasks.
  • What's notable: Released on August 26, 2026. Designed for efficiency and long-context reasoning without requiring the compute resources of 2T+ parameter models.

Source image
Source image

aireleasetracker.com

aireleasetracker.com

aireleasetracker.com

aireleasetracker.com


Claude Opus 5 by Anthropic

  • Type: Closed-source
  • Key benchmarks: Intelligence Index score of 63 (Adaptive Reasoning, Max Effort). Leads among 143 reasoning models evaluated by Artificial Analysis.
  • vs. Previous best: Maintains its lead over GPT-5.6 and Gemini 3.1 Pro in complex reasoning tasks, though GPT-5.6 Luna is noted for cost-efficiency.
  • What's notable: Continues to dominate the "Intelligence Index" leaderboard, particularly in tasks requiring multi-step adaptive reasoning.

Gemini 3.1 Pro by Google

  • Type: Closed-source (Preview)
  • Key benchmarks: Intelligence Index score of 57. Strong performance across general reasoning and multimodal tasks.
  • vs. Previous best: Trails Claude Opus 5 in pure reasoning but offers competitive speed and integration capabilities.
  • What's notable: Part of Google's broader Gemini 3 series, which includes Flash variants optimized for speed.

GPT-5.6 by OpenAI

  • Type: Closed-source
  • Key benchmarks: Intelligence Index scores vary by effort level; GPT-5.6 Luna is highlighted for low cost per task ($0.0047 per Intelligence Index task).
  • vs. Previous best: Offers a balance between high intelligence (scores ~59-60) and cost-effectiveness, challenging Anthropic's dominance in value-per-intelligence metrics.
  • What's notable: The "Luna" variant is specifically noted for its energy efficiency and low operational cost.

Kimi K3 by Moonshot AI

  • Type: Open-weights, 2.8T total parameters (104B active)
  • Key benchmarks: Terminal-Bench 2.1 score of 88.3. 1M context window.
  • vs. Previous best: Currently ranks as one of the top open-source models for coding and agentic tasks.
  • What's notable: Released with open weights on July 27, 2026, but continues to be a benchmark leader in August evaluations.

Leaderboard Snapshot


Frontier Models (Closed-Source)

ModelProviderNotable StrengthsKey Score
Claude Opus 5AnthropicComplex Reasoning, Adaptive Effort63 (Intelligence Index)
GPT-5.6OpenAIBalanced Performance, Cost Efficiency60 (Intelligence Index)
Gemini 3.1 ProGoogleMultimodal, Speed57 (Intelligence Index)
Grok 4.6xAIReal-time Data IntegrationN/A (Top Tier)
GPT-5.6 LunaOpenAILow Cost, Energy Efficiency$0.0047/task

Open-Source Leaders

ModelParametersNotable StrengthsKey Score
Kimi K32.8T (104B active)Coding, Agentic Tasks88.3 (Terminal-Bench)
Qwen3.8-Max2.4TGeneral Reasoning, Long ContextHigh (Specifics vary)
GLM-5.2N/AMultilingual, ReasoningHigh
DeepSeek V4-ProN/ACost-effective ReasoningHigh
Nemotron 3.5N/ANVIDIA-optimized efficiencyN/A

Benchmark Deep Dive: The Rise of "Effort-Based" Intelligence

The most significant trend in recent benchmarking is the shift from static model comparisons to "effort-based" evaluation. Artificial Analysis's latest data highlights that models like Claude Opus 5 and GPT-5.6 do not have a single intelligence score; instead, their performance varies dramatically based on the computational "effort" allowed during inference. For instance, Claude Opus 5 achieves an Intelligence Index of 63 only when set to "Max Effort," whereas lower-effort settings yield scores comparable to smaller, faster models.

This development is crucial for practitioners because it decouples model capability from inference cost more granularly than ever before. Previously, choosing a "frontier" model meant accepting a fixed high cost. Now, users can dial in the exact level of reasoning depth required for a task. For simple classification, a low-effort setting on a frontier model might outperform a mid-tier model while costing less. For complex legal or medical reasoning, the "Max Effort" setting justifies the premium.

Furthermore, the emergence of specialized variants like GPT-5.6 Luna suggests a bifurcation in the market. We are seeing distinct models optimized for speed (Celeris-1 at 1,574 t/s), cost (Granite 4.2 3B at $0.0047/task), and peak intelligence (Claude Opus 5). This fragmentation means "best model" is no longer a valid question without specifying the constraint: budget, latency, or accuracy.


Analysis & Trends

  • State of the art: Claude Opus 5 leads in pure reasoning complexity. GPT-5.6 Luna leads in cost-performance ratio. Celeris-1 leads in raw token throughput.
  • Open vs. Closed gap: The gap is narrowing rapidly. Kimi K3 and Qwen3.8-Max are achieving scores within striking distance of closed frontier models on specific benchmarks like Terminal-Bench, often with larger context windows.
  • Cost-performance: Granite 4.2 3B is noted for having the lowest cost per Intelligence Index task, indicating that small language models (SLMs) are becoming viable for many enterprise tasks previously reserved for LLMs.
  • Emerging patterns: Long-context capabilities are becoming standard, with new releases like Qwen3.8-Flash-Next offering 262k tokens natively.

What to Watch Next

  • Qwen3.8 Technical Report: Full benchmark details for the newly released Qwen3.8 series are expected soon, which may disrupt current open-source rankings.
  • Artificial Analysis Grader Updates: The platform recently upgraded its graders for HLE and AA-LCR using GPT-5.6 Luna, which may cause slight shifts in reported scores for all models as grading becomes more accurate.
  • NVIDIA Nemotron Ecosystem: Continued releases from NVIDIA's local AI community are focusing on efficient, open-weight models optimized for consumer hardware, potentially democratizing access to high-performance AI.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Claude Opus 5 achieve its top score?
  • QWhat is the running cost of Qwen3.8-Flash-Next?
  • QHow do open-source models compare to GPT-5.6?
  • QWhat tasks does Gemini 3.1 Pro handle best?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.