CrewCrew
FeedSignalsMy Subscriptions
Get Started
AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-07-28

  1. Signals
  2. /
  3. AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-07-28

AI Benchmarks & Leaderboard|July 28, 2026(2h ago)5 min read9.3AI quality score — automatically evaluated based on accuracy, depth, and source quality
43 subscribers

July 20-28 delivered a landmark week: four frontier models launched in rapid succession (GPT-5.6 Sol, Claude Opus 5, Gemini 3.6 Flash, Qwen 3.8 Max), reshaping the competitive landscape. Chinese open-weight models including Kimi K3 are closing the gap with U.S. labs. Coding performance leaders now include GPT-5.6 Sol (96.2% on SWE-bench) and Claude Fable 5 (95.0%), with notable efficiency gains across the board.

AI Benchmarks & Leaderboard — 2026-07-28


New Model Releases & Updates

Source image
Source image

felloai.com

felloai.com


GPT-5.6 Sol by OpenAI

  • Type: Closed-source, frontier model
  • Key benchmarks: 96.2% SWE-bench Verified (independent evaluation); top-tier performance on reasoning and coding tasks
  • vs. Previous best: Surpasses GPT-5.5, solidifying OpenAI's lead in coding-focused evaluation
  • What's notable: Highest independent SWE-bench score achieved; aggressive pricing strategy undercuts competitors

Source image
Source image

digitalapplied.com

digitalapplied.com

digitalapplied.com

digitalapplied.com


Claude Opus 5 & Claude Fable 5 by Anthropic

  • Type: Closed-source, frontier models
  • Key benchmarks: Claude Opus 5 (Adaptive Reasoning, Max Effort) achieves 61 on Artificial Analysis Intelligence Index; Claude Fable 5 reaches 95.0% SWE-bench for coding
  • vs. Previous best: Opus 5 ranks among highest-intelligence models; Fable 5 competes directly with Sol
  • What's notable: Dual-release strategy targets both reasoning (Opus) and efficiency (Fable); Fable launched July 1

Kimi K3 by Moonshot AI (Open-Weight)

  • Type: Open-source, 2.8T parameters
  • Key benchmarks: Ranks #1 on Arena Elo (1,486); beats Fable 5 on select benchmarks; open-weights release imminent (July 27, 2026)
  • vs. Previous best: First Chinese open-weight to seriously challenge frontier models
  • What's notable: Demonstrates China's capability surge; open-sourcing will democratize access. Kimi models lead on Arena community evaluations.

Qwen 3.8 Max by Alibaba

  • Type: Closed-source, competitive with frontier
  • Key benchmarks: Strong performance on reasoning and general tasks; competitive with Gemini 3.6 Flash
  • vs. Previous best: Closes gap with OpenAI/Anthropic on multi-modal reasoning
  • What's notable: Alibaba's second aggressive push in July; Chinese labs now releasing models weekly

Leaderboard Snapshot


Frontier Models (Closed-Source)

ModelProviderNotable StrengthsKey Score
Claude Opus 5 (Max)AnthropicAdaptive reasoning, instruction following61 Intelligence Index
GPT-5.6 Sol (max)OpenAICoding, math, agentic tasks96.2% SWE-bench
Claude Fable 5AnthropicEfficient reasoning, coding95.0% SWE-bench
GPT-5.5 (xhigh)OpenAIGeneral reasoning, coding60 Intelligence Index
Gemini 3.6 FlashGoogleSpeed + reasoning balanceFaster than 3.5 Flash
Mercury 2UnknownExtreme speed953.1 tokens/sec
Qwen 3.8 MaxAlibabaMulti-modal, reasoningCompetitive with Gemini 3.6

Open-Source Leaders

ModelParametersNotable StrengthsKey Score
Kimi K32.8TReasoning, instruction-follow, Arena #11,486 Arena Elo
GLM-5.2UnknownGeneral reasoning, multi-modalTop open-source
Qwen 3.6-27B27BReasoning + coding on modest hardwareRuns on single RTX 4090
Gemma 4 31B31BCost-effective reasoningStrong Apache 2.0 alternative
DeepSeek V4-FlashUnknownFast, cost-efficient$0.14/M input tokens

Benchmark Deep Dive

SWE-bench Verified (Software Engineering): The most significant shift this week is in independent coding evaluation. GPT-5.6 Sol achieved 96.2% on SWE-bench Verified, the highest score to date on this rigorous benchmark designed to evaluate autonomous software engineering capability. Claude Fable 5 reached 95.0%, demonstrating that Anthropic's efficiency-focused model remains competitive with OpenAI's flagship architecture on task completion.

What this reveals: The frontier gap on coding is narrowing because both labs prioritized engineering capability during model training. SWE-bench remains the most reliable independent test because it measures end-to-end task completion (fixing real GitHub issues) rather than multiple-choice accuracy. The 96.2% vs 95.0% gap suggests Sol has an edge in complex multi-step reasoning, while Fable's 1-point gap is within error margins.

For practitioners: If your workload is autonomous code generation, GPT-5.6 Sol and Claude Fable 5 are now interchangeable—choose based on cost and API availability rather than capability. Open-source teams should note that neither Llama nor Mistral appear in the SWE-bench top 5, creating an opportunity for Qwen or DeepSeek to fill the gap.


Analysis & Trends

  • State of the art:

    • Reasoning: Claude Opus 5 leads (61 on Artificial Analysis index); GPT-5.5 xhigh (60)
    • Coding: GPT-5.6 Sol (96.2% SWE-bench) > Claude Fable 5 (95.0%)
    • Speed/Cost: Mercury 2 (953 tok/sec), Gemini 3.5 Flash-Lite (417 tok/sec); Qwen 0.8B ($0.01/M)
    • Arena (community voting): Kimi K3 (#1, 1,486 Elo)
  • Open vs. Closed gap: Kimi K3's Arena #1 ranking is a watershed moment. For the first time, an open-weight model beats closed-source on community preference (though not necessarily on standardized benchmarks). This signals that practitioners care about usability and alignment, not just raw capability. Gap to frontier still ~3-5 percentage points on MMLU-Pro, but closing.

  • Cost-performance: Gemini 3.5 Flash-Lite halves latency vs Flash (from ~30ms to 15ms) while improving intelligence by 11 points. Qwen 3.5 0.8B at $0.01/M tokens is now the cost floor. DeepSeek V4-Flash at $0.14/M offers 10x better value than GPT-4 baseline ($1.50/M).

  • Emerging patterns:

    • China surge: Kimi, Qwen, and DeepSeek released in parallel—no longer sequential. Suggests R&D acceleration and confidence in competitive parity.
    • Multimodal maturity: Gemini 3.6 Flash includes vision; Claude models added video understanding.
    • Efficiency obsession: 7 of the top 10 new models prioritize speed/cost over pure capability.
    • Open-weight momentum: Kimi K3 (2.8T parameters, open) will ship July 27—a full frontier-scale model, not a toy.

What to Watch Next

  • Kimi K3 open-weight release (July 27, 2026): If reproducible on public benchmarks, this validates China's R&D model and forces U.S. labs to compete on accessibility, not just performance. Watch for immediate fine-tuning experiments on LLaMA-style architectures.

  • GPT-5.6 Sol pricing & availability: OpenAI's aggressive cost positioning (implied in July release timing) may force Anthropic to reduce Claude pricing. Monitor API cost wars as the primary competitive lever for Q3 2026.

  • Gemini 3.6 Flash multimodal benchmarks: Google claims no intelligence regression vs 3.5, but hasn't released vision benchmark scores. When MMVP, SEED-Bench, and video understanding results land (expected early August), this will clarify whether multimodal capacity comes "free" or at a cost.

This report covers releases and benchmark updates from July 21–28, 2026. Previous weeks' models (Gemini 3.5 Flash, Laguna S 2.1, FLUX 3) are tracked in prior issues.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Kimi K3's open-weights release impact security?
  • QWhat makes Claude Opus 5's reasoning unique?
  • QHow will OpenAI's new pricing affect market competition?
  • QDoes Kimi K3 match frontier models in real-world usage?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.