AI Benchmarks & Leaderboard — 2026-09-29
OpenAI's DevDay 2026 launched the Managed Agents platform on September 29, while BenchLM expanded its benchmark suite to 507 models. September saw five major frontier releases in ten days (Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, Muse Spark 1.3, DeepSeek V4.1-Flash), followed by a late-September wave including Claude Opus 5.5 at 40% lower cost and GPT-6 Sol and Luna from OpenAI.
AI Benchmarks & Leaderboard — 2026-09-29
New Model Releases & Updates
Claude Opus 5.5 by Anthropic
- Type: Closed-source, frontier-scale language model
- Key benchmarks: Achieves highest intelligence scores on Artificial Analysis leaderboard (alongside Claude Opus 5.5 with fallback options)
- vs. Previous best: 40% lower cost than previous Claude Opus iteration while maintaining frontier capabilities
- What's notable: Expanded Claude Marketplace integration; represents significant cost-performance improvement for enterprise users
GPT-6 Sol and Luna by OpenAI
- Type: Closed-source, dual-track frontier models
- Key benchmarks: GPT-6 Luna (low) achieves lowest cost per task according to Artificial Analysis leaderboard
- vs. Previous best: Luna reaches new price floor at $0.10 per million tokens, creating 119x price spread across frontier models
- What's notable: Launched as part of OpenAI DevDay 2026 on September 29; Sol and Luna represent different optimization targets (efficiency vs. reasoning)
Gemini 3.8 Flash by Google
- Type: Closed-source, fast inference model
- Key benchmarks: Achieves low latency with Gemini 2.5 Flash-Lite variants now on leaderboard
- vs. Previous best: Improves on Gemini 3.6 Flash; maintains speed advantage with flash-optimized inference
- What's notable: Part of September's rapid release cadence; focuses on sub-100ms latency performance
Leaderboard Snapshot
Frontier Models (Closed-Source) — By Intelligence
| Model | Provider | Notable Strengths | Primary Use |
|---|---|---|---|
| Claude Opus 5.5 (max with fallback) | Anthropic | Highest reasoning, 40% lower cost than prior | Complex reasoning, code generation |
| Claude Opus 5.5 (xhigh with fallback) | Anthropic | Top intelligence scores, adaptive effort levels | Enterprise applications |
| GPT-6 Luna (low) | OpenAI | Lowest cost per task, $0.10/M token pricing | Cost-sensitive deployments |
| Claude Sonnet 5.5 (max with fallback) | Anthropic | Strong general performance, balanced speed/cost | Balanced production use |
| Gemini 3.8 Flash | Sub-100ms latency, low cost per query | Real-time inference applications |
Speed & Latency Leaders
| Model | Provider | Tokens/Second | Latency Profile |
|---|---|---|---|
| Celeris-1 | Multiple | 1,491.1 t/s | Fastest frontier model |
| Mercury 2 | Multiple | 812.0-895.7 t/s | Top-tier speed |
| Mercury 2.5 | Multiple | 762.5 t/s | Consistent fast inference |
| Gemini 2.5 Flash-Lite (Non-reasoning) | High throughput | Lowest latency category | |
| Step 3.7 Flash | Multiple | 407.0 t/s | Fast general purpose |
Benchmark Deep Dive
BenchLM Reaches 507-Model Scale; Cost-Performance Becomes Central Metric
BenchLM's expansion to evaluate 507 models reflects a fundamental shift in how the field measures AI progress. Rather than focusing solely on capability benchmarks like MMLU-Pro or GPQA, September 2026 benchmarking now emphasizes the cost-performance trade-off—a practical concern for deployed systems.
The expanded benchmark pipeline now tracks API runtime metrics and evaluation tooling across closed and open-source models. This methodological shift reveals three critical patterns: (1) frontier models cluster tightly on raw intelligence but diverge sharply on cost (a 119x price spread across the frontier), (2) open-source models have closed the capability gap faster than cost gap, and (3) latency and throughput matter as much as per-token pricing for production workloads.
Drivenbench, a finance-sector focused benchmark released by Driven AI, evaluated 11 frontier models across capability, cost, and latency to match models to domain-specific workflows. This pattern—domain-specific evaluation—is becoming standard in September 2026. Generic leaderboards remain useful for orientation, but practitioners now require task-aligned benchmarks to make deployment decisions.
Analysis & Trends
-
State of the art: Claude Opus 5.5 and GPT-6 Sol lead on raw reasoning capability (per Artificial Analysis Intelligence Index); GPT-6 Luna dominates cost-efficiency category. Gemini 3.8 Flash and Mercury 2 lead on latency/throughput. Open-source models (Qwen, DeepSeek, Llama variants) compete in the sub-50B parameter range but remain behind frontier models on complex reasoning.
-
Open vs. Closed gap: The frontier has widened slightly on reasoning benchmarks (GPQA, MATH) but narrowed dramatically on cost. Open models like Qwen 3.5 and DeepSeek V4 now occupy the dominant cost-per-task position for inference-heavy workloads. September's releases show closed-source models optimizing for enterprise reasoning tasks, while open-source gains traction in volume inference (chatbots, retrieval-augmented generation).
-
Cost-performance: The $0.10 per million token price floor (GPT-6 Luna) represents a 12x reduction from Claude Opus 4.7 pricing. This shift has made cost-aware model selection a routine practice. Artificial Analysis reports a 119x price spread across frontier models, with Luna at bottom and Opus 5.5 in high-effort reasoning modes at top.
-
Emerging patterns: (1) Domain-specific benchmarking (finance via Drivenbench, code via SWE-bench) is replacing generic leaderboards for production decisions. (2) Multi-tier model architectures (reasoning vs. speed variants like GPT-6 Sol vs. Luna) allow cost optimization by task. (3) Runtime metrics (latency, throughput) now weigh equally with benchmark scores. (4) September's release pace (5 frontier launches in 10 days) indicates competitive saturation in frontier capability—differentiation now happens on cost and deployment efficiency.
What to Watch Next
-
OpenAI Managed Agents deployment metrics: The Managed Agents platform launched September 29 with 20+ new products; measure adoption rates and cost-per-task improvements for agentic workflows over next 4 weeks.
-
Claude Marketplace traction: Anthropic expanded the Claude Marketplace alongside Opus 5.5's 40% price cut; track whether cost reduction drives adoption of custom Claude implementations vs. generic GPT deployments.
-
Open-source reasoning models: DeepSeek V4.1-Flash and Qwen 3.8 variants are narrowing the reasoning gap on GPQA/MATH; measure if any reach >85% GPQA accuracy by end of Q4 2026 to signal true commodity access to frontier reasoning.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.