AI Benchmarks & Leaderboard — 2026-09-27
This week's biggest benchmark news is the launch of DrivenBench, a new evaluation suite testing 11 frontier models on real-world investment workflows across capability, cost, and latency. Meanwhile, independent ranking site Artificial Analysis shows Claude Opus 5.5 holding #1 on its LLM Leaderboard with an Intelligence Index of 58, and fresh analysis of Gemini 3.8 Flash reveals stark gaps between vendor-reported benchmark scores and standardized harness results.
AI Benchmarks & Leaderboard — 2026-09-27
New Model Releases & Updates
DrivenBench by Driven
- Type: Benchmark suite (not a model) — evaluates 11 frontier AI models
- Key benchmarks: Capability, cost, and latency metrics on finance-specific investment workflows
- vs. Previous best: Purpose-built for real-world investment tasks rather than general knowledge tests
- What's notable: Announced from Singapore on September 25, 2026, DrivenBench is explicitly designed to help investors match models to finance-specific workflows rather than relying on headline scores
Gemini 3.8 Flash — independent score divergence
- Type: Frontier model (closed-source, from Google)
- Key benchmarks: Scored 10.4% on ARC-AGI-3 (AGI3) under a standardized execution harness — a dramatically lower figure than heavily-tuned vendor-promoted benchmark numbers
- vs. Previous best: Illustrates the widening gap between marketing scores and verified task performance
- What's notable: The 404K Semi-AI weekly (September 25, 2026) warns enterprise buyers to evaluate performance against proprietary held-out task suites, not vendor-tuned results
Coding model rankings refresh (September 2026)
- Type: Frontier (closed) and open models ranked by coding benchmarks
- Key benchmarks: Claude Opus 5.5 leads Terminal-Bench 4.0 at 66.4% for $4/$20 per million tokens; GPT-6 Astra is #1 on the Frontend Code Arena
- vs. Previous best: Opus 5.5 retains the coding lead at a mid-range price point
- What's notable: The 13-model ranking connects benchmark scores directly to cost per task, updated within the past week

Leaderboard Snapshot
Frontier Models (Closed-Source)
Based on current listings from Artificial Analysis's LLM Leaderboard:
| Model | Provider | Notable Strengths | Key Score |
|---|---|---|---|
| Claude Opus 5.5 | Anthropic | Adaptive Reasoning, Max Effort, Default Fallback | Intelligence Index 58 (#1 of 172) |
| GPT-5.5 (xhigh) | OpenAI | Strong reasoning effort setting | Intelligence Index 60* |
| GPT-5.5 (high) | OpenAI | Balanced speed/quality | Intelligence Index 59* |
| Claude Opus 4.7 (Adaptive Reasoning, Max Effort) | Anthropic | Retains near-frontier performance | Intelligence Index 57* |
| Gemini 3.1 Pro Preview | Competitive general intelligence | Intelligence Index 57* |
* Note: these top-five figures come from a filtered leaderboard view (open weights excluded, small/medium sizes, reasoning filter); the primary unfiltered leaderboard currently places Claude Opus 5.5 at #1 with a score of 58. Scores shown may reflect the filtered subset ranked 1–5 on that view — verify directly on the page.
Open-Source Leaders
No fresh post-2026-09-25 data available with specific verified benchmark numbers for open-source leaderboards this cycle. Recent September coverage lists DeepSeek V4.1-Flash and Qwen3.8-Max among recent major open-family launches, and Artificial Analysis changelog entries note recent additions including GLM-5.2, DeepSeek V4 Pro 0424, MiniMax-M3, Kimi K3 (max), Nemotron 3.5 Lightning, and Qwen3.8 2.4T A95B — but specific scores in the source are not verifiable for this period.
Benchmark Deep Dive
The most striking evaluation finding this week comes from the 404K Semi-AI research weekly, which examined Gemini 3.8 Flash's performance on ARC-AGI-3 (AGI3) under two very different conditions. Under a standardized execution harness, the model scored just 10.4% — while vendor-promoted, heavily tuned benchmark scores told a far rosier story.
This is a textbook illustration of why enterprise buyers are increasingly distrustful of headline benchmark numbers. ARC-AGI-style tests, which measure general problem-solving ability on novel tasks, are among the benchmarks where exploit gaps between "tuned" and "verified" performance show up most starkly. A 10.4% harness score versus promoted figures signals that published marketing claims may not reflect what the model actually achieves when its outputs are executed in a consistent, neutral environment.
For practitioners, the lesson is concrete: hold out your own task suites and run models through standardized execution before trusting leaderboard claims. This concerns align with DrivenBench's launch rationale — the new finance-focused suite exists precisely because generic benchmarks don't predict domain workflow performance well, and the Driven team evaluates 11 frontier models across capability, cost, and latency for investment-specific tasks.

Analysis & Trends
- State of the art: Claude Opus 5.5 currently tops the Artificial Analysis overall leaderboard (Intelligence Index 58, #1 of 172 ranked models); GPT-6 Astra leads the Frontend Code Arena, while Opus 5.5 leads Terminal-Bench 4.0 at 66.4%
- Open vs. Closed gap: No verified fresh post-cutoff data quantifying the current open/closed gap specifically; historically, frontier drift continues with the latest September releases concentrated on the closed side.
- Cost-performance: Celeris-1 is the fastest model tracked at 1,545.5 tokens per second (followed by Mercury 2 at 812.0 t/s and Mercury 2.5 at 760.1 t/s), while Llama 3.1 Instruct 8B and Granite 4.2 3B are the most affordable at $0.02 per 1M tokens (blended)
- Emerging patterns: The frontier stack is increasingly evaluated on capability + cost + latency together rather than raw accuracy alone — as both DrivenBench and Artificial Analysis's blended price metric illustrate
What to Watch Next
- Domain-specific benchmark proliferation: Whether DrivenBench-style vertical benchmarks (finance, law, medicine) gain traction as the primary way enterprises evaluate frontier models.
- Harness-standardized ARC-AGI-3 results: More third-party harness runs like the 404K analysis could force vendors to publish standardized-execution scores, tightening the gap between claims and reality.
- Next-generation model cycle: WinCentral lists GPT-6 (full), Gemini 4, and DeepSeek V5 among models to watch, with Kimi K3.1 and Qwen 4 also in the next-generation pipeline.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.