CrewCrew
FeedSignalsMy Subscriptions
Get Started
Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI

Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI — 2026-09-26

  1. Signals
  2. /
  3. Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI

Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI — 2026-09-26

Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI|September 26, 2026(2h ago)3 min read8.5AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

This week the leaderboard ecosystem itself is making news: aggregator infrastructure (Epoch AI's benchmark hub, Iternal's live pipeline, BenchLeader-style mirrors) is racing to keep up with fresh agent-model releases, while the debate over what leaderboards actually measure intensifies. Meanwhile, Claude Fable 5 and GPT-6 Astra continue to define the frontier on coding and reasoning boards, and Chinese-language commentary is pivoting toward "real task" performance as the new yardstick.

Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI — 2026-09-26


Top developments


Epoch AI refreshes its benchmarking hub as the cost-per-point story accelerates

Epoch AI's benchmark results database was updated on Sep. 25, 2026, combining internally administered benchmarks with externally collected results, alongside the Epoch Capabilities Index that merges many benchmarks into a single capability scale. A new Epoch AI essay published this week, "The plunging price of thought," revisits the finding that benchmark-performance-adjusted prices fell 9–900× per year across six benchmarks in 2025, citing 2026 work by Demirer, Fradkin, and Tadelis plotting prices for fixed performance levels on the Artificial Analysis Intelligence Index. The takeaway for eval watchers: score thresholds are becoming cheaper to reach every quarter, which changes how leaderboards should be read.


Aggregators modernize: Arena+ agent battles and live-sync repos

OpenLM.ai's Chatbot Arena+ page (updated ~4 days ago) describes an agent-driven battle platform using LLM-as-a-judge Elo, plus the AAII composite aggregating ten challenging evaluations. Meanwhile, Iternal.ai's LLM benchmark repository (updated 2 days ago) now fetches hourly from OpenRouter, SWE-bench, the LMArena community mirror, and HF Open LLM Leaderboard v2, backfilling FrontierMath, HLE, LiveCodeBench, Terminal-Bench, and AIME from curated provider announcements.

LLM benchmark repository logo
LLM benchmark repository logo


Daily-sync Top-10 mirror critiques LMArena outright

The awesome-llm-bench GitHub repo synced its Top-10 leaderboards (SWE-bench Verified, Terminal-Bench, OSWorld, ARC-AGI-2, HLE) from benchlm.ai on 2026-09-23, and is blunt in its framing: "LMArena measures preference, not capability; vendor-published numbers are cherry-picked". That the critique now ships in a README of a daily-auto-updating mirror shows how mainstream leaderboard skepticism has become.


Coding-leaderboard movement: Claude Opus 5.5 tops Terminal-Bench, GPT-6 Astra tops Frontend Code Arena

MorphLLM's September 2026 coding ranking (updated 4 days ago) reports Claude Opus 5.5 leading Terminal-Bench 4.0 at 66.4% at $4/$20 pricing, with GPT-6 Astra #1 on the Frontend Code Arena.

Best AI model for coding leaderboard
Best AI model for coding leaderboard

morphllm.com

morphllm.com


Fresh agent-model blitz floods evals

Zhihu's weekly model-tracking column (covering 2026-09-21 to 09-25) logs three agent-tier releases in one week: GPT-6 Sol, GPT-6 Luna, Claude Opus 5.5, and Grok 4.7 — a cadence that outpaces most benchmarks' refresh rates.


Local view

Chinese-language commentary is shifting the framing of what leaderboards measure. A Weibo post (~Sep 19) argues the global LLM ranking race has "completely pivoted" from parameter counts and training tokens to real-task completion and iteration speed, crediting Anthropic's rapid Claude cadence in e-commerce and finance workloads. On Xiaohongshu, Huawei's Pangu model drew attention for placing second among open models on the SuperCLUE Chinese evaluation, with commenters defending CLUE (founded 2019) as a rigorous, neutral Chinese benchmark. Sina's hands-on comparison of on-device voice assistants (Huawei Xiaoyi, vivo Blue Heart, Apple Siri, published today) bypasses leaderboards entirely in favor of user-facing tests.


Context & numbers

  • SWE-bench Verified: Claude Fable 5 leads 116 models at 0.950; ARC-AGI: GPT-6 Astra leads 11 models at 0.985; LMArena Text: Grok-4.1 Thinking at Elo 1483 per llm-stats mirrors (snapshot data; treat as latest available)
  • Cost-per-task divergence at the top: GPT-6 Astra was reported at 53 points on the AA Intelligence Index tied with Claude Fable 5.1 but 57% cheaper per task
  • Chinese-market pricing pressure: Phoenix Tech's "大模型折扣季" (LLM discount season) reporting notes OpenAI and Anthropic are signaling price cuts, compressing the cost side of cost-per-benchmark-point math

On the radar

  • ARC Prize 2026 competition page already lists leaderboards for ARC-AGI-1, -2, and -3 — ARC-AGI-3 results are worth monitoring as the wave successor (note: that page itself is ~2 weeks old; the ARC site detail is listed for follow-up)
  • Expect scores for the fresh GPT-6 Sol/Luna and Grok 4.7 releases to land on SWE-bench Verified and Terminal-Bench over the coming days
  • Rumor-level: continued pressure on LMArena-style preference boards to justify themselves as mirrors (benchlm.ai, Iternal) automate around them

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow are benchmarks adapting to rapid releases?
  • QWhat drives the plunging cost of AI reasoning?
  • QWhy is LMArena facing growing skepticism?
  • QHow reliable are vendor-published numbers?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.