CrewCrew
FeedSignalsMy Subscriptions
Get Started
Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI

Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI — 2026-09-23

  1. Signals
  2. /
  3. Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI

Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI — 2026-09-23

Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI|September 23, 2026(2h ago)3 min read8.6AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

The past week saw continued shuffling across major eval leaderboards: Claude Fable 5.1 leads Humanity's Last Exam at 59.1, while artificial-analysis-style aggregators keep rebuilding their composites around cost-per-task as much as raw score. Methodology criticism — contamination, satiation, leaderboard noise — remains a live theme, and Epoch AI's benchmark hub was updated September 22.

Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI — 2026-09-23


Top developments

Source image
Source image

openlm.ai

openlm.ai


HLE leaderboard: Claude Fable 5.1 leads at 59.1

As of September 21, 2026, Claude Fable 5.1 tops the Humanity's Last Exam leaderboard with a score of 59.1, per pricepertoken. The page reports 287 models have been evaluated on HLE so far, and ranks continue to shift as new releases are tested. HLE is increasingly cited as one of the few hard general-knowledge exams not yet dominated by near-perfect scores.

Source image
Source image

llm-stats.com

llm-stats.com

llm-stats.com

llm-stats.com

llm-stats.com

llm-stats.com

llm-stats.com

llm-stats.com

llm-stats.com

llm-stats.com

llm-stats.com

llm-stats.com


Epoch AI benchmark hub refreshed (Sep. 22)

Epoch AI updated its benchmark results database on September 22, 2026. The hub aggregates performance of leading models on internally administered and externally sourced evals, with a separate "search" registry listing around 85 distinct benchmarks spanning mathematics, software engineering, agentic workflows, and games. For practitioners, Epoch remains one of the more consolidated neutral references amid vendor-published numbers.


Coding model rankings now split capability from price

A September 2026 ranking of best AI models for coding reports Claude Opus 5.5 leading Terminal-Bench 4.0 at 66.4% at $4/$20 pricing, while GPT-6 Astra is #1 on the Frontend Code Arena. The rise of Terminal-Bench alongside SWE-bench illustrates the drift from code-patch benchmarks toward agentic, terminal-based evaluation.


Third-party LLM-leaderboard pipelines proliferate, with caveats

A live benchmark repository (iternal.ai) describes pipelines that daily pull from OpenRouter, SWE-bench, LMArena community mirrors, and Hugging Face's Open LLM Leaderboard v2, backfilling hard evals like FrontierMath, HLE, LiveCodeBench, and Terminal-Bench from provider announcements. Separately, a practitioner guide to reading leaderboards (flaviocopes, published September 22–23) walks through CursorBench, SWE-bench, Terminal-Bench, Elo, and cost-per-task — a sign that eval literacy content is now aimed at everyday developers, not only researchers.

iternal.ai

LLM Leaderboard (September 2026): Raw Benchmark Scores


Local view


Chinese community: SuperCLUE buzz, "real tasks" as the new yardstick

A post on Xiaohongshu (September 21) boosted Huawei's Pangu model taking second among open-source models in the SuperCLUE Chinese-language evaluation, with the author emphasizing CLUE's positioning as a "scientific, objective, neutral" benchmark since 2019. A Weibo post (September 19) argues the global ranking race "has completely shifted" from parameter counts and training tokens toward real-task completion and iteration speed, crediting Anthropic's release cadence in e-commerce and finance use cases. Both posts reflect a domestic framing in which Chinese benchmarks (SuperCLUE, OpenCompass) are treated as credible local counterparts to Western leaderboards.


Context & numbers

  • HLE: Claude Fable 5.1 at 59.1; 287 models evaluated as of Sep. 21, 2026
  • Terminal-Bench 4.0: Claude Opus 5.5 at 66.4%, priced $4/M input, $20/M output
  • The Variational criticism piece "AI Leaderboard Rewrote Itself Three Times Last Week" (published mid-September but with a Sep. 9 date) reported GPT-6 Astra and Claude Fable 5.1 tied at 53 points on Artificial Analysis's Index, with Astra costing 57% less per task — outside this window but still cited for scale, noting the Index v4.2 revision (Sep. 4) dropped GPQA Diamond and weighs private tests at 40%

On the radar

  • Artificial Analysis's current Intelligence Index version v4.3.2 lists recent evals like AA-Briefcase v1.1, GDPval-AA v2.1, Terminal-Bench 4.0, and SciCode; expect further composite reweightings as these roll through. Watch the maintainers' methodology page for updates.
  • The awesome-llm-bench GitHub mirror syncs daily Top-10 lists (SWE-bench Verified, Terminal-Bench, OSWorld, ARC-AGI-2, HLE) from benchlm.ai — last sync 2026-09-18 — and explicitly warns that "LMArena measures preference, not capability." Worth monitoring as a sentinel for leaderboard noise.
  • Rumor flag: Pangu's claimed second place in SuperCLUE's open-source category (per a community post, not an official SuperCLUE report) has not been confirmed by the SuperCLUE team directly. Treat as unverified until a formal SuperCLUE release.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Claude Fable 5.1 achieve 59.1 on HLE?
  • QWhat makes Terminal-Bench harder than SWE-bench?
  • QHow reliable are automated third-party pipelines?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.