Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI — 2026-09-21
Anthropic’s Claude Fable 5.1 has solidified its lead across major benchmarks, topping Terminal-Bench and Humanity's Last Exam, while GPT-6 Astra dominates the ARC-AGI leaderboard. Meanwhile, the industry grapples with "benchmark saturation," as top models score in the mid-90s on legacy tests like SWE-bench Verified, prompting a shift toward agentic and real-task evaluations.
Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI — 2026-09-21
Top developments
Claude Fable 5.1 Leads Cross-Discipline and Agentic Benchmarks
In early September 2026, Anthropic released Claude Fable 5.1, which has taken the number one spot on key leaderboards including Terminal-Bench 4.0 (agentic coding) and Humanity's Last Exam (HLE). While earlier models like GPT-5.6 Sol matched Claude on GPQA Diamond, Fable 5.1's performance on HLE distinguishes it in cross-discipline reasoning. This release underscores a trend where "global best" titles are increasingly contested between Anthropic, OpenAI, and xAI based on specific task domains rather than aggregate scores.

GPT-6 Astra Tops ARC-AGI with Near-Perfect Score
GPT-6 Astra currently leads the ARC-AGI leaderboard with a score of 0.985, significantly ahead of other frontier models. ARC-AGI is designed to test abstract reasoning and novel problem-solving without relying on pre-trained pattern matching, making it one of the few benchmarks that remains unsaturated for top-tier models. This high score suggests that the latest generation of models is beginning to crack the "reasoning gap" that previously separated LLMs from human-like general intelligence on these specific tasks.
Benchmark Saturation Crisis: When 90% Scores Mean Nothing
Recent analyses highlight a growing crisis in AI evaluation: "benchmark saturation." With frontier models scoring in the mid-90s on tests like SWE-bench Verified and MMLU, these leaderboards are losing their discriminatory power. A systematic study by Lacuna notes that benchmarks are "breaking" as models approach the ceiling of the test, making small score differences statistically insignificant for real-world capability assessment. The industry is responding by shifting focus to dynamic benchmarks like Terminal-Bench and LiveCodeBench, which update frequently to prevent contamination and measure ongoing learning.

LMArena ELO Rankings Show Tight Top-Tier Competition
The LMArena (Chatbot Arena) leaderboard continues to be the primary metric for human-preference evaluation, with over 6.8 million blind votes recorded across 360+ models. Currently, the top tier of models is separated by only ~55 ELO points, indicating that user preference is no longer a clear differentiator between the leading closed-source models (like GPT-6, Claude Fable 5, and Grok-4.1). Critics argue that LMArena measures style and preference rather than raw capability, a sentiment echoed in recent methodological guides warning against relying solely on ELO ratings for technical selection.
Local view
Chinese tech media outlets are closely tracking the shift toward "real-task completion ability" as the new standard for model ranking. Weibo discussions highlight that Anthropic's rapid iteration cycle has allowed Claude to maintain a lead in practical applications like e-commerce and finance, despite intense competition from domestic models like DeepSeek and Kimi. The narrative in local media suggests that while global leaderboards remain important, domestic stakeholders are increasingly prioritizing benchmarks that reflect specific industrial use cases over general academic scores.
Context & numbers
- ARC-AGI Leaderboard: GPT-6 Astra leads with 0.985 accuracy among 11 evaluated models.
- SWE-bench Verified: Claude Fable 5 holds the top spot with a score of 0.950 across 116 models.
- Humanity's Last Exam (HLE): Claude Fable 5 leads Scale SEAL's HLE board at 55.5%, while GPT-5.6 Sol scores 44.4%.
- LMArena Votes: The platform has accumulated over 6.8 million blind human preference votes.
On the radar
- Epoch AI Updates: Epoch AI continues to update its Capabilities Index and benchmark database daily, providing independent verification of model performance outside vendor claims. Their data hub was last updated on September 20, 2026.
- Contamination Detection Tools: New methodologies like "Kernel Divergence Score" are being proposed to quantify dataset leakage more accurately, aiming to resolve disputes over whether high scores are due to capability or contamination.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.