Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI — 2026-10-10
Arena Intelligence, the company behind LMArena, has raised $200 million to nearly double its valuation to $3.1 billion, signaling a shift toward measuring AI alignment and truthfulness alongside raw capability. Meanwhile, Epoch AI’s latest data reveals that the cost of achieving high-level reasoning benchmarks has plummeted by 750-fold in 18 months, fundamentally altering the economics of model evaluation.
Evals and Leaderboards: LMArena, SWE-bench, ARC-AGI — 2026-10-10
Top developments
Arena Intelligence raises $200M, expands into alignment metrics
On October 8, 2026, TechCrunch reported that Arena Intelligence, the entity behind the popular LMArena leaderboard, secured a $200 million funding round led by Lightspeed and Khosla Ventures. This investment nearly doubles the company's valuation to $3.1 billion in just ten months. The capital will be used to expand the platform's evaluation suite beyond simple preference-based Elo ratings to include rigorous assessments of alignment issues such as lying and deception. This move underscores the market's growing demand for trustworthy AI metrics, moving beyond "which model is smarter" to "which model is safer and more honest."

Epoch AI reports 750-fold drop in cost per benchmark task
Epoch AI updated its benchmarking database on October 9, 2026, highlighting a dramatic decrease in the cost of achieving specific capability levels. According to their analysis, the price of equivalent capability on the GPQA Diamond benchmark has declined roughly 750-fold in eighteen months. For instance, OpenAI’s GPT-5.6 Luna now matches the 75% score previously achieved by GPT o3 at a cost of just $0.0004 per question, compared to $0.30 previously. This "plunging price of thought" suggests that high-level reasoning is becoming commoditized, potentially rendering some static leaderboards less relevant as cost-efficiency becomes the new frontier.

BenchLM updates Artificial Analysis Index with GPT-5.6 Sol leading
As of October 10, 2026, the independent aggregator BenchLM updated its Artificial Analysis Intelligence Index leaderboard. OpenAI’s GPT-5.6 Sol currently holds the top position with a score of 58.9%. The index aggregates results from 218 AI models across various challenging evaluations. This update provides a current snapshot of the frontier, contrasting with preference-based metrics by focusing on objective capability scores across math, science, and coding tasks.
Chinese media highlights October model cross-comparisons
Chinese tech outlet AI Puzi published a detailed cross-evaluation on October 8, 2026, comparing Claude Opus 5.5, Gemini 4 Argon, GPT-6 Astra, and DeepSeek V4.1 Flash. The article synthesizes data from LMSys Arena, Artificial Analysis, and LiveBench to provide selection advice for programming, mathematics, and agent tasks. This reflects the intense competitive landscape in China, where domestic models like DeepSeek are being rigorously benchmarked against Western frontier models using international standards.
Local view
AI Puzi, a prominent Chinese AI tutorial and news site, released a comprehensive "October 2026 Large Model Cross-Evaluation" on October 8. The report focuses on practical selection criteria for developers, analyzing the trade-offs between Claude Opus 5.5, Gemini 4 Argon, GPT-6 Astra, and DeepSeek V4.1 Flash. The outlet emphasizes that while GPT-6 Astra leads in general reasoning, DeepSeek V4.1 Flash offers superior cost-performance ratios for agent-based workflows, reflecting local stakeholders' sensitivity to API pricing and efficiency.
Context & numbers
- Valuation: Arena Intelligence (LMArena) valuation increased to $3.1 billion following a $200 million raise.
- Cost Efficiency: The cost to match GPQA Diamond performance levels dropped ~750x in 18 months; GPT-5.6 Luna achieves 75% score at $0.0004/question vs. $0.30 for previous leaders.
- Leaderboard Scores: GPT-5.6 Sol leads the Artificial Analysis Intelligence Index with 58.9% as of Oct 10, 2026.
- Model Updates: Recent weeks saw updates from Claude (Haiku 5.5), Google (Gemini 3.6 Flash Image), and Mistral (new Agent models), all tracked in daily leaderboard syncs.
On the radar
- Epoch Capabilities Index Update: Epoch AI continues to update its ECI scores daily; the latest CSV was updated on October 10, 2026, providing granular confidence intervals for model releases.
- Benchmark Saturation Debate: Discussions continue regarding "benchmark saturation," where top models approach ceiling scores on tests like MMLU, prompting a shift toward harder, dynamic benchmarks like FrontierMath and HLE.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.