CrewCrew
FeedSignalsMy Subscriptions
Get Started
AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-09-18

  1. Signals
  2. /
  3. AI Benchmarks & Leaderboard

AI Benchmarks & Leaderboard — 2026-09-18

AI Benchmarks & Leaderboard|September 18, 2026(2h ago)5 min read9.3AI quality score — automatically evaluated based on accuracy, depth, and source quality
43 subscribers

MLCommons has released MLPerf Inference v6.1, setting a new participation record and introducing benchmarks for Agentic Inference. Meanwhile, Artificial Analysis continues to track the rapid evolution of open-source models, with new entrants like MiniMax-M3 and Nemotron 3 Ultra joining the leaderboard alongside established leaders like Claude Fable 5.1 and GPT-6 Astra. The latest data highlights a narrowing gap in intelligence scores for smaller, faster models, particularly in reasoning tasks.

AI Benchmarks & Leaderboard — 2026-09-18


New Model Releases & Updates


MLPerf Inference v6.1 by MLCommons

  • Type: Benchmark Suite Update (Open Standard)
  • Key benchmarks: Introduces two new tests: Agentic Inference and Multimodal Generation.
  • vs. Previous best: Sets a new participation record with broader hardware vendor support compared to v6.0.
  • What's notable: The new "Agentic Inference" test evaluates models on their ability to perform complex, multi-step actions rather than just static text generation, reflecting the industry shift toward autonomous agents.

MLCommons MLPerf Inference v6.1 Announcement
MLCommons MLPerf Inference v6.1 Announcement

mlcommons.org

mlcommons.org


Union Alpha AI by Union Labs

  • Type: Closed-source (Limited Preview)
  • Key benchmarks: Verified specs currently under evaluation; available via OpenRouter and Cloudflare.
  • vs. Previous best: Positioned as a free anonymous model in limited preview, aiming to disrupt cost structures for entry-level inference.
  • What's notable: Offers free access through specific platforms, though full benchmark performance details remain pending wider release.

Recent Additions to Artificial Analysis Leaderboard

  • Models: MiniMax-M3, Kimi K2.7 Code, DeepSeek V4 Flash Vision, Nemotron 3 Ultra 550B A55B.
  • Key benchmarks: Gemini 3.8 Flash recently scored 59 on the Artificial Analysis Intelligence Index.
  • vs. Previous best: These models are competing directly with frontier closed-source models in specific niches like coding (Kimi K2.7 Code) and vision-reasoning (DeepSeek V4 Flash).

Leaderboard Snapshot


Frontier Models (Closed-Source)

ModelProviderNotable StrengthsKey Score
Claude Fable 5.1AnthropicHighest Intelligence Index53
GPT-6 AstraOpenAIStrong Reasoning (xhigh/max)N/A
Gemini 3.8 FlashGoogleSpeed/Intelligence Balance59
Claude Opus 4.8AnthropicAdaptive Reasoning (Max Effort)61
GPT-5.5OpenAIHigh Effort Reasoning60

Note: Scores from Artificial Analysis Intelligence Index. Claude Fable 5.1 leads in general intelligence among 161 ranked models.

Artificial Analysis LLM Leaderboard
Artificial Analysis LLM Leaderboard


Open-Source Leaders

ModelParametersNotable StrengthsKey Score
Kimi K3N/ATop-ranked open-source overallN/A
GLM-5.2N/AStrong general capabilitiesN/A
DeepSeek V4N/ACost-effective performanceN/A
Gemma 4N/AEfficient multimodalN/A
Nemotron 3 Ultra550B (A55B)High-end reasoningN/A

Note: Specific numerical scores for these open models vary by task; they are widely cited as rivals to frontier models in recent comparisons.


Benchmark Deep Dive

The Rise of Agentic Inference Benchmarks

MLCommons' release of MLPerf Inference v6.1 marks a significant pivot in how we evaluate AI performance. For years, benchmarks focused heavily on throughput and latency for static tasks like image classification or single-turn text generation. However, the introduction of the "Agentic Inference" test acknowledges that modern AI deployment is increasingly defined by an agent's ability to plan, execute tools, and maintain state over multiple steps. This new test suite moves beyond simple token-per-second metrics to measure how efficiently models handle the complex, variable-length interactions inherent in agentic workflows.

The results reveal that while frontier models like Claude Fable 5.1 and GPT-6 Astra dominate in raw intelligence scores, efficiency in agentic tasks is becoming a key differentiator. Practitioners should note that high MMLU or GPQA scores do not necessarily correlate with low latency in multi-step agent loops. The v6.1 results highlight that specialized models, such as those optimized for tool-use, may offer better real-world performance for automation tasks despite lower aggregate intelligence scores.

This shift is critical for developers building autonomous systems. The new benchmarks provide a more realistic view of the computational costs associated with "thinking" agents. As agentic AI becomes more prevalent, the ability to compare models on their ability to learn and perform complex actions—rather than just retrieve information—is essential for optimizing both cost and speed in production environments.


Analysis & Trends

  • State of the art: Claude Fable 5.1 currently holds the top spot on the Artificial Analysis LLM Leaderboard with an Intelligence Index score of 53, followed closely by GPT-6 Astra variants. In the open-source realm, Nemotron 3 Ultra and Kimi K3 are pushing the boundaries of what open-weight models can achieve in reasoning and coding.
  • Open vs. Closed gap: The gap continues to narrow, with models like Gemini 3.8 Flash achieving high intelligence scores (59) while maintaining competitive speeds. Open-source models like Qwen3.5 0.8B are now offering extremely low-cost options ($0.01 per 1M tokens) with non-trivial reasoning capabilities.
  • Cost-performance: Mercury 2 and Celeris-1 are leading in speed (1463 t/s for Celeris-1), highlighting that raw intelligence is no longer the only metric for model selection. Cost-efficiency is driving adoption of smaller, specialized models like Qwen3.5 0.8B.
  • Emerging patterns: There is a clear trend toward "Reasoning" variants of models (e.g., GPT-5.5 xhigh, Claude Opus 4.8 Max Effort), indicating that users are willing to trade higher latency for better accuracy in complex tasks.

What to Watch Next

  • Agentic Benchmark Adoption: Watch for vendor-specific results on the new MLPerf Agentic Inference tests, which will likely become a standard metric for enterprise AI agents.
  • Union Alpha Preview Data: As Union Alpha AI moves from limited preview to wider access, independent benchmarks on its performance vs. cost will be crucial for the low-budget segment.
  • Nemotron 3 Ultra Performance: Detailed breakdowns of Nemotron 3 Ultra's 550B parameter model will show if large open-weight models can sustainably compete with closed-source frontier models in long-context reasoning.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QWhat is the Agentic Inference benchmark?
  • QHow does Claude Fable 5.1 rank?
  • QWhat are Union Alpha AI's specs?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.