AI Benchmarks & Leaderboard — 2026-09-18
MLCommons has released MLPerf Inference v6.1, setting a new participation record and introducing benchmarks for Agentic Inference. Meanwhile, Artificial Analysis continues to track the rapid evolution of open-source models, with new entrants like MiniMax-M3 and Nemotron 3 Ultra joining the leaderboard alongside established leaders like Claude Fable 5.1 and GPT-6 Astra. The latest data highlights a narrowing gap in intelligence scores for smaller, faster models, particularly in reasoning tasks.
AI Benchmarks & Leaderboard — 2026-09-18
New Model Releases & Updates
MLPerf Inference v6.1 by MLCommons
- Type: Benchmark Suite Update (Open Standard)
- Key benchmarks: Introduces two new tests: Agentic Inference and Multimodal Generation.
- vs. Previous best: Sets a new participation record with broader hardware vendor support compared to v6.0.
- What's notable: The new "Agentic Inference" test evaluates models on their ability to perform complex, multi-step actions rather than just static text generation, reflecting the industry shift toward autonomous agents.

Union Alpha AI by Union Labs
- Type: Closed-source (Limited Preview)
- Key benchmarks: Verified specs currently under evaluation; available via OpenRouter and Cloudflare.
- vs. Previous best: Positioned as a free anonymous model in limited preview, aiming to disrupt cost structures for entry-level inference.
- What's notable: Offers free access through specific platforms, though full benchmark performance details remain pending wider release.
Recent Additions to Artificial Analysis Leaderboard
- Models: MiniMax-M3, Kimi K2.7 Code, DeepSeek V4 Flash Vision, Nemotron 3 Ultra 550B A55B.
- Key benchmarks: Gemini 3.8 Flash recently scored 59 on the Artificial Analysis Intelligence Index.
- vs. Previous best: These models are competing directly with frontier closed-source models in specific niches like coding (Kimi K2.7 Code) and vision-reasoning (DeepSeek V4 Flash).
Leaderboard Snapshot
Frontier Models (Closed-Source)
| Model | Provider | Notable Strengths | Key Score |
|---|---|---|---|
| Claude Fable 5.1 | Anthropic | Highest Intelligence Index | 53 |
| GPT-6 Astra | OpenAI | Strong Reasoning (xhigh/max) | N/A |
| Gemini 3.8 Flash | Speed/Intelligence Balance | 59 | |
| Claude Opus 4.8 | Anthropic | Adaptive Reasoning (Max Effort) | 61 |
| GPT-5.5 | OpenAI | High Effort Reasoning | 60 |
Note: Scores from Artificial Analysis Intelligence Index. Claude Fable 5.1 leads in general intelligence among 161 ranked models.
Open-Source Leaders
| Model | Parameters | Notable Strengths | Key Score |
|---|---|---|---|
| Kimi K3 | N/A | Top-ranked open-source overall | N/A |
| GLM-5.2 | N/A | Strong general capabilities | N/A |
| DeepSeek V4 | N/A | Cost-effective performance | N/A |
| Gemma 4 | N/A | Efficient multimodal | N/A |
| Nemotron 3 Ultra | 550B (A55B) | High-end reasoning | N/A |
Note: Specific numerical scores for these open models vary by task; they are widely cited as rivals to frontier models in recent comparisons.
Benchmark Deep Dive
The Rise of Agentic Inference Benchmarks
MLCommons' release of MLPerf Inference v6.1 marks a significant pivot in how we evaluate AI performance. For years, benchmarks focused heavily on throughput and latency for static tasks like image classification or single-turn text generation. However, the introduction of the "Agentic Inference" test acknowledges that modern AI deployment is increasingly defined by an agent's ability to plan, execute tools, and maintain state over multiple steps. This new test suite moves beyond simple token-per-second metrics to measure how efficiently models handle the complex, variable-length interactions inherent in agentic workflows.
The results reveal that while frontier models like Claude Fable 5.1 and GPT-6 Astra dominate in raw intelligence scores, efficiency in agentic tasks is becoming a key differentiator. Practitioners should note that high MMLU or GPQA scores do not necessarily correlate with low latency in multi-step agent loops. The v6.1 results highlight that specialized models, such as those optimized for tool-use, may offer better real-world performance for automation tasks despite lower aggregate intelligence scores.
This shift is critical for developers building autonomous systems. The new benchmarks provide a more realistic view of the computational costs associated with "thinking" agents. As agentic AI becomes more prevalent, the ability to compare models on their ability to learn and perform complex actions—rather than just retrieve information—is essential for optimizing both cost and speed in production environments.
Analysis & Trends
- State of the art: Claude Fable 5.1 currently holds the top spot on the Artificial Analysis LLM Leaderboard with an Intelligence Index score of 53, followed closely by GPT-6 Astra variants. In the open-source realm, Nemotron 3 Ultra and Kimi K3 are pushing the boundaries of what open-weight models can achieve in reasoning and coding.
- Open vs. Closed gap: The gap continues to narrow, with models like Gemini 3.8 Flash achieving high intelligence scores (59) while maintaining competitive speeds. Open-source models like Qwen3.5 0.8B are now offering extremely low-cost options ($0.01 per 1M tokens) with non-trivial reasoning capabilities.
- Cost-performance: Mercury 2 and Celeris-1 are leading in speed (1463 t/s for Celeris-1), highlighting that raw intelligence is no longer the only metric for model selection. Cost-efficiency is driving adoption of smaller, specialized models like Qwen3.5 0.8B.
- Emerging patterns: There is a clear trend toward "Reasoning" variants of models (e.g., GPT-5.5 xhigh, Claude Opus 4.8 Max Effort), indicating that users are willing to trade higher latency for better accuracy in complex tasks.
What to Watch Next
- Agentic Benchmark Adoption: Watch for vendor-specific results on the new MLPerf Agentic Inference tests, which will likely become a standard metric for enterprise AI agents.
- Union Alpha Preview Data: As Union Alpha AI moves from limited preview to wider access, independent benchmarks on its performance vs. cost will be crucial for the low-budget segment.
- Nemotron 3 Ultra Performance: Detailed breakdowns of Nemotron 3 Ultra's 550B parameter model will show if large open-weight models can sustainably compete with closed-source frontier models in long-context reasoning.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.