오늘의 AI 모델 벤치마크 및 성능 비교
2026년 8월 2-3일 AI 벤치마크의 핵심은 GPT-5.6의 대폭적인 가격 인하와 DeepSeek V4-Flash의 재훈련 성과입니다. OpenAI가 GPT-5.6 Luna 가격을 80% 낮추고 Sol Fast 티어를 출시해 가성비를 높였으며, DeepSeek V4-Flash는 9개 에이전트 벤치마크에서 V4-Pro-Preview를 뛰어넘는 성능을 보여주었습니다.


Get the latest daily updates, rankings, and performance breakdowns for the newest AI models at a glance.
2026년 8월 2-3일 AI 벤치마크의 핵심은 GPT-5.6의 대폭적인 가격 인하와 DeepSeek V4-Flash의 재훈련 성과입니다. OpenAI가 GPT-5.6 Luna 가격을 80% 낮추고 Sol Fast 티어를 출시해 가성비를 높였으며, DeepSeek V4-Flash는 9개 에이전트 벤치마크에서 V4-Pro-Preview를 뛰어넘는 성능을 보여주었습니다.

DeepSeek V4-Flash outperformed its own V4-Pro-Preview across all 9 agent benchmarks, highlighted by a 645% leap in DeepSWE. OpenAI has slashed prices for its GPT-5.6 Luna and Terra models and launched a faster Sol Fast tier. Meanwhile, Thinking Machines released Inkling-Small, a multimodal model that packs the same punch as the original Inkling at just one-quarter the size.

Between July 28 and 30, big news hit the AI scene: Exabase’s M-1 memory engine topped the BEAM benchmark, and Microsoft’s MAI-Cyber-1-Flash hit 95.95% accuracy in cybersecurity tasks. Also, the latest Model Context Protocol (MCP) spec and the new ‘mcpbench’ test are really pushing the limits of GPT-5.6 and Claude Opus.

2026년 7월 26-27일 기준, Claude 5 패밀리와 OpenAI의 GPT-5.6 시리즈가 최고 성능을 유지하는 가운데 Moonshot의 Kimi K3가 뒤를 잇고 있습니다. 코딩 역량 면에서는 GPT-5.6 Sol이 96.2%, Claude Fable 5가 95.0%를 기록하며 치열하게 경쟁 중입니다.

The AI model landscape since July 21st has been defined by intensifying open-source competition and security concerns among frontier models. Moonshot’s Kimi K3 (2.8T parameters) has outperformed Claude Fable 5 in specific benchmarks, while evaluation standards are becoming increasingly stringent to combat benchmark saturation.
중국 Moonshot AI의 2.8조 파라미터 모델 Kimi K3가 프론트엔드 코드 벤치마크에서 Anthropic의 Claude Fable 5를 앞질렀습니다. 이는 중국 모델이 미국 최상위권 모델과 대등하게 경쟁할 수 있음을 보여주는 사례로, AI 업계의 판도를 흔들고 있습니다.

Over the last 24 hours, Claude Fable 5 has maintained its dominance in the SWE-bench evaluation with 95% accuracy, while the 97.5B parameter open-source model Inkling from Thinking Machines Lab has emerged as a new competitor with video and audio understanding capabilities. BenchLM.ai’s July 17 update now provides a comprehensive benchmarking platform comparing 284 different models.

The past 24 hours were dominated by the wide release of OpenAI’s GPT-5.6, while Claude Fable 5 continues to lead coding tasks with a 95.0% success rate on SWE-bench. GPT-5.6 Sol has achieved a 54% improvement in token efficiency for agentic coding.

The last 24 hours in the AI world have been buzzing with OpenAI’s official launch of GPT-5.6 and fresh benchmark results for new models. Competition in China is heating up, with Tencent’s Hy3 model matching the performance of GLM-5.2 and DeepSeek-V4, while a tight race continues at the top between Claude Fable 5 and GPT-5.6.

"OpenAI가 GPT-5.6 프리뷰를 통해 새로운 세 계층(Sol, Terra, Luna) 모델 체계를 도입했습니다. 한편, Claude Opus 4.8은 SWE-bench Verified에서 88.6%를 기록하며 코딩 실무 분야에서 압도적인 선택지로 떠오르고 있습니다."

The release of Google's Gemini 2.5 Pro with Deep Think on June 22 is shaking up the leaderboard. Claude Opus 4.8 currently leads with an AA Index of 61.4, while intense global competition continues between GPT-5.5, GLM-5.2, and other top-tier models.

6월 19일부터 21일까지 AI 모델 벤치마크 분야의 핵심 이슈는 **DeepSWE 벤치마크의 등장**입니다. 이를 통해 그동안 가려졌던 모델 간 성능 차이가 분명해졌습니다. 더불어 중국 스타트업 Z.ai가 자사 모델 GLM-5.2가 GPT-5.5를 주요 지표에서 앞섰다고 주장하며 글로벌 AI 경쟁이 더욱 뜨거워지고 있습니다.

OpenAI’s new LifeSciBench and China-based Z.ai’s claim that their GLM-5.2 model beats GPT-5.5 are shaking up the AI rankings. Meanwhile, Codex + GPT-5.5 is leading the Terminal-Bench coding agent race with 83.4%, NVIDIA’s Blackwell is crushing it in MLPerf Training 6.0, and Nature Medicine finds that general-purpose LLMs are actually outperforming specialized medical AI.

2023-2024년에 출시된 주요 AI 벤치마크들이 포화 상태에 이르렀습니다. 최근 평가에서 NVIDIA가 에이전틱 AI 코딩 성능에서 앞서가는 모습을 보였고, 오픈소스 모델 중에는 GLM-5(85점)가 선두를 달리고 있습니다. 이제 단일 모델보다는 작업별 특화 모델로 시장 흐름이 바뀌고 있네요.

2026년 6월 13일 이후, Anthropic의 최상위 모델 해외 접근 제한과 NVIDIA의 새로운 에이전틱 AI 벤치마크 성과가 업계의 주요 화두로 떠올랐습니다. Anthropic은 Mythos 5와 Fable 5의 해외 지원을 중단했고, NVIDIA는 업계 최초의 에이전틱 코딩 AI 평가에서 뛰어난 성적을 기록했습니다.

The most notable shift in AI benchmarking over the past 24 hours is that the major evaluation metrics released in 2023-2024 have reached a saturation point. Benchmarks like METR, SWE-Bench, CORE-Bench, MLE-Bench, and PostTrainBench are either already maxed out or rapidly approaching their ceiling, highlighting how fast AI capabilities are actually advancing.

지난 24시간 동안 가장 눈에 띄는 AI 소식은 마이크로소프트의 새로운 MAI 모델 시리즈 공개와 트럼프 행정부의 사이버보안 벤치마킹 행정명령입니다. 마이크로소프트의 MAI-Thinking-1은 복잡한 문제 해결을 위해 설계된 첫 추론 전문 모델이며, 미 연방정부는 AI 보안 평가를 위한 표준화 작업을 본격화하고 있습니다.

마이크로소프트가 Build 2026에서 MAI(Microsoft AI) 패밀리의 첫 추론 모델인 MAI-Thinking-1을 선보이며 주목받고 있습니다. 한편, 트럼프 행정부는 첨단 AI 모델의 사이버보안 성능을 평가하는 새로운 벤치마크 프로세스 도입을 위한 행정령에 서명했으며, 2026년 AI 추론 비용이 급격히 낮아지면서 업계 내 경쟁이 한층 더 뜨거워지고 있습니다.

2026년 6월 4일 기준, Microsoft Build 2026에서 발표된 MAI-Thinking-1이 화제입니다. Microsoft의 첫 추론 전용 모델로 높은 효율성과 비용 절감을 내세우네요. 한편, 백악관 행정명령에 따라 고급 AI의 사이버 보안을 평가하는 새로운 정부 차원의 벤치마킹 프로세스도 도입되었습니다.

GPT-5.6이 이번 주 출시를 앞두고 있으며 Mythos 수준의 성능을 제공할 것으로 보입니다. 현재 벤치마크에서는 Claude Opus 4.7이 코딩, GPT-5.5가 에이전트, Gemini 3.1이 추론 분야에서 각각 두각을 나타내고 있습니다.

Create custom signals on any topic. AI curates and delivers 24/7.
Create Signal