Reasoning Models and RL Research: Test-Time Compute — 2026-09-26
The week's biggest story came from DeepSeek, which published a 31-page paper on agent (AI) RL training infrastructure with CEO Liang Wenfeng among 130+ co-authors, revealing a production sandbox platform running 3 million sandboxes per day. Meanwhile, Chinese media reported DeepSeek is reportedly preparing an 8-trillion-parameter MoE model, and a fresh September 2026 benchmark roundup shows Claude Opus 5.5 and GPT-6 Astra leading coding and agentic tasks.
Reasoning Models and RL Research: Test-Time Compute — 2026-09-26
Top developments
DeepSeek publishes agent training infrastructure paper with Liang Wenfeng as co-author
On September 23, DeepSeek released a 31-page arXiv paper describing its production-grade sandbox platform, DSec (DeepSeek Elastic Compute), aimed at training AI agents more efficiently while curbing anomalous agent behavior — a growing global concern. The paper lists over 130 co-authors, with founder Liang Wenfeng signing last, and reports running 3 million sandboxes per day on a single cluster, including defenses against AI "cheating" in environments. For the RL research community, this is a rare production-level disclosure of the infrastructure layer that verifiable-reward agent training depends on.

DeepSeek reportedly joins the ultra-large parameter race with 8T-parameter MoE plans
Chinese outlets reported on September 22–23 that DeepSeek is planning to train an 8-trillion-parameter model, with the parameter competition shifting toward MoE architectures. 21财经 framed it as "even Liang Wenfeng is competing on 'ultra-large'". If confirmed, this would mark a notable pivot for the efficiency-focused lab and could reshape compute budgets for reasoning-model post-training at scale. (Reported plans, not an official announcement.)

September 2026 coding leaderboard: Claude Opus 5.5 leads Terminal-Bench 4.0 at 66.4%
A September 2026 roundup ranks 13 models on coding benchmarks and cost per task, finding Claude Opus 5.5 leading Terminal-Bench 4.0 at 66.4% at $4/$20 per million input/output tokens, while GPT-6 Astra takes the top spot on the Frontend Code Arena. These terminal-level agentic coding scores are increasingly the practical yardstick for reasoning models trained with RL on verifiable rewards.
Local view
Chinese-language coverage of the DeepSeek paper was extensive and fast: 澎湃新闻 (via 21财经) produced multiple pieces within a day, emphasizing the paper's scale (~130 co-authors), the 3-million-sandboxes-per-day figure, and the anti-"cheating" mechanisms. Sohu highlighted that "if large-model training is about teaching AI to think, agent training is about getting AI to actually do the work — starting with solving the 'getting started' problem". On Zhihu, a weekly model tracker (2026/09/21–25) noted agent-model updates across closed-source labs: GPT-6 Sol and GPT-6 Luna, Claude Opus 5.5, and Grok 4.7, alongside domestic open-source MiMo updates.
Context & numbers
- DeepSeek's DSec paper: 31 pages, 100+ (about 130) co-authors, 3 million sandboxes run daily on one cluster
- Claude Opus 5.5: 66.4% on Terminal-Bench 4.0 at $4/$20 per million tokens; GPT-6 Astra: #1 on Frontend Code Arena
- Reported (unconfirmed) DeepSeek next-model target: 8 trillion parameters, MoE architecture
- A four-model comparison of Chinese open-weight labs (DeepSeek, Qwen, GLM, Kimi) published this week covers current model lineup, pricing, and self-hosting requirements for production teams
On the radar
- The 8T-parameter DeepSeek model remains a rumor flagged by Chinese tech media; watch for an official confirmation or periodic benchmark leaks
- DSec's anti-reward-hacking ("cheating" prevention) techniques may set a template for how verifiable-reward agent RL environments are built; independent reproductions would be worth tracking
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.