Retrieval, Memory and Long-Context Research — 2026-09-26
This week's story is million-token context going mainstream in open source: Xiaomi's freshly released MiMo-V2.6 series ships with 1M-token context and full multimodality, while Moonshot AI's Kimi roadmap leaks point toward K3.1 in October. On the research side, a new alphaXiv paper characterizes system-level requirements for stateful, long-horizon agent memory workloads — a reminder that persistence, not window size, is the real bottleneck.
Retrieval, Memory and Long-Context Research — 2026-09-26
Top developments
Xiaomi MiMo-V2.6 launches with 1M-token context across all tiers
On September 22, Xiaomi released and open-sourced the MiMo-V2.6 series — flagship MiMo-V2.6-Pro and efficiency-focused MiMo-V2.6-Flash — described as native full-modality models, with a Pro-UltraSpeed variant that the company says achieves up to 20× the output speed of Pro while maintaining quality. On September 23, the B.AI platform confirmed the entire series carries 1M-token context plus text, image, video and audio understanding, live via API and web chat. Xiaomi also claims MiMo-V2.6 has surpassed Kimi K3 and GLM-5.3 to become the highest-ranked open-source model on the AA index. For RAG builders, another fully open 1M-context multimodal frontier competitor widens routing options at the top of the leaderboard.

New paper: system characterization of stateful long-horizon agent workloads
A new alphaXiv paper, "Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads" (posted within the last week), analyzes how agents deployed on long-horizon tasks must persistently store and retrieve information over extended interaction histories — and what that demands from the underlying systems stack. This is directly relevant to agent memory system design: it frames memory as a systems problem (storage, indexing, retrieval latency) rather than just a prompt-engineering one.

Moonshot roadmap: Kimi K3.1 reportedly landing in October
On September 23–24, Wccftech reported (relayed by Chinese press) that Moonshot AI is preparing Kimi K3.1 with three reasoning-efficiency tiers — Low, High, Max — expected in October 2026. A separate, more speculative leak circulating this week names Kimi K4, GLM-5.5/5.4 Flash and DeepSeek V4.1P as unannounced models; treat that list as rumor only. Moonshot's Kimi line is closely watched for long-context capability, so K3.1's tiered serving could matter for cost-per-million-token comparisons.
Builder guidance: short-term vs long-term agent memory architecture
Activepieces published a fresh guide (2 days ago) covering how short-term vs long-term memory architectures determine how agents manage working context and retrieve historical data. Together with Unite.AI's explainer on short-term, long-term, episodic and semantic memory as a designed system — not an endlessly growing prompt — it signals that the memory-architecture market is consolidating around layered schemas.

Local view
Chinese-language coverage is dominated by the MiMo launch: Sohu Tech reported Lei Jun spent roughly RMB 23 million in six days on training the new model, and quoted researcher Luo Fuli saying the innovation and engineering difficulty exceeded DeepSeek R1. Zhihu's weekly model tracker (updated 2026/09/23) lists this week's agent-model refreshes across GPT-6 Sol/Luna, Claude Opus 5.5, Grok 4.7 and domestic open models. Huxiu's commentary (4 days ago) argues that strong model benchmarks are only an "entry ticket" for Moonshot, which still needs to prove itself commercially.

Context & numbers
- MiMo-V2.6: 1M-token context, full-modality (text/image/video/audio), API + web chat via B.AI since September 23
- MiMo-V2.6-Pro-UltraSpeed: up to 20× Pro's output speed, per Xiaomi
- Xiaomi claims MiMo-V2.6 is currently the top-ranked open-source model on the AA index, overtaking Kimi K3 and GLM-5.3
- Kimi K3.1: three reasoning-strength tiers (Low, High, Max), expected October 2026
On the radar
- Kimi K3.1 launch, expected October 2026 — watch its long-context and reasoning-tier pricing
- The unannounced Kimi K4 / GLM-5.5 / DeepSeek V4.1P leak — rumor, unverified, but worth monitoring for long-context claims
- Aging-but-still-relevant caveat to keep in mind when judging the 1M-token claims above: RULER, MRCR v2 and NoLiMa scores have shown advertised vs. effective context diverging by 30–60 points for multi-fact retrieval past 200K tokens
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.