에이전트 하네스 엔지니어링 리포트 살펴보기
This week's agent harness engineering update covers Anthropic's new harness patterns, Kestra's orchestration framework comparison, and structural scaffolding solutions for five recurring failure modes. Production teams are actively discussing memory corrosion, exploration stagnation, and the "prevent-persist-redirect" three-principle structure to solve them.
에이전트 하네스 엔지니어링 주간 리포트 — 2026-09-30
This Week's Headlines
-
Anthropic's 3 Harness Design Patterns — Anthropic's agent harness design guide, originally published in April 2026, is gaining renewed attention in the AI community. The core focus revolves around "building tool-based foundations Claude already understands," "removing harness assumptions as capabilities improve," and "carefully setting UX, cost, and safety boundaries," guiding teams on which scaffolding to maintain, add, or remove over time.
-
Kestra's 12-Framework Comparison Analysis — A comparison of major orchestration platforms (LangGraph, MAF 1.0, Claude SDK, OpenAI Agents, Google ADK, CrewAI, etc.) based on control, state management, approval, and self-hosted deployment. It emphasizes that choices in control plane design and state checkpointing can impact production reliability by up to 30 points.
-
Analytics Insight's Latest Framework Evaluation — A comparison released three days ago details the features, pros, cons, and use cases of CrewAI, OpenAI Agents SDK, Microsoft Agent Framework, and Google ADK. It specifically covers how framework selection in multi-agent orchestration affects deployment complexity and operational costs.
-
220+ Experiment Data: 5 Recurring Failure Modes and 3-Principle Solutions — Literature triage from the AgentCollaborationSystems project two days ago identified consistent failure modes across two independently developed systems (infrastructure vulnerability, agent memory corrosion, exploration direction stagnation, iteration cost asymmetry, and metric fixation). A structural design solving these through a "prevent-persist-redirect" three-principle scaffolding was presented.
Framework & Tooling Updates
Kestra AI Agent Orchestration Framework — Orchestrator 2026 Edition
- What's new: Integrated comparison interface for 12 major frameworks, state checkpointing validation tools, approval workflow templates, and MCP-integrated monitoring dashboard.
- Why it matters: Enables production teams to clearly evaluate the trade-offs between control domains and abstraction levels when choosing a framework. Since the state resilience of long-running agents depends entirely on the chosen framework's checkpointing mechanism, the importance of initial architectural decisions is maximized.
- Migration notes: When transitioning from existing LangChain-based agents, pay attention to signature changes in memory systems and tool registration APIs. Migrating from CrewAI to OpenAI Agents requires redesigning the role definition systems as they differ.

Research & Evaluation
Five Recurring Failure Modes in Agent Harnesses: Empirical Study from 220+ Experiments
- Authors / Org: AgentCollaborationSystems Contributors (2026-09-28)
- Core finding: Discovery of 5 recurring failure modes consistently observed across 220+ iteration experiments in two independently designed book recommendation pipeline agent systems: (1) Infrastructure vulnerability — intermittent failures due to missing timeout/retry policies, (2) Agent memory corrosion — loss of early learnings when exceeding the context window, (3) Exploration direction stagnation — getting stuck in local optima due to lack of tool call diversity, (4) Iteration cost asymmetry — low initial iteration costs followed by an explosion of LLM calls in later stages, (5) Metric fixation — optimization of evaluation metrics diverging from actual user experience.
- Implication for harness design: Structural scaffolding design countering each failure mode is necessary: Prevent — infrastructure circuit breakers, pre-configured tool call sampling policies; Persist — memory compression/summarization mechanisms, context prioritization; Redirect — reward signal reset upon detecting exploration stagnation, rollbacks based on cost thresholds. This design scales to handle increasing iteration costs linearly.
GuardianAgentBench: Comprehensive Safety and Guardrail Evaluation for Production Agent Platforms
- Authors / Org: AI Safety Evaluation Team (July 2026)
- Core finding: Testing 580 safety scenarios across 6 production agent platforms (LangChain, LlamaIndex, CrewAI, Mastra, Haystack, Strands) revealed that guardrail effectiveness heavily relies on platform-based implementations and framework defaults. Platforms with weak tool usage validation and permission systems showed privilege escalation defense rates below 40%.
- Implication for harness design: Implementing a 3-tier validation layer (syntax, permission, intent) prior to tool calls is essential. Framework defaults are insufficient, and integrating application-level policy engines (e.g., Lakera Guard, AWS Bedrock Guardrails) is recommended.
Production Patterns & Practitioner Insights
Context Window Management: Preventing Memory Corrosion Learned from 220+ Agent Iterations
- Context: Book recommendation agent losing initial user intent recognition after more than 30 tool interactions.
- Problem: In a fixed-size context window (4K–128K tokens), as it transitions to mid-stages (iterations 10–50), new tool call results accumulate while initial instructions and user feedback are progressively pushed out. As a result, the agent gradually drifts away from its original goal.
- Solution / Takeaway: Introduce a 3-tier memory structure: (1) Active memory (2–4K tokens) — last 3 interactions + current task, (2) Summary memory (1–2K) — compressing the first 5 interactions into a single LLM summary, (3) Index memory (search-based) — embedding all tool call results and retrieving relevant learnings via similarity-based search. While this structure increases total token count by 20–30%, it cuts iteration costs by 40% and improves final quality by 14%.
Production Templates and Cost Optimization: Reducing 16 Hours of Development with AgentKit
- Context: The 5th agent building team selecting and designing after comparing LangGraph, CrewAI, and AutoGen.
- Problem: The first 4 agents each chose different frameworks, leading to steep learning curves and having to re-implement tool registration, memory initialization, and error handling patterns from scratch every time. The repetitive scaffolding cost alone took 16 hours for the 5th agent.
- Solution / Takeaway: Operating a framework-agnostic template repository providing (1) standard project structure (config, tools, memory, evaluators directories), (2) tool registration adapters (LangChain, CrewAI, OpenAI compatible), (3) memory interface abstraction (SQLite, Redis backends), and (4) reusable evaluation tools (BLEU, RougeL, cost tracking). This reduced setup time for the 5th agent to 2 hours, and future agents can configure basic structures within 30 minutes after selecting the initial template.

AgentScaffold: Memory Systems for Peer Review and Continuous Improvement
- Context: Code generation agent unable to re-evaluate the rationale of past decisions after complex multi-file refactoring tasks.
- Problem: The agent only logged tool call logs, while the reasoning behind each decision, expected vs. actual outcomes, and alternative review processes were not explicitly saved. Consequently, tracking "why it did that" after a failure became impossible.
- Solution / Takeaway: learnings_tracker mechanism: For every iteration, record (1) Intent — specific goal to solve in this iteration, (2) Plan — expected tool call sequence, (3) Execution — actual call results and diffs, (4) Retrospective — gap between expectation and reality, unexpected findings, and improvements for the next iteration. Storing this as structured JSON/Markdown allows future or similar tasks to query these learning logs and avoid similar mistakes. Production teams adopting this saw an average 35% reduction in debugging time per iteration.

Trending OSS Repositories
-
awesome-harness-engineering — A curated list combining Anthropic's 4 harness design patterns, tool call schemas, MCP integration cases, and memory architecture comparisons. Updated 1 day ago, community-driven content.
-
AgentCollaborationSystems — A repository implementing the 5 recurring failure modes and "prevent-persist-redirect" 3-principle scaffolding design derived from 220+ experiments. Active discussions ongoing via literature triage issues.
-
OC-TL — Bain & Company's internal agent collaboration platform development project retrospective. Completed monorepo scaffolding, BFF with Entra SSO, agent skeleton, and initial data collection during Sprint 1 (Sep 1–14). Progressing via open collaboration.
Deep Dive: Validation of 5 Recurring Failure Modes and Structural Solutions
Over the past two years, agent harness engineering has moved beyond the formal understanding of "model + prompt + tools" to identify structural problems that repeatedly surface in production workloads. The 220+ experiment data from the AgentCollaborationSystems team is concrete proof of this transition.
Stage 1: Infrastructure Vulnerability → Prevention Principle Every time an agent calls a tool, external failures such as network timeouts, API rate limits, and out-of-memory errors can occur. Early harnesses handled this with simple retries, but unlimited retries lead to skyrocketing costs and infinite recursion. Solution: Circuit breaker pattern — wait 30 seconds after 3 consecutive failures, then attempt recovery with a single test request, and return a failure after 5 total retries. Configuring this per tool ensures that a failure in one tool does not take down the entire agent.
Stage 2: Memory Corrosion → Persistence Principle Initial user intent ("This book must be readable by elementary school students", "Must be historical fiction, not sci-fi") disappears after 30 tool calls. In a fixed 4K token window, as recent call results (200–300 tokens each) accumulate, the very first instructions get pushed out. Solution: Hierarchical memory — (1) Never remove system prompts, (2) Summarize initial user intent (1000 tokens down to 100 via 1–2 LLM calls), (3) Index tool call results via embeddings, keeping only the last 3 results + top 2 query-similar items in the active context. Experiments showed that while total token usage increased by 20%, final quality (user intent fulfillment) improved by 14%.
Stage 3: Exploration Stagnation → Redirection Principle The agent falls into local optimization by repeatedly calling the same tool (e.g., "search_books") with only minor tweaks to search queries. Other tools (e.g., "get_author_bio", "list_related_genres") remain uncalled. Solution: Tool call diversity policy — track tool distribution across the last N calls, and if a tool's call frequency exceeds 80%, block its usage and force alternative tool exploration. Alternatively, if the agent calls the same tool a set number of times (e.g., 15 times), reset reward signals to encourage exploring new directions.
Stage 4: Iteration Cost Asymmetry → Cost Threshold Management Early iterations (1–5) are low cost ($0.01–$0.05 each), mid-stage iterations (10–30) increase ($0.1–$0.3/call), and later iterations (after 40) surge ($0.5–$1.0/call) due to context window accumulation. Solution: Dynamic cost tracking — calculate cumulative costs per iteration, and upon hitting budget thresholds (e.g., $5), immediately terminate or switch to a cost-saving mode (switching to smaller models, enforcing stronger context compression).
Stage 5: Metric Fixation → Transition to User-Centric Evaluation Automated metrics like BLEU and RougeL improve, but actual user satisfaction drops by 37%, showing a gap between evaluation metrics and actual tasks. Solution: Multi-layer evaluation — (1) Automated metrics (speed, token efficiency), (2) Heuristics (user intent fulfillment checklist), (3) Sample human evaluation (inspecting 10 samples every 100 iterations). Adopting this allowed production teams to align metric improvements with user satisfaction once again.
These 5 failure modes occur regardless of framework selection, and solutions must be implemented at the application harness level, not the framework level. Whether using LangGraph or CrewAI, layering this 3-principle structure on top ensures both stability and efficiency.
What to Watch Next Week
-
Anthropic's Claude Agent SDK v2.1 Release Expected — Anticipated official support for memory compression APIs and standardized context selective retention mechanisms. A framework-level implementation of the "Persistence Principle" above is likely.
-
OpenAI Agents SDK — Cost Tracking Dashboard — Official tools expected to be released for real-time tracking of token usage per iteration and cumulative costs (currently in internal testing). This will bring operational visibility to the "iteration cost asymmetry" issue.
-
GuardianAgentBench v2 Public Release — Expanding from 580 to 1,000+ scenarios, publishing safety scores for each framework/platform on a public leaderboard. This will enable quantitative comparisons of safety during framework selection.
Reader Action Items
-
If building a production agent, immediately adopt the "prevent-persist-redirect" 3-principle checklist — especially if more than 10 tool calls are expected, include a hierarchical memory structure in your initial design. Refer to templates (awesome-harness-engineering, AgentCollaborationSystems) and adapt them to your organization's framework.
-
Review Kestra's 12 comparative analyses before choosing a framework — specifically if you have requirements for state checkpointing, context management, and multi-agent coordination, explicitly test the chosen framework's capabilities in these three areas. Initial selection errors carry high migration costs (abandoning existing learnings, retraining teams), so choose carefully.
-
Integrate cost tracking and memory monitoring into your production monitoring pipeline — create dashboards tracking token growth rates per iteration for early detection of "iteration cost asymmetry." A signal where costs surge 3x or more starting from the 5th iteration is a concrete sign of memory corrosion or exploration stagnation, which should automatically trigger dynamic responses (context compression, enforcing tool diversity).
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.