SOURCE-LINKED INTELLIGENCE
Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents
arXiv · AI, language, vision and robotics · article · Sep 10, 2026 · UTC
Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing failure-driven approaches often treat observed agent failures as direct evidence for harness modification. A key challenge in failure-driven harness evolution is that observed failures can reflect either limitations of the underlying model or systematic deficiencies of the harness. Directly optimizing against individual failures can therefore induce model-specific accommodation and impair generalization across
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
First collected: 2026-09-20T19:02:05.452Z. This is not the publication date.
Observed changes
AIIC observation times, not verified publisher revision times. Up to eight recent revisions.
2026-09-24T23:02:23.507Z
- summary:
Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on iterative search, repeatedly evaluating and revising candidate harnesses based on execution feedback from task instances. While this paradigm enables continuous harness optimization, it incurs substantial time overhead due to repeated agent executions and code modifications, and may overfit to observed tasks and specific failure patterns, resulting in degraded general → Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing failure-driven approaches often treat observed agent failures as direct evidence for harness modification. A key challenge in failure-driven harness evolution is that observed failures can reflect either limitations of the underlying model or systematic deficiencies of the harness. Directly optimizing against individual failures can therefore induce model-specific accommodation and impair generalization across