EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
Abstract
The study introduces a benchmark to evaluate LLM agents under evolving tool, skill, and agent harnesses, revealing persistent gaps in retention, adaptation, and harness-induced forgetting.
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EVOHARNESSBENCH, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents). Unlike existing continual-learning benchmarks for agents, which typically place non-stationarity (i.e., what changes over time) in the task stream while keeping the harness fixed, EVOHARNESSBENCH places non-stationarity in the externally supplied harness itself. It contains 17 multi-stage harness streams constructed deterministically from verifier-based benchmarks, comprising 802 tasks, 520 tools, 42 skills, and 62 agents. We evaluate two complementary settings corresponding to the central challenges of harness evolution: deployment evaluation, which isolates retention of previously accessible competence as the harness expands, and self-evolving adaptation evaluation, which tests whether accumulated experience remains useful as new capabilities are introduced. Our results reveal three persistent gaps. First, harness expansion alone can degrade performance on previously solved tasks, producing harness-induced forgetting. Second, gains from self-evolving adaptation remain inconsistent across stages of harness evolution, capability axes, and environments. Third, retention and adaptation can pull in different directions: preserving earlier competence does not necessarily improve adaptation to newly introduced capabilities, and vice versa. These results establish harness evolution as a distinct challenge for building agents that can keep pace with an evolving harness while preserving previously effective behavior.
Community
With coding agents like Claude Code or Codex, we add tools and skills whenever they seem useful. Each addition feels like an upgrade. But what if an expanding harness quietly makes the agent worse at tasks it could already solve?
Our results show that it can.
More surprisingly, even today’s memory-based and other self-evolving methods cannot reliably adapt to new capabilities while retaining earlier competence.
Introducing EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents (2026)
- Demystifying Agent Skills: Why They Work-Until They Don't (2026)
- MemoHarness: Agent Harnesses That Learn from Experience (2026)
- Harness Continual Learning: Continual Adaptation Beyond Model Parameters (2026)
- StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments (2026)
- JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution (2026)
- ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.04280 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 3
ZixuanKe/evovling_skills
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper