recursive-improve

作者 kayba-ai已验证

🪞 Make your agents recursively self-improve

250
Stars
25
Forks
Python
语言
2026/8/23
添加时间

⚠️ 第三方软件声明

本 Skill 为第三方开源软件,独立托管于 GitHub。SkillTip 仅为信息目录,不控制或维护底层仓库。所显示的安全检查为自动化且范围有限,安装前请自行审查源码。

阅读服务条款

安装

添加到你的 Claude Code skills 目录:

# Add to your Claude Code skills
git clone https://github.com/kayba-ai/recursive-improve

快速入门

使用 recursive-improve 等 Skills 的指南。

安全报告

已验证

上次扫描:—

{
  "status": "PASSED",
  "issues": []
}

README.md

recursive improve

Discord Twitter Follow kayba.ai

make your agents recursively self-improve

90% of Claude's code is now written by Claude. Recursive self-improvement is already happening at Anthropic. What if you could do the same for your own agents?

Closing the Loop

You have an agent. It works, most of the time. But it could be better. Solving harder problems, handling more edge cases, wasting fewer tokens. What if it could improve itself, recursively, every time it runs?

Right now, it can't. Your agent is stateless. Every run starts from scratch. The only way to improve it is to manually improve it. There is no compounding of improvements.

recursive-improve closes this loop:

recursive improvement loop

Your agent runs. Every LLM call is captured. Your coding agent analyzes the traces, identifying common failure patterns across runs, and applies targeted fixes. You run it again. It's better.


Get Started

1. Install

uv tool install "recursive-improve[all] @ git+https://github.com/kayba-ai/recursive-improve.git"

Then in your agent's project directory:

cd /path/to/your/agent
recursive-improve init

This creates the /recursive-improve skill files and the eval/traces/ directory.

2. Add tracing to your agent

Add the tracing dependency to your project:

uv add "recursive-improve @ git+https://github.com/kayba-ai/recursive-improve.git"

Two lines. Your agent code stays unchanged, but now your agents execution traces get saved locally.

import recursive_improve as ri

ri.patch()  # auto-captures openai, anthropic, litellm calls

with ri.session("./eval/traces") as run:
    result = my_agent("book a flight to Paris")
    run.finish(output=result, success=True)

Already have traces? Drop them in eval/traces/ and skip to step 4.

3. Run your agent a few times to generate traces

4. Run the improvement loop

Open Claude Code or Codex in your project directory:

/recursive-improve

5. Re-run your agent

Clear old traces and run your agent again so the benchmark measures your improved code:

rm -f eval/traces/*.json
# run your agent the same way as step 3

6. Benchmark

Measure whether your changes actually solved the problems:

/benchmark

Results are stored in eval/benchmark_results.json and auto-compared against the previous run on the same dynamic metrics that were generated for your agent.

CLI alternative: recursive-improve benchmark --label "v1-baseline" and recursive-improve benchmark list

7. Dashboard

Start the interactive dashboard to visualize your improvement cycles:

recursive-improve dashboard          # default: http://localhost:8420
recursive-improve dashboard -p 8080  # custom port

Each improvement cycle lives on its own branch. The dashboard shows before/after metrics for every cycle. See exactly what improved, merge the wins, discard the rest.

Dashboard

8. Run it overnight

/ratchet

An autoresearch-style autonomous loop. It asks you what to optimize, then repeats: improve → run agent → eval → keep or revert. Only improvements survive. Check eval/ratchet_summary.md when you wake up.

[!TIP] Want deeper analysis? Kayba offers managed recursive agent improvement at scale, tailored to your agent.


How It Works

When you run the /recursive-improve skill, it walks through a structured pipeline:

  1. Build context: detects your agent's architecture, tools, and system prompt
  2. Analyze traces: reads your traces, surfaces failure patterns, missed opportunities, recurring errors
  3. Measure: runs built-in detectors (loops, give-ups, errors, recovery) and generates custom domain-specific evaluations from your insights, then computes baselines
  4. Plan: triages each insight into discard / code fix / prompt fix, prioritized by impact
  5. Review: presents the plan for your approval before anything changes
  6. Fix: implements approved changes on a dedicated branch

Every fix traces back to a specific insight, linked to a specific metric.


Architecture

your agent  ──>  ri.patch() + ri.session()  ──>  eval/traces/*.json
                                                        │
                                                        ▼
                                                  /recursive-improve
                                                        │
                                                        ▼
                                              improved agent code  ──>  repeat
                                                        │
                                                        ▼
                                                    benchmark  ──>  recursive-improve dashboard

                              ┌──────────────────────────────┐
                              │  /ratchet (autonomous loop)   │
                              │  improve → run → eval →       │
                              │  keep or revert → repeat      │
                              └──────────────────────────────┘
  • ri.patch(): monkey-patches OpenAI, Anthropic, and LiteLLM clients to capture every call
  • ri.session(): context manager that writes structured trace JSON files
  • /recursive-improve: Claude Code / Codex skill that analyzes traces and applies fixes
  • recursive-improve benchmark: snapshot metric quality, store, and compare over time
  • recursive-improve dashboard: web UI to visualize runs and compare branches
  • /ratchet: autonomous keep-or-revert loop that runs /recursive-improve repeatedly overnight

Star this repo if you find it useful!

Built with ❤️ by Kayba and the open-source community.

常见问题

What is recursive-improve?

recursive-improve is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by kayba-ai. 🪞 Make your agents recursively self-improve. It has 250 GitHub stars.

Is recursive-improve safe to use?

Yes. recursive-improve passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.

How do I install recursive-improve?

Clone the repository with "git clone https://github.com/kayba-ai/recursive-improve" and add it to your Claude Code skills directory (see the Installation section above).

What programming language is recursive-improve written in?

recursive-improve is primarily written in Python. It is open-source under kayba-ai on GitHub, so you can review or fork the full source.

Are there alternatives to recursive-improve?

Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh recursive-improve against similar tools.

评论 (0)

暂无评论,成为第一个分享想法的人!

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI 智能体ai-agentsbrainstorming
查看详情

hermes-agent

by NousResearch

10

The agent that grows with you

234,43747,175Python
AI 智能体ai-agentsagent-orchestration
查看详情

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI 智能体claude-codeai-tools
查看详情

claude-code

by anthropics

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

120,03119,897Shell
AI 智能体
查看详情

开发者还喜欢

基于喜欢此 Skill 的开发者投票和收藏

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI 智能体ai-agentsbrainstorming
查看详情

hermes-agent

by NousResearch

10

The agent that grows with you

234,43747,175Python
AI 智能体ai-agentsagent-orchestration
查看详情

n8n

by n8n-io

12

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

201,88160,308TypeScript
MCP 服务器apisai-tools
查看详情

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI 智能体claude-codeai-tools
查看详情