rocketplaneIO

作者 olemeyer已验证

Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable.

157
Stars
4
Forks
Go
语言
2026/8/23
添加时间

⚠️ 第三方软件声明

本 Skill 为第三方开源软件,独立托管于 GitHub。SkillTip 仅为信息目录,不控制或维护底层仓库。所显示的安全检查为自动化且范围有限,安装前请自行审查源码。

阅读服务条款

安装

添加到你的 Claude Code skills 目录:

# Add to your Claude Code skills
git clone https://github.com/olemeyer/rocketplaneIO

快速入门

使用 rocketplaneIO 等 Skills 的指南。

安全报告

已验证

上次扫描:—

{
  "status": "PASSED",
  "issues": []
}

README.md

Full-stack Kubernetes observability —
that can safely fix things, too.

Self-hosted · eBPF traces, logs & metrics, zero instrumentation · a safe MCP interface for AI agents · Apache-2.0 · air-gapped capable

Status License Go

Quick start · Is it safe? · Under the hood


rocketplaneIO walkthrough: connect a cluster and the live service map draws itself in, then an external AI agent opens a transaction over MCP, acts on the cluster, a human approves the hard part — and one click rolls everything back

Alpha. The full loop works end-to-end today, developed against minikube. APIs and schemas still change without notice — don't point it at production yet.

Three things you don't get anywhere else

1 · Traces for services you never instrumented

One outbound-only agent plus an eBPF DaemonSet (OTel eBPF Instrumentation, née Grafana Beyla). HTTP/gRPC spans with cross-service context propagation — including compiled Go binaries — plus SQL, Redis and Kafka client spans. No SDKs, no sidecars, no code changes.

A real 500 investigated: the failure cascades across three services, correlated error logs on the right
A real 500 on GET /checkout: the failure cascades frontdoor → checkout → catalog, the exact ERROR log lines are correlated on the right. No SDK in any of these services.

2 · Logs, traces, metrics and topology — one dataset, not four tools

It's not four products bolted together. A log line is two clicks from its distributed trace. The service map is drawn from real eBPF traffic, not a static config — tech logos matched from container images, health live from RED metrics. PromQL runs on the actual Prometheus evaluation engine, embedded, over ClickHouse. And the complete Kubernetes inventory — Services, Ingress, ConfigMaps, policies, volumes — is synced and searchable right beside it. Everything self-hosted; your telemetry never leaves your infrastructure.

Live service map with automatic technology detection, drawn from real eBPF traffic
Service map — topology from real traffic flows, not config. Healthy is calm grey; only anomalies light up. The whole UI follows one instrument-panel design system (RETICLE).

3 · And it can fix what it finds — safely

Observability that stops at "here's the problem" leaves the hard part to you. rocketplaneIO doesn't ship its own AI — it's the MCP interface your existing agent (Claude Code, Cursor, anything that speaks MCP) connects to. The agent reads everything above, and acts through kubectl-shaped primitives (k8s_get / k8s_patch / k8s_apply / k8s_delete / k8s_exec / Starlark workflows) on any resource, CRDs included.

The guardrail is a transaction. Nothing mutates outside one — every change is snapshotted durably before it commits, so cancel, timeout or a later Revert restores every before-state in reverse (LIFO). Each operation is risk-classified by rocketplaneIO, not by the model:

LevelExamplesDefault policy
◎ readk8s_get, k8s_list, logs, traces, PromQLruns immediately, no transaction
↺ reversiblek8s_patch, k8s_apply (snapshot-backed)runs inside the transaction
◇ disruptivek8s_delete on Pods/Jobsa human approves in the UI
△ destructiveany other delete, k8s_exec, workflows, scale-to-0a human approves in the UI

Classification is parameter-aware (replicas: 3 is reversible; : 0 is destructive) and fail-closed. Disruptive and destructive operations park until a human approves them — and an API token can never approve its own proposals.

A transaction timeline: every tool call, the parked approval, the human decision, and the rollback — one auditable spine
Transactions — the audit trail. What the agent read, what it changed, who approved what, and ↺ revert at run- or transaction-level. Cancel always terminates — a reaper finalizes anything a dead agent leaves behind and drives the rollback to its end.

Quick start

You need Docker and a Kubernetes cluster to point it at (minikube is fine).

1 — run the platform (one command)

curl -fsSL https://rocketplane.io/install.sh | sh

That's the whole install: it downloads the compose bundle, generates real secrets, pulls the published images and starts everything. The UI comes up on http://localhost:4173 (create the owner account there), the control plane on :8090. Re-running the same command later is the update path — secrets are kept, images are bumped to the newest release.

Prefer to see every step? The manual compose variant.
curl -O https://raw.githubusercontent.com/olemeyer/rocketplaneIO/main/deploy/compose/docker-compose.prod.yml
curl -o .env https://raw.githubusercontent.com/olemeyer/rocketplaneIO/main/deploy/compose/.env.example

# REQUIRED: a real session secret — the control plane refuses to start without one.
echo "RP_SESSION_SECRET=$(openssl rand -hex 32)" >> .env

docker compose --env-file .env -f docker-compose.prod.yml up -d

Production mode by default (no anonymous login). For a throwaway localhost trial add RP_ENV=dev to .env to skip account setup.

2 — connect your cluster

Open the UI, hit Connect cluster — it hands you one copy-paste command that installs the agent and the Beyla DaemonSet (a rendered kubectl apply, or Helm). When the service map draws your namespaces and spans appear under Traces, you're live — without touching a line of your code. The installer auto-detects a LAN address your cluster can dial back to; for a local minikube that is http://host.minikube.internal:8090.

3 — connect your AI agent

Create an API token under Settings → API keys (role admin for mutations), then:

claude mcp add --transport http rocketplaneio \
  http://localhost:8090/mcp/orgs/<org>/clusters/<cluster> \
  --header "Authorization: Bearer rp_..."

The Settings → MCP tab generates this snippet (and the .mcp.json / Cursor variants) with your real IDs filled in. From that moment your agent can investigate freely — and every change it wants to make runs through a transaction you can watch, gate and roll back live under Transactions.

Images ship as pinned releases (see RP_VERSION in .env); edge tracks main for the latest unreleased changes. A platform Helm chart is the next milestone. Want a demo workload? A Python + Redis shop behind nginx — the one in every screenshot here — ships in deploy/dev/ (kubectl apply -f deploy/dev/shop-realistic.yaml -f deploy/dev/frontdoor.yaml).

Is it safe?

The section every platform team reads first:

  • Outbound-only. The agent dials out over HTTPS; nothing connects into your cluster, nothing listens. Actions are claimed by the agent — never pushed in.
  • Auth. In production the dev-login bypass is off: the first account is created at /setup, or you wire up Google SSO. Sessions are HMAC-signed with RP_SESSION_SECRET, and the control plane refuses to start with a missing or weak one — no forgeable cookies, no shipped default.
  • Network / TLS. The control plane speaks plain HTTP on :8090 — put it behind a TLS reverse proxy (or the platform Helm chart's ingress) for anything past localhost. The OTLP ingest ports 4317/4318 are unauthenticated by design; keep them on a trusted network (in-cluster or behind the proxy), not on the public internet.
  • RBAC, two blocks (deploy/install.yaml): observe is enumerated read-only; act is deliberately broad — the generic operation set works on any kind, CRDs included. The guardrails are architectural, not a resource list: no mutation outside a transaction, risk classification, human approval for hard operations, durable pre-mutation snapshots, LIFO rollback, full audit. Delete the act block (or set rbac.actions=false in Helm) for a strictly observe-only agent.
  • Secrets: the inventory shows names, types and key counts — never values. That restraint is enforced in agent code; if you'd rather enforce it with RBAC too, delete the one secrets rule and the agent degrades gracefully.
  • The agent is fenced, not trusted. It authenticates with an org-scoped API token against the same RBAC you use; mutations demand an admin token AND an open transaction; disruptive and destructive operations park until a human approves them in the UI — and a token can never approve its own proposals. Every tool call, decision and result lands in the transaction timeline.
  • eBPF: Beyla runs as a privileged DaemonSet (kernel ≥ 5.8 with BTF). The capture layer is upstream OpenTelemetry tooling — rocketplaneIO is the investigation-and-action loop on top.
  • Footprint (measured on single-node minikube under continuous synthetic load — indicative, not a production benchmark): the rocketplaneIO agent is one lightweight pod at ~2m CPU and ~10Mi memory, cluster-wide. Beyla is the real cost and scales with request volume — ~0.02 core idle, spiking to ~0.25 core under load, ~280Mi memory per node (bounded; its default limit is 512Mi). One agent for the cluster, one Beyla per node.

What else is in the box

Alerts with auto-remediation · PromQL on ClickHouse · full K8s inventory · approval gate · logs · nodes
Alert rule firing and dispatching an auto-remediation workflow Alerts — typed checks or PromQL conditions with for-durations and sparklines. A firing rule can dispatch a remediation workflow: once per transition, fully audited, still subject to verify-or-rollback.



PromQL editor on ClickHouse PromQL — the actual Prometheus evaluation engine, embedded and pointed at ClickHouse (internal/promqlx); editor built on codemirror-promql. Custom metrics are named PromQL expressions, validated at save.



Full Kubernetes inventory, tabbed by resource kind Resources — the complete cluster inventory (Services, Ingress, ConfigMaps, batch, policies, volumes, quotas), synced every 60s. The same data an agent reads via k8s_list.



A parked destructive operation waiting for human approval inside a transaction Approvals — a destructive operation proposed by an external agent parks as awaiting approval; the agent waits, you decide. Approve releases it to the cluster, reject cancels it — either way it's on the timeline forever.



Log stream with severity histogram Logs — severity histogram, brushing, and the two-click path to the distributed trace.



Node infrastructure with kubelet stats Nodes — kubelet-level stats for every node: CPU, memory, disk pressure and schedulability at a glance.



The UI follows a strict instrument-panel design system (RETICLE) — healthy is calm, only anomalies speak.

Under the hood

 YOUR CLUSTER   agent (Go, outbound-only, enumerated RBAC)      beyla eBPF DaemonSet
                topology · logs · inventory · action pipelines   HTTP/gRPC/Go · SQL/Redis/Kafka
                        │ HTTPS                                          │ OTLP
                        ▼                                                ▼
 CONTROL PLANE  control plane (Go, single binary)                OTel collector → ClickHouse
                MCP endpoint (Streamable HTTP) · transactions + approval gate
                API · auth/orgs · alert evaluator · operation queue · rollback reaper
                embedded Prometheus engine (PromQL → ClickHouse) · SSE broker
                        │
 DATA           PostgreSQL (state, rules, transactions, runs)  ·  ClickHouse (logs, traces, metrics)
 ACCESS         web (Next.js) — service map · transactions · logs · traces · metrics · alerts · incidents

The MCP endpoint lives in the control plane (/mcp/orgs/{org}/clusters/{cluster}, Streamable HTTP, stateless — no sticky sessions across replicas). Read tools cover logs, traces (incl. single-trace waterfalls and trace↔log correlation), PromQL, the service map and any Kubernetes object; the mutating tools are the transaction-gated primitives above. Auth is a bearer API token; every tool call reuses the exact HTTP handlers, validation and RBAC the browser uses.

Monorepo layout
├── agent/                      # in-cluster agent (Go): sync, logs, inventory, snapshot substrate + generic restore
├── services/controlplane/      # control plane (Go): API, auth, alerts, MCP endpoint, PromQL engine
│   └── internal/
│       ├── api/                #   REST + SSE + mcp_* (Streamable HTTP, tools, transactions, approvals)
│       ├── promqlx/            #   embedded Prometheus engine on ClickHouse
│       └── alerts/ · telemetry/ · events/ · store/ · migrations/
├── apps/web/                   # Next.js UI (RETICLE design system)
└── deploy/                     # compose (platform + dev stores), helm chart, install.yaml, demo shop
Develop from source

Prereqs: Go 1.25+, Node 22+ with pnpm, Docker.

git clone https://github.com/olemeyer/rocketplaneIO && cd rocketplaneIO

docker compose -f deploy/compose/docker-compose.yml up -d          # dev data stores + collector
go run ./services/controlplane/cmd/controlplane                    # control plane on :8090
cd apps/web && pnpm install && pnpm dev                            # UI on :4173

Building the platform images yourself (what CI publishes):

docker build -t rocketplaneio/controlplane -f services/controlplane/Dockerfile .
docker build -t rocketplaneio/web          -f apps/web/Dockerfile .
docker build -t rocketplaneio/agent        -f agent/Dockerfile .

Status

Works end-to-end todayOn the roadmap
eBPF traces incl. compiled Go, log→trace correlationTagged releases + platform Helm chart
Live service map, PromQL + custom metrics on ClickHouseHosted demo
Full K8s inventory, alerts with auto-remediationProduction-scale overhead benchmarks
MCP endpoint: transactions, classification, approval gateMulti-user RBAC hardening
LIFO snapshot rollback on cancel/expiry — CRDs includedNamespace-scoped transaction policies
Transaction timeline, revert at run and transaction level
Published multi-arch images (edge) + agent install.yaml

Contributing

Reports from real clusters are the contribution we want most — open an issue. See CONTRIBUTING.md. If this resonates, a star helps others find it.

License

Apache-2.0.


Built with Go, Next.js, eBPF and ClickHouse · rocketplaneIO

常见问题

What is rocketplaneIO?

rocketplaneIO is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by olemeyer. Self-hosted AI SRE for Kubernetes — zero-instrumentation eBPF observability plus a copilot that fixes issues through guardrailed, self-verifying actions. BYO-LLM, air-gapped capable. It has 157 GitHub stars.

Is rocketplaneIO safe to use?

Yes. rocketplaneIO passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.

How do I install rocketplaneIO?

Clone the repository with "git clone https://github.com/olemeyer/rocketplaneIO" and add it to your Claude Code skills directory (see the Installation section above).

What programming language is rocketplaneIO written in?

rocketplaneIO is primarily written in Go. It is open-source under olemeyer on GitHub, so you can review or fork the full source.

Are there alternatives to rocketplaneIO?

Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh rocketplaneIO against similar tools.

评论 (0)

暂无评论,成为第一个分享想法的人!

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI 智能体ai-agentsbrainstorming
查看详情

hermes-agent

by NousResearch

10

The agent that grows with you

234,43747,175Python
AI 智能体ai-agentsagent-orchestration
查看详情

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI 智能体claude-codeai-tools
查看详情

claude-code

by anthropics

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

120,03119,897Shell
AI 智能体
查看详情

开发者还喜欢

基于喜欢此 Skill 的开发者投票和收藏

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI 智能体ai-agentsbrainstorming
查看详情

hermes-agent

by NousResearch

10

The agent that grows with you

234,43747,175Python
AI 智能体ai-agentsagent-orchestration
查看详情

n8n

by n8n-io

12

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

201,88160,308TypeScript
MCP 服务器apisai-tools
查看详情

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI 智能体claude-codeai-tools
查看详情