lumen

作者 ory已验证

Save 30% token costs when using Claude Code, Codex, OpenCode for free - with open source, local semantic search. Works for small and large codebases and monorepos! Enterprise-ready and fully compliant via Ollama and SQLite-vec.

251
Stars
31
Forks
Go
语言
2026/8/23
添加时间

⚠️ 第三方软件声明

本 Skill 为第三方开源软件,独立托管于 GitHub。SkillTip 仅为信息目录,不控制或维护底层仓库。所显示的安全检查为自动化且范围有限,安装前请自行审查源码。

阅读服务条款

安装

添加到你的 Claude Code skills 目录:

# Add to your Claude Code skills
git clone https://github.com/ory/lumen

快速入门

使用 lumen 等 Skills 的指南。

安全报告

已验证

上次扫描:—

{
  "status": "PASSED",
  "issues": []
}

README.md

Ory Lumen: Semantic code search for AI agents

CI Go Report Card Go Reference Coverage Status License

Claude reads entire files to find what it needs. Lumen gives it a map.

Lumen is a 100% local semantic code search engine for AI coding agents. No API keys, no cloud, no external database, just open-source embedding models (Ollama or LM Studio), SQLite, and your CPU. A single static binary and your own local embedding server.

The payoff is measurable and reproducible: across 9 benchmark runs on 9 languages and real GitHub bug-fix tasks, Lumen cuts cost in every single language — up to 39%. Output tokens drop by up to 66%, sessions complete up to 53% faster, and patch quality is maintained in every task. All verified with a transparent, open-source benchmark framework that you can run yourself.

With LumenBaseline (no Lumen)
Cost (avg, bug-fix)$0.29 (-26%)$0.40
Time (avg, bug-fix)125s (-28%)174s
Output tokens (avg)5,247 (-37%)8,323
JavaScript (marked)$0.32, 119s (-33%, -53%)$0.48, 255s
Rust (toml)$0.38, 204s (-39%, -34%)$0.61, 310s
PHP (monolog)$0.14, 34s (-27%, -34%)$0.19, 52s
TypeScript (commander)$0.14, 56s (-27%, -33%)$0.19, 84s
Svelte (chat-ui)$0.10, 56s (-26%, -31%)$0.14, 80s
Patch qualityMaintained in all 9 tasks

Table of contents

Demo

Lumen demo

Claude Code asking about the Prometheus codebase. Lumen's semantic_search finds the relevant code without reading entire files.

Quick start

Prerequisites:

Platform support: Linux, macOS, and Windows. File locking for background indexing coordination uses flock(2) on Unix and LockFileEx on Windows (via gofrs/flock).

  1. Ollama installed and running, then pull the default embedding model:
    ollama pull ordis/jina-embeddings-v2-base-code
    
  2. One of: Claude Code, Cursor, Codex, or OpenCode

Note: Installation differs by platform. Claude Code and Codex install from plugin marketplaces. OpenCode installs from npm. Cursor packaging is shipped in this repository and is ready for Cursor's plugin distribution workflow.

Install:

Claude Code

/plugin marketplace add ory/claude-plugins
/plugin install lumen@ory

Verify by starting a new Claude session and running /lumen:doctor.

Cursor

Lumen ships a native Cursor plugin bundle in this repository:

  • .cursor-plugin/plugin.json - plugin manifest
  • mcp.json - local lumen MCP server wiring
  • hooks/hooks-cursor.json - SessionStart hook
  • skills/ - shared doctor and reindex skills

Use Cursor's plugin installation or distribution workflow with this bundle. Detailed packaging notes: .cursor-plugin/INSTALL.md

Verify by opening a new Cursor agent session and asking it to use the doctor skill or the Lumen semantic_search tool.

Codex

Codex CLI 0.147.0 or newer installs Lumen as a native plugin:

codex plugin marketplace add ory/claude-plugins
codex plugin add lumen@ory

If the marketplace already exists, run codex plugin marketplace upgrade ory before installing. Legacy manual-clone and broken-plugin repair instructions: .codex/INSTALL.md.

Verify with:

codex mcp get lumen --json

OpenCode

Add @ory/lumen-opencode to the plugin array in your opencode.json:

{
  "plugin": ["@ory/lumen-opencode"]
}

Detailed docs: .opencode/INSTALL.md

Verify with:

opencode mcp list

Updating

  • Claude Code - update through Claude's plugin marketplace
  • Cursor - refresh or reinstall the bundled plugin through Cursor after updating this repository or the published package
  • Codex - upgrade the ory marketplace, reinstall lumen@ory, and restart
  • OpenCode - update the version pin in opencode.json (e.g. @ory/lumen-opencode@0.0.29) and restart OpenCode

On first Claude Code or Cursor session start, Lumen:

  1. Downloads the binary automatically from the latest GitHub release
  2. Indexes your project in the background using Merkle tree change detection
  3. Registers a semantic_search MCP tool that the host can use automatically

In Codex and OpenCode, the same binary download and index seeding happen on the first semantic_search call. Codex stores the downloaded binary in the plugin's writable data directory rather than the read-only package cache.

Two shared skills are also available: doctor (health check) and reindex (forced re-indexing). Claude exposes them as /lumen:doctor and /lumen:reindex; the other hosts discover the same shared skill content through their native skill systems.

The same semantic_search, health_check, and index_status MCP tools plus the shared doctor and reindex skills are exposed through the Codex, Cursor, and OpenCode surfaces as well. The first semantic_search call seeds or refreshes the index automatically.

What you get

  • Semantic vector search — Claude finds relevant functions, types, and modules by meaning, not keyword matching
  • Auto-indexing — indexes on session start, only re-processes changed files via Merkle tree diffing
  • Incremental updates — re-indexes only what changed; large codebases re-index in seconds after the first run
  • 12 language families — Go, Python, TypeScript, JavaScript, Svelte, Rust, Ruby, Java, PHP, C/C++, C#, Dart
  • Git worktree support — worktrees share index data automatically; a new worktree seeds from a sibling's index and only re-indexes changed files, turning minutes of embedding into seconds
  • Zero cloud — embeddings stay on your machine; no data leaves your network
  • Ollama and LM Studio — works with either local embedding backend

How it works

Lumen sits between your codebase and Claude as an MCP server. When a session starts, it walks your project and builds a Merkle tree over file hashes: only changed files get re-chunked and re-embedded. Each file is split into semantic chunks (functions, types, methods) using Go's native AST or tree-sitter grammars for other languages. Chunks are embedded and stored in SQLite + sqlite-vec using cosine-distance KNN for retrieval.

Files → semantic chunks → vector embeddings → SQLite/sqlite-vec → KNN search

When Claude needs to understand code, it calls semantic_search instead of reading entire files. The index is stored outside your repo (~/.local/share/lumen/<hash>/index.db). Git worktrees from the same repository use one collection for a compatible model, vector-storage, and chunking profile; non-Git projects use private collections. Different profiles never collide.

Benchmarks

Lumen is evaluated using bench-swe: a SWE-bench-style harness that runs Claude on real GitHub bug-fix tasks and measures cost, time, output tokens, and patch quality — with and without Lumen. All results are reproducible: raw JSONL streams, patch diffs, and judge ratings are committed to this repository.

Key results — 9 runs across 9 languages, hard difficulty, real GitHub issues (ordis/jina-embeddings-v2-base-code, Ollama):

LanguageCost ReductionTime ReductionOutput Token ReductionQuality
Rust-39%-34%-31% (18K → 12K)Poor (both)
JavaScript-33%-53%-66% (14K → 5K)Perfect (both)
TypeScript-27%-33%-64% (5K → 1.8K)Good (both)
PHP-27%-34%-59% (1.9K → 0.8K)Good (both)
Ruby-24%-11%-9% (6.1K → 5.6K)Good (both)
Python-20%-29%-36% (1.7K → 1.1K)Perfect (both)
Go-12%-9%-10% (11K → 10K)Good (both)
C++-8%-3%+42% (feature task)Good (both)
Svelte-26%-31%-26% (4.0K → 3.0K)Poor (both)

Cost was reduced in every language tested. Quality was maintained in every task — zero regressions. JavaScript and TypeScript show the most dramatic efficiency gains: same quality fixes in half the time with two-thirds fewer tokens. Even on tasks too hard for either approach (Rust, Svelte), Lumen cuts the cost of failure by 26–39%.

See docs/BENCHMARKS.md for all 9 per-language deep dives, judge rationales, and reproduce instructions.

Supported languages

Supports 12 language families with semantic chunking (10 benchmarked):

LanguageParserExtensionsBenchmark status
GoNative AST.goBenchmarked: -12% cost, Good quality
Pythontree-sitter.pyBenchmarked: Perfect quality, -36% tokens
TypeScript / TSXtree-sitter.ts, .tsxBenchmarked: -64% tokens, -33% time
JavaScript / JSXtree-sitter.js, .jsx, .mjsBenchmarked: -66% tokens, -53% time
Darttree-sitter.dartBenchmarked: -76% cost, -82% tokens, -79% time
Rusttree-sitter.rsBenchmarked: -39% cost, -34% time
Rubytree-sitter.rbBenchmarked: -24% cost, -11% time
PHPtree-sitter.phpBenchmarked: -59% tokens, -34% time
C / C++tree-sitter.c, .h, .cpp, .cc, .cxx, .hppBenchmarked: -8% cost (C++ feature task)
Sveltetree-sitter.svelteBenchmarked: -26% cost, -31% time
Javatree-sitter.javaSupported
C#tree-sitter.csSupported

Go uses the native Go AST parser for the most precise chunks. All other languages use tree-sitter grammars. See docs/BENCHMARKS.md for all 10 per-language benchmark deep dives.

Configuration

All configuration is via environment variables:

VariableDefaultDescription
LUMEN_EMBED_MODELsee note ¹Embedding model; use with LUMEN_EMBED_DIMS for unlisted models
LUMEN_BACKENDollamaEmbedding backend (ollama or lmstudio)
OLLAMA_HOSThttp://localhost:11434Ollama server URL
LM_STUDIO_HOSThttp://localhost:1234LM Studio server URL
LUMEN_MAX_CHUNK_TOKENS512Max tokens per chunk before splitting
LUMEN_VECTOR_STORAGEint8Vector precision (int8 or float32)
LUMEN_EMBED_DIMSOverride embedding dimensions (required for unlisted models)
LUMEN_EMBED_CTX8192 (unlisted models)Override context window length

¹ ordis/jina-embeddings-v2-base-code (Ollama), nomic-ai/nomic-embed-code-GGUF (LM Studio)

Supported embedding models

Dimensions and context length are configured automatically per model:

ModelBackendDimsContextRecommended
ordis/jina-embeddings-v2-base-codeOllama7688192Best default — lowest cost, no over-retrieval
qwen3-embedding:8bOllama409640960Best quality — strongest dominance (7/9 wins), very slow indexing
nomic-ai/nomic-embed-code-GGUFLM Studio35848192Usable — good quality, but TypeScript over-retrieval raises costs
qwen3-embedding:4bOllama256040960Not recommended — highest costs, severe TypeScript over-retrieval
nomic-embed-textOllama7688192Untested
qwen3-embedding:0.6bOllama102432768Untested
all-minilmOllama384512Untested

Switching models creates a separate index automatically. The model name is part of the database path hash, so different models never collide.

Caveat: the DB path hash includes the model name but not the backend. If the same model name is configured on two backends (e.g. an Ollama and an LM Studio entry both named foo), they share the same index — use distinct model names per backend to avoid collisions.

Selecting a server per invocation

lumen index and lumen search accept --model/-m and --backend/-b to pick from a multi-server config.yaml. The selection filters the configured servers to those matching both fields; failover still works within the filtered subset.

# Index with the Ollama server matching this model name.
lumen index --model ordis/jina-embeddings-v2-base-code .

# Same model name hosted on LM Studio (present in YAML, not in the
# static registry) — accepted because the name is configured.
lumen index --model text-embedding-jina-embeddings-v2-base-code .

# Disambiguate when the same model is configured on two backends.
lumen index --model my-embed --backend lmstudio .

# Pick the first configured Ollama server regardless of model.
lumen search --backend ollama "…"

If --model is not configured in YAML but is a known registry model (and --backend is unset), Lumen falls back to mutating the default server's model — preserving lumen index --model all-minilm . for users with no YAML.

Using a custom or unlisted model

If your model is not in the registry above, set LUMEN_EMBED_DIMS to bypass the registry check. LUMEN_EMBED_CTX is optional and defaults to 8192.

Both variables can also override values for known models — useful when running a model variant with a longer context window or different output dimensions.

LUMEN_BACKEND=lmstudio
LM_STUDIO_HOST=http://localhost:8801
LUMEN_EMBED_MODEL=mlx-community/Qwen3-Embedding-8B-4bit-DWQ
LUMEN_EMBED_DIMS=4096
LUMEN_EMBED_CTX=40960   # optional, defaults to 8192

Controlling what gets indexed

Lumen filters files through six layers: built-in directory and lock file skips → .gitignore.lumenignore.gitattributes (linguist-generated) → supported file extension. Only files that pass all layers are indexed.

.lumenignore uses .gitignore syntax. Place it in your project root (or any subdirectory) to exclude files that aren't in .gitignore but are noise for code search — generated protobuf files, test snapshots, vendored data, etc.

Built-in skips (always excluded)

Directories: .git, node_modules, vendor, dist, .cache, .venv, venv, __pycache__, target, .gradle, _build, deps, .idea, .vscode, .next, .nuxt, .build, .output, bower_components, .bundle, .tox, .eggs, testdata, .hg, .svn

Lock files: package-lock.json, yarn.lock, pnpm-lock.yaml, bun.lock, bun.lockb, go.sum, composer.lock, poetry.lock, Pipfile.lock, Gemfile.lock, Cargo.lock, pubspec.lock, mix.lock, flake.lock, packages.lock.json

Database location

Index databases are stored outside your project:

~/.local/share/lumen/<hash>/index.db

Where <hash> identifies the Git common directory (or the absolute path for a non-Git project), indexed scope, embedding model and dimensions, vector precision, chunking profile, and index version. Worktrees in one repository share content-addressed file revisions and vectors while retaining independent project memberships. Vectors use int8 storage by default; set LUMEN_VECTOR_STORAGE=float32 to opt out. No files are added to your repo.

You can safely delete the entire lumen directory to clear all indexes, or let Lumen reclaim the space for you:

lumen clean            # remove indexes unused for 30 days or whose project is gone
lumen clean --days 7   # tighten the cutoff to a week
lumen clean --days 0   # remove every eligible index except actively locked indexes

An index counts as used every time Lumen opens it (search, indexing, status, or session start), so indexes for projects you still work on are never removed. Indexes with an indexer currently running are always kept.

Git worktrees are detected automatically. A new worktree attaches unchanged path-and-content revisions directly from the repository collection and embeds only missing chunk inputs. Removing an old worktree drops its memberships; shared revisions and vectors remain until their final reference disappears. Legacy per-worktree indexes migrate lazily, reusing unchanged float32 vectors without contacting the embedding backend.

For the complete storage key, sharing rules, status metrics, migration process, and cleanup lifecycle, see Index storage and lifecycle.

CLI Reference

Download the binary from the GitHub releases page or let the plugin install it automatically.

lumen help

Troubleshooting

Ollama not running / "connection refused"

Start Ollama and verify the model is pulled:

ollama serve
ollama pull ordis/jina-embeddings-v2-base-code

Run /lumen:doctor inside Claude Code to confirm connectivity.

In Cursor, Codex, or OpenCode, use the shared doctor skill or call health_check and index_status directly.

Stale index after large refactor

Run /lumen:reindex inside Claude Code to force a full re-index, or:

lumen index --force .

In Codex, use the bundled reindex skill to refresh the index through the MCP server, or run the same CLI commands for a clean rebuild. The same shared reindex skill is available in Cursor and OpenCode as well.

LM Studio: embedding model appears under LLMs instead of Embeddings

LM Studio classifies embedding models by matching the GGUF arch field against a hardcoded allowlist (bert, nomic-bert). Models built on other architectures — including Qwen2-based models like nomic-embed-code — are misclassified as LLMs. This affects lms ls output and the /v1/embeddings REST endpoint.

Fix (GGUF, v0.3.16+): Open LM Studio → My Models, click the gear icon next to the model, set Override Domain TypeText Embedding.

macOS / Apple Silicon: MLX format models are significantly faster on Apple Silicon. However, LM Studio removed the domain type override for MLX in v0.3.30+, so MLX embedding models cannot be reclassified. Use the GGUF variant to retain the override option, or switch to Ollama (ordis/jina-embeddings-v2-base-code or qwen3-embedding:8b).

Switching embedding models

Set LUMEN_EMBED_MODEL to a model from the supported table above. Each model gets its own database; the old index is not deleted automatically.

Changing LUMEN_VECTOR_STORAGE, LUMEN_EMBED_DIMS, or LUMEN_MAX_CHUNK_TOKENS also selects a separate collection. Run lumen clean after the old profile is no longer in use if you want to reclaim its disk space.

Understanding index size and deduplication

Call index_status for the project. It reports project-local file and chunk counts alongside collection-wide unique vectors, shared references, deduplication ratio, vector precision, database size, and currently reclaimable SQLite pages. See Index storage and lifecycle for definitions and examples.

Slow first indexing

The first run embeds every file. Subsequent runs only process changed files (typically a few seconds). For large projects (100k+ lines), first indexing can take several minutes — this is a one-time cost.

Development

git clone https://github.com/ory/lumen.git
cd lumen

# Build locally (CGO required for sqlite-vec)
make build-local

# Run tests
make test

# Run linter
make lint

# Load as a Claude Code plugin from source
make plugin-dev

See CLAUDE.md for architecture details, design decisions, and contribution guidelines, and AGENTS.md for repo-specific agent instructions.

常见问题

What is lumen?

lumen is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by ory. Save 30% token costs when using Claude Code, Codex, OpenCode for free - with open source, local semantic search. Works for small and large codebases and monorepos! Enterprise-ready and fully compliant via Ollama and SQLite-vec. It has 251 GitHub stars.

Is lumen safe to use?

Yes. lumen passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.

How do I install lumen?

Clone the repository with "git clone https://github.com/ory/lumen" and add it to your Claude Code skills directory (see the Installation section above).

What programming language is lumen written in?

lumen is primarily written in Go. It is open-source under ory on GitHub, so you can review or fork the full source.

Are there alternatives to lumen?

Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh lumen against similar tools.

评论 (0)

暂无评论,成为第一个分享想法的人!

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI 智能体ai-agentsbrainstorming
查看详情

hermes-agent

by NousResearch

10

The agent that grows with you

234,43747,175Python
AI 智能体ai-agentsagent-orchestration
查看详情

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI 智能体claude-codeai-tools
查看详情

claude-code

by anthropics

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

120,03119,897Shell
AI 智能体
查看详情

开发者还喜欢

基于喜欢此 Skill 的开发者投票和收藏

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI 智能体ai-agentsbrainstorming
查看详情

hermes-agent

by NousResearch

10

The agent that grows with you

234,43747,175Python
AI 智能体ai-agentsagent-orchestration
查看详情

n8n

by n8n-io

12

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

201,88160,308TypeScript
MCP 服务器apisai-tools
查看详情

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI 智能体claude-codeai-tools
查看详情