claude-video-vision

作者 jordanrendric已验证

Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis

1,255
Stars
153
Forks
TypeScript
语言
2026/8/23
添加时间

⚠️ 第三方软件声明

本 Skill 为第三方开源软件,独立托管于 GitHub。SkillTip 仅为信息目录,不控制或维护底层仓库。所显示的安全检查为自动化且范围有限,安装前请自行审查源码。

阅读服务条款

安装

添加到你的 Claude Code skills 目录:

# Add to your Claude Code skills
git clone https://github.com/jordanrendric/claude-video-vision

快速入门

使用 claude-video-vision 等 Skills 的指南。

安全报告

已验证

上次扫描:—

{
  "status": "PASSED",
  "issues": []
}

README.md

Claude Code Video Vision

Give Claude the ability to watch and understand videos.

A Claude Code plugin that extracts frames via ffmpeg and processes audio via multiple backends (Gemini API, local Whisper, or OpenAI API). Claude receives frames as images and audio transcription with timestamps — the plugin is a perception layer, not an interpretation layer.

Features

  • Multimodal perception — Claude sees video frames directly and reads audio transcriptions with timestamps

  • YouTube URL support — Pass a YouTube URL directly; the MCP server downloads it with yt-dlp, preserving source metadata and captions for context

  • Flexible backends — Choose between cloud APIs or fully local processing

  • Adaptive extraction — Claude adjusts fps, time range, and resolution based on your question

  • Auto-installation — Whisper models download automatically on first use

  • Interactive setup wizard/setup-video-vision walks you through configuration

Quick Start

1. Install the plugin

Inside Claude Code, run these commands one at a time:

/plugin marketplace add https://github.com/jordanrendric/claude-video-vision

Then:

/plugin install claude-video-vision

The MCP server will auto-install via npx from npm on first use — no build step required.

Alternative: local development

git clone https://github.com/jordanrendric/claude-video-vision.git
claude --plugin-dir /path/to/claude-video-vision

2. Configure

Inside Claude Code, run the interactive wizard:

/claude-video-vision:setup-video-vision

It will walk you through backend selection, whisper configuration (if local), frame options, and dependency verification.

Usage

Slash command

/watch-video path/to/video.mp4
/watch-video tutorial.mp4 "what language is used in this tutorial?"
/watch-video https://www.youtube.com/watch?v=... "summarize this video"

Conversational

Just mention a video file or YouTube URL — Claude will detect it:

"analyze this video for me: ~/Downloads/demo.mp4"

"take a look at the first second of ~/videos/bug-report.mov"

"summarize this YouTube Short: https://www.youtube.com/shorts/..."

Claude adapts parameters automatically:

  • "the first second" → extracts at original fps from 00:00:00 to 00:00:01

  • "summarize this 1h lecture" → low fps, full duration

  • "what text is on screen at 1:30?" → high resolution, narrow time window

Backends

Backend Audio processing Cost Setup

Gemini API Native (speech + non-speech events) Free tier: 1500 req/day GEMINI_API_KEY env var

Local (Whisper) whisper.cpp or Python openai-whisper Free, fully offline brew install whisper-cpp + auto model download

OpenAI API OpenAI Whisper API Paid per usage OPENAI_API_KEY env var

All backends extract video frames via ffmpeg — Claude always has direct visual access.

Architecture

┌───────────────────────────────────────────────────────┐
│ Claude Code (your session)                            │
│                                                       │
│  /watch-video  ──→  Skill: video-perception          │
│                        │                              │
│                        ▼                              │
│                  MCP tool: video_watch                │
│                        │                              │
└────────────────────────┼──────────────────────────────┘
                         │
                         ▼
      ┌────────────────────────────────────┐
      │ MCP Server (Node.js)               │
      │                                    │
      │  ┌──────────┐    ┌──────────────┐  │
      │  │ ffmpeg   │    │ Audio backend│  │
      │  │ frames   │ ║  │ (parallel)   │  │
      │  └──────────┘    └──────────────┘  │
      │       │                 │          │
      └───────┼─────────────────┼──────────┘
              ▼                 ▼
        base64 images     transcription
        + timestamps      + audio events
              │                 │
              └────────┬────────┘
                       ▼
              Claude receives both

Requirements

  • Node.js 20+ (for the MCP server)

  • ffmpeg (auto-detected, install instructions provided by setup wizard)

  • yt-dlp (optional, required only for YouTube URLs; brew install yt-dlp on macOS)

  • Backend-specific:

  • Gemini API: free API key from ai.google.dev

  • Local: brew install whisper-cpp (macOS) or equivalent

  • OpenAI: API key from OpenAI

MCP Tools

The plugin exposes 6 MCP tools:

  • video_watch — Extract frames + process audio (main tool)

  • video_analyze — Analyze video structure with ffmpeg filters before extraction

  • video_detail — Drill into specific cached or newly extracted moments

  • video_info — Get video metadata without processing

  • video_configure — Change settings

  • video_setup — Check and guide dependency installation

Slash Commands

  • /watch-video <path> [question] — Analyze a video

  • /setup-video-vision — Interactive configuration wizard

Configuration

Settings are stored in ~/.claude-video-vision/config.json:

{
  "backend": "local",
  "whisper_engine": "cpp",
  "whisper_model": "auto",
  "whisper_at": false,
  "frame_mode": "images",
  "frame_format": "jpeg",
  "frame_resolution": 512,
  "default_fps": "auto",
  "max_frames": 100,
  "frame_describer_model": "sonnet",
  "enable_index": false,
  "session_max_age_days": 7,
  "downloads_max_age_days": 7
}

frame_format can be jpeg, png, or webp. jpeg remains the default for backwards compatibility; png is useful for screen recordings where text and sharp UI edges should stay lossless.

Whisper models auto-download to ~/.claude-video-vision/models/ on first use. Available: tiny, base, small, medium, large-v3-turbo, large-v3, auto (picks best for your RAM).

YouTube Transcripts

For YouTube URLs, the server uses this transcript order:

  • Manual YouTube subtitles when an English track is available.

  • YouTube automatic captions when manual subtitles are not available.

  • The configured audio backend when captions are missing, empty, or cover too little of a longer video.

Audio results label provenance with transcription_source, for example youtube_subtitles or youtube_auto_captions, so Claude can treat manual subtitles as stronger evidence than auto-captions.

Status

v1.0.0 — Initial release. Tested on macOS (Apple Silicon) with Local backend (whisper.cpp).

License

MIT — see LICENSE.

Author

Jordan Vasconcelos

Star History

常见问题

What is claude-video-vision?

claude-video-vision is an open-source mcp servers skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by jordanrendric. Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis. It has 1,255 GitHub stars.

Is claude-video-vision safe to use?

Yes. claude-video-vision passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.

How do I install claude-video-vision?

Clone the repository with "git clone https://github.com/jordanrendric/claude-video-vision" and add it to your Claude Code skills directory (see the Installation section above).

What programming language is claude-video-vision written in?

claude-video-vision is primarily written in TypeScript. It is open-source under jordanrendric on GitHub, so you can review or fork the full source.

Are there alternatives to claude-video-vision?

Yes. SkillsLLM lists many other MCP Servers skills you can browse and compare side by side. Open the MCP Servers category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh claude-video-vision against similar tools.

评论 (0)

暂无评论,成为第一个分享想法的人!

n8n

by n8n-io

12

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

201,88160,308TypeScript
MCP 服务器apisai-tools
查看详情

Scrapling

by D4Vinci

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

75,9137,581Python
MCP 服务器
查看详情

TrendRadar

by sansan0

⭐AI-driven public opinion & trend monitor with multi-platform aggregation, RSS, and smart alerts.🎯 告别信息过载,你的 AI 舆情监控助手与热点筛选工具!聚合多平台热点 + RSS 订阅,支持关键词精准筛选。AI 智能筛选新闻 + AI 翻译 + AI 分析简报直推手机,也支持接入 MCP 架构,赋能 AI 自然语言对话分析、情感洞察与趋势预测等。支持 Docker ,数据本地/云端自持。集成微信/飞书/钉钉/Telegram/邮件/ntfy/bark/slack 等渠道智能推送。

61,65224,883Python
MCP 服务器
查看详情

context7

by upstash

Context7 Platform -- Up-to-date code documentation for LLMs and AI code editors

61,0602,938TypeScript
MCP 服务器
查看详情

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

39,9393,219C
MCP 服务器
查看详情

开发者还喜欢

基于喜欢此 Skill 的开发者投票和收藏

ECC

by affaan-m

10

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

242,21936,702JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情
15

An agentic skills framework & software development methodology that works.

234,96620,863Shell
AI 智能体ai-agentsbrainstorming
查看详情

hermes-agent

by NousResearch

10

The agent that grows with you

234,43747,175Python
AI 智能体ai-agentsagent-orchestration
查看详情

n8n

by n8n-io

12

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

201,88160,308TypeScript
MCP 服务器apisai-tools
查看详情

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

185,94028,768JavaScript
AI 智能体ai-agentsanthropicclaude-code
查看详情

cc-switch

by farion1231

3

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

128,8688,826Rust
AI 智能体claude-codeai-tools
查看详情
claude-video-vision — Claude Code AI Skill | SkillTip