Skip to content
 
 

Repository files navigation

claude-video-vision

Codex Video Vision

Give Codex the ability to watch and understand videos.

A Codex-ready fork of claude-video-vision. It extracts frames via ffmpeg and processes audio via multiple backends (Gemini API, local Whisper, or OpenAI API). Codex receives frames as images plus audio transcription with timestamps — the plugin is a perception layer, not an interpretation layer.

Features

  • Multimodal perception — Codex sees video frames directly and reads audio transcriptions with timestamps
  • YouTube URL support — pass a YouTube URL and the MCP server downloads it with yt-dlp, preserving source metadata and captions for context
  • Flexible backends — Choose between cloud APIs or fully local processing
  • Adaptive extraction — Codex adjusts fps, time range, and resolution based on your question
  • Auto-installation — Whisper models download automatically on first use
  • Codex skill workflow — video-perception guides frame extraction, analysis, and drill-down

Quick Start

Codex local setup

Clone and build the fork:

git clone https://github.com/miketuckman/codex-video-vision.git
cd codex-video-vision/mcp-server
npm ci
npm run build
npm test

Add the MCP server to ~/.codex/config.toml:

[mcp_servers.codex-video-vision]
command = "node"
args = ["/absolute/path/to/codex-video-vision/mcp-server/dist/index.js"]
cwd = "/absolute/path/to/codex-video-vision/mcp-server"
startup_timeout_sec = 30
tool_timeout_sec = 600

Then copy or install the skill from skills/video-perception/ into your Codex skills path, or install this repo as a local Codex plugin once plugin support is available in your environment.

See docs/codex.md for full Codex setup, verification, and recommended defaults.

npm fallback

Until this fork publishes its own npm package, you can run the upstream MCP server directly:

npx -y claude-video-vision@latest

Use this Codex config when you prefer the npm package over a local build:

[mcp_servers.codex-video-vision]
command = "npx"
args = ["-y", "claude-video-vision@latest"]
startup_timeout_sec = 30
tool_timeout_sec = 600

Configure

Use the video_configure MCP tool from Codex. Recommended first-pass local settings:

{
  "backend": "local",
  "whisper_engine": "cpp",
  "whisper_model": "auto",
  "whisper_at": false,
  "frame_mode": "images",
  "frame_resolution": 512,
  "default_fps": "auto",
  "max_frames": 100,
  "enable_index": false,
  "session_max_age_days": 7
}

Usage

Conversational

Just mention a video file — the Codex skill will guide the workflow:

"analyze this video for me: ~/Downloads/demo.mp4"

"take a look at the first second of ~/videos/bug-report.mov"

"analyze and summarize this YouTube video: https://www.youtube.com/watch?v=..."

Codex adapts parameters automatically:

  • "the first second" → extracts at original fps from 00:00:00 to 00:00:01
  • "summarize this 1h lecture" → low fps, full duration
  • "what text is on screen at 1:30?" → high resolution, narrow time window

Backends

Backend Audio processing Cost Setup
Gemini API Native (speech + non-speech events) Free tier: 1500 req/day GEMINI_API_KEY env var
Local (Whisper) whisper.cpp or Python openai-whisper Free, fully offline brew install whisper-cpp + auto model download
OpenAI API OpenAI Whisper API Paid per usage OPENAI_API_KEY env var

All backends extract video frames via ffmpeg — Codex always has direct visual access.

Architecture

┌───────────────────────────────────────────────────────┐
│ Codex (your session)                                  │
│                                                       │
│  video request ──→  Skill: video-perception          │
│                        │                              │
│                        ▼                              │
│                  MCP tool: video_watch                │
│                        │                              │
└────────────────────────┼──────────────────────────────┘
                         │
                         ▼
      ┌────────────────────────────────────┐
      │ MCP Server (Node.js)               │
      │                                    │
      │  ┌──────────┐    ┌──────────────┐  │
      │  │ ffmpeg   │    │ Audio backend│  │
      │  │ frames   │ ║  │ (parallel)   │  │
      │  └──────────┘    └──────────────┘  │
      │       │                 │          │
      └───────┼─────────────────┼──────────┘
              ▼                 ▼
        base64 images     transcription
        + timestamps      + audio events
              │                 │
              └────────┬────────┘
                       ▼
              Codex receives both

Requirements

  • Node.js 20+ (for the MCP server)
  • ffmpeg (auto-detected, install instructions provided by setup wizard)
  • yt-dlp for YouTube URLs (brew install yt-dlp on macOS)
  • Backend-specific:
    • Gemini API: free API key from ai.google.dev
    • Local: brew install whisper-cpp (macOS) or equivalent
    • OpenAI: API key from OpenAI

MCP Tools

The server exposes 6 MCP tools:

  • video_watch — Extract frames + process audio (main tool)
  • video_analyze — Analyze video structure with ffmpeg filters before extraction
  • video_detail — Drill into specific cached or newly extracted moments
  • video_info — Get video metadata without processing
  • video_configure — Change settings
  • video_setup — Check and guide dependency installation

Configuration

Settings are stored in ~/.claude-video-vision/config.json:

{
  "backend": "local",
  "whisper_engine": "cpp",
  "whisper_model": "auto",
  "whisper_at": false,
  "frame_mode": "images",
  "frame_resolution": 512,
  "default_fps": "auto",
  "max_frames": 100,
  "frame_describer_model": "sonnet",
  "enable_index": false,
  "session_max_age_days": 7
}

Whisper models auto-download to ~/.claude-video-vision/models/ on first use. Available: tiny, base, small, medium, large-v3-turbo, large-v3, auto (picks best for your RAM).

Status

Codex fork: initial enablement. Upstream MCP package is currently claude-video-vision v1.2.1. Tested on macOS (Apple Silicon) with local build, ffmpeg, and whisper.cpp.

License

MIT — see LICENSE.

Author

Jordan Vasconcelos

Star History

Star History Chart

About

Codex-ready fork of claude-video-vision: MCP video perception for Codex

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages