Give Codex the ability to watch and understand videos.
A Codex-ready fork of claude-video-vision. It extracts frames via ffmpeg and processes audio via multiple backends (Gemini API, local Whisper, or OpenAI API). Codex receives frames as images plus audio transcription with timestamps — the plugin is a perception layer, not an interpretation layer.
- Multimodal perception — Codex sees video frames directly and reads audio transcriptions with timestamps
- YouTube URL support — pass a YouTube URL and the MCP server downloads it with
yt-dlp, preserving source metadata and captions for context - Flexible backends — Choose between cloud APIs or fully local processing
- Adaptive extraction — Codex adjusts fps, time range, and resolution based on your question
- Auto-installation — Whisper models download automatically on first use
- Codex skill workflow —
video-perceptionguides frame extraction, analysis, and drill-down
Clone and build the fork:
git clone https://github.com/miketuckman/codex-video-vision.git
cd codex-video-vision/mcp-server
npm ci
npm run build
npm testAdd the MCP server to ~/.codex/config.toml:
[mcp_servers.codex-video-vision]
command = "node"
args = ["/absolute/path/to/codex-video-vision/mcp-server/dist/index.js"]
cwd = "/absolute/path/to/codex-video-vision/mcp-server"
startup_timeout_sec = 30
tool_timeout_sec = 600Then copy or install the skill from skills/video-perception/ into your Codex skills path, or install this repo as a local Codex plugin once plugin support is available in your environment.
See docs/codex.md for full Codex setup, verification, and recommended defaults.
Until this fork publishes its own npm package, you can run the upstream MCP server directly:
npx -y claude-video-vision@latestUse this Codex config when you prefer the npm package over a local build:
[mcp_servers.codex-video-vision]
command = "npx"
args = ["-y", "claude-video-vision@latest"]
startup_timeout_sec = 30
tool_timeout_sec = 600Use the video_configure MCP tool from Codex. Recommended first-pass local settings:
{
"backend": "local",
"whisper_engine": "cpp",
"whisper_model": "auto",
"whisper_at": false,
"frame_mode": "images",
"frame_resolution": 512,
"default_fps": "auto",
"max_frames": 100,
"enable_index": false,
"session_max_age_days": 7
}Just mention a video file — the Codex skill will guide the workflow:
"analyze this video for me: ~/Downloads/demo.mp4"
"take a look at the first second of ~/videos/bug-report.mov"
"analyze and summarize this YouTube video: https://www.youtube.com/watch?v=..."
Codex adapts parameters automatically:
- "the first second" → extracts at original fps from
00:00:00to00:00:01 - "summarize this 1h lecture" → low fps, full duration
- "what text is on screen at 1:30?" → high resolution, narrow time window
| Backend | Audio processing | Cost | Setup |
|---|---|---|---|
| Gemini API | Native (speech + non-speech events) | Free tier: 1500 req/day | GEMINI_API_KEY env var |
| Local (Whisper) | whisper.cpp or Python openai-whisper |
Free, fully offline | brew install whisper-cpp + auto model download |
| OpenAI API | OpenAI Whisper API | Paid per usage | OPENAI_API_KEY env var |
All backends extract video frames via ffmpeg — Codex always has direct visual access.
┌───────────────────────────────────────────────────────┐
│ Codex (your session) │
│ │
│ video request ──→ Skill: video-perception │
│ │ │
│ ▼ │
│ MCP tool: video_watch │
│ │ │
└────────────────────────┼──────────────────────────────┘
│
▼
┌────────────────────────────────────┐
│ MCP Server (Node.js) │
│ │
│ ┌──────────┐ ┌──────────────┐ │
│ │ ffmpeg │ │ Audio backend│ │
│ │ frames │ ║ │ (parallel) │ │
│ └──────────┘ └──────────────┘ │
│ │ │ │
└───────┼─────────────────┼──────────┘
▼ ▼
base64 images transcription
+ timestamps + audio events
│ │
└────────┬────────┘
▼
Codex receives both
- Node.js 20+ (for the MCP server)
- ffmpeg (auto-detected, install instructions provided by setup wizard)
- yt-dlp for YouTube URLs (
brew install yt-dlpon macOS) - Backend-specific:
- Gemini API: free API key from ai.google.dev
- Local:
brew install whisper-cpp(macOS) or equivalent - OpenAI: API key from OpenAI
The server exposes 6 MCP tools:
video_watch— Extract frames + process audio (main tool)video_analyze— Analyze video structure with ffmpeg filters before extractionvideo_detail— Drill into specific cached or newly extracted momentsvideo_info— Get video metadata without processingvideo_configure— Change settingsvideo_setup— Check and guide dependency installation
Settings are stored in ~/.claude-video-vision/config.json:
{
"backend": "local",
"whisper_engine": "cpp",
"whisper_model": "auto",
"whisper_at": false,
"frame_mode": "images",
"frame_resolution": 512,
"default_fps": "auto",
"max_frames": 100,
"frame_describer_model": "sonnet",
"enable_index": false,
"session_max_age_days": 7
}Whisper models auto-download to ~/.claude-video-vision/models/ on first use. Available: tiny, base, small, medium, large-v3-turbo, large-v3, auto (picks best for your RAM).
Codex fork: initial enablement. Upstream MCP package is currently claude-video-vision v1.2.1. Tested on macOS (Apple Silicon) with local build, ffmpeg, and whisper.cpp.
MIT — see LICENSE.
Jordan Vasconcelos
- GitHub: @jordanrendric
- LinkedIn: jordanvasconcelos
- Instagram: @jordanvasconcelos__
- X/Twitter: @jordanrendric
