2.1.0
Added
- The tool has its own icon, in a dark and a light version. Codex shows it on the plugin card and in the skill list, and the toolkit site shows it on the tool's page and in share previews.
Watch / analyze a video or audio source the agent cannot natively play. A URL from YouTube, TikTok, Instagram, X, Facebook, Vimeo, Reddit, Twitch, LinkedIn or 1750+ other sites, or a local file. Downloads it, pulls captions (or transcribes locally when there are none), and extracts frames so the agent can Read the transcript + frames and describe pacing, hooks, on-screen text, format, or answer questions about the content. No API key needed.
Once it is installed, that is the whole interface. Your agent picks the skill up on its own and runs whatever it needs to. The commands further down are there for anyone who would rather drive it themselves.
bash scripts/setup.sh --check # what's installed, what's missing
bash scripts/setup.sh # shows a plan, asks, installs
bash scripts/setup.sh --yes # non-interactive
bash scripts/setup.sh --with-whisper # add free offline speech-to-text
Detects your platform and package manager — Homebrew, apt, dnf, pacman, zypper, apk, pipx/pip — and prints every command before it runs. Nothing is installed without a confirmation or an explicit --yes, and anything needing root is shown in full first.
| Tool | Needed for | Notes |
|---|---|---|
ffmpeg |
always | Frames and media metadata |
yt-dlp |
URLs only | Installed via pip where distro packages go stale |
| a whisper engine | optional | Free and local. Used automatically when a source has no captions. setup.sh picks mlx-whisper on Apple Silicon, whisper-ctranslate2 elsewhere |
| Option | Default | Meaning |
|---|---|---|
--frames N |
30 (max 60) | Target frame count |
--width W |
480 | Frame width in px |
--start SEC / --end SEC |
— | Sample a time range only |
--lang CODE |
en.* |
Caption language, e.g. ar.*, all |
--scenes |
off | One frame per visual cut |
--no-whisper |
off | Skip local transcription even if an engine is installed |
--whisper-model M |
base |
Speed vs accuracy, e.g. large-v3 |
--yes |
off | Accept prompts, including installing a transcriber |
--cookies BROWSER |
— | Login-gated videos, e.g. chrome |
--playlist N |
single | Allow a playlist, take first N |
--outdir DIR |
temp | Where output lands |
--keep-video |
off | Keep the download |
yt-dlp fetches the media (capped at 1920px tall) plus auto or uploaded subtitles, converted to .srt. Local files skip this step. Playlists are refused unless --playlist N is given, so one URL means one video.ffprobe reads duration and dimensions. A source with no video stream is handled as audio-only: transcript, no frames.ffmpeg extracts frames, either at an even interval or on scene changes with --scenes. Scene mode falls back to even sampling when a source has too few cuts to be useful.--cookies reuses a session you are already signed into; it does not defeat access control.--start / --end.--no-whisper skips it entirely.setup.sh updates it./watch-video:watch-video, with the plugin name in front. This has been true since 1.3.0, and this release marks it as breaking. Use the new name, or ask in words, which works as before. A copy in a skills folder keeps the plain name /watch-video.agents/openai.yaml inside the skill folder. Other agents ignore that file.SKILL.md. It tells the agent how to check for a newer version, and to report what changed before it updates.An agent that has this installed reads CHANGELOG.md in the skill folder.
The same history is at changelog.json,
and every tool's releases are at /updates.json.
Built at Itqan Lab, a design and technology studio.
إتقان — itqan, the Arabic word for mastery: doing a thing precisely, and completely.