video-use

A video editing skill for Claude Code and Codex: the agent cuts pauses and filler words, grades color, burns subtitles and adds animated overlays

SkillEditors’ pick

Medium risk

We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.

Why this level

  • The agent runs Python scripts and ffmpeg on your machine and, on first use, installs HyperFrames or Remotion through npx for animations
  • Audio from every source goes to the ElevenLabs cloud, and each transcription spends account credits
  • The ElevenLabs key is stored in plain text in the .env file inside the skill folder
All reasons and checks

browser-use/video-use

Install

In your terminal, with SkillFoxx CLI

npx skillfoxx add skills/video-use

Detects the agents on your machine, checks the risk and pins the version.

Other ways to install

Run in a terminal in the project folder

npx skills add browser-use/video-use --skill video-use -a claude-code -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add browser-use/video-use --skill video-use -a cursor -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add browser-use/video-use --skill video-use -a github-copilot -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add browser-use/video-use --skill video-use -a codex -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add browser-use/video-use --skill video-use -a gemini-cli -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add browser-use/video-use --skill video-use -a cline -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add browser-use/video-use --skill video-use -a roo -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

A fork of Roo Code, same .roo folders.

Run in a terminal in the project folder

npx skills add browser-use/video-use --skill video-use -a opencode -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

You will need: Node.js

Checked against the repository on Oct 1, 2026, commit b877063.

Text for your agent

Set up https://github.com/browser-use/video-use. Read install.md first: clone the repo to ~/Developer/video-use, install dependencies with uv sync or pip install -e ., check ffmpeg and register the skill by symlinking the whole folder into your agent's skills directory. Ask me for the ElevenLabs API key and write it to .env, and ask before running brew install. After install, do not transcribe anything; tell me it is ready and wait for a folder of footage.

Other ways from the author
Text for your agent

Set up https://github.com/browser-use/video-use for me. Read install.md first to install this repo, wire up ffmpeg, register the skill with whichever agent you're running under, and set up the ElevenLabs API key — ask me to paste it when you need it. Then read SKILL.md for daily usage, and always read helpers/ because that's where the editing scripts live. After install, don't transcribe anything on your own — just tell me it's ready and wait for me to drop footage into a folder.

The setup prompt from the README for Claude Code, Codex and other agents with shell access. The agent clones the repo, installs dependencies and ffmpeg, registers the skill and asks for the ElevenLabs key once.

Security check

How we review

Medium riskWe rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.

  • The agent runs Python scripts and ffmpeg on your machine and, on first use, installs HyperFrames or Remotion through npx for animations
  • Audio from every source goes to the ElevenLabs cloud, and each transcription spends account credits
  • The ElevenLabs key is stored in plain text in the .env file inside the skill folder

This is third-party code. Review the repository files before installing.

What it does

video-use edits raw takes into a finished video through a chat with your agent. You drop the footage into a folder, start Claude Code or Codex there and describe the job, for example "edit these into a launch video". The agent transcribes the audio with ElevenLabs Scribe at word-level timestamps, reads a compact transcript and cuts on word boundaries, dropping filler words, bad takes and dead air. It looks at the picture only at decision points, through a filmstrip with a waveform. Then ffmpeg scripts apply the color grade, 30 ms audio fades at every cut and subtitles, while sub-agents build animated overlays in parallel with HyperFrames, Remotion, Manim or PIL.

Who it is for. For video creators, marketers and developers who record talking heads, tutorials or product demos and want an agent to handle the rough cut.

Good fit when

  • You have several takes of the same script and need one video assembled from the best parts
  • You need to remove pauses, ums and slips from an interview or lesson recording
  • The video needs subtitles, a color grade and animated overlays without manual work in an editing app
  • The edit spans several sessions: the agent picks up earlier decisions from project.md

Not a fit when

  • You have no ElevenLabs key or cannot send the recording's audio to an external service: the skill does nothing without a transcript
  • The footage has almost no speech: the skill finds cuts from words and pauses in the audio and checks the picture only selectively
  • You want hands-on control of every cut on a timeline, as in an editing app

Example request

Edit these takes into a 60-second launch video: cut the slips and pauses, add subtitles and a warm color grade

Limitations

You need Python 3.10 or newer, ffmpeg with ffprobe and an ElevenLabs API key. Transcription runs in the cloud Scribe service, every call spends account credits, and nothing works without the key. HyperFrames and Remotion overlays need Node.js, version 22 or newer for HyperFrames, and Manim also needs LaTeX. The agent cannot listen to the result, so it measures loudness with ffmpeg and reports the numbers.

How to disable. Remove the ~/.claude/skills/video-use or ~/.codex/skills/video-use symlink, then the ~/Developer/video-use folder. The ElevenLabs key sits in that folder's .env file; revoke it in your ElevenLabs account settings if needed. Finished videos and transcripts stay in the edit folder next to your sources.

FAQ

Does the agent watch the whole video?

No. It reads a transcript with word-level timestamps, and all takes fit into roughly 12 KB of text. The agent renders a filmstrip with a waveform only at ambiguous moments and to check the cuts after rendering.

Where does the skill write files?

Everything lands in an edit folder next to your sources: transcripts, cut decisions in edl.json, subtitles, the preview and final.mp4. Source files and the skill folder stay untouched.

Can I use it without ElevenLabs?

Not in this repository: ElevenLabs Scribe handles all transcription, and the skill relies on its word timestamps, speaker labels and audio events such as laughter.

Skills
Editors’ pick

Skills and CLI from HeyGen: the agent writes HTML compositions and renders them into MP4 videos, captions, motion graphics and decks

SkillMedium risk54.4KRepository stars

Blender MCP

MCP for Blender

Editors’ pick

An MCP server and Blender addon: the agent creates and edits objects, materials and scenes, pulls assets and runs Python in Blender

MCP serverHigh risk29.8KRepository stars
Editors’ pick

Official Remotion skills for making videos in React: compositions, animation, captions, maps and rendering

SkillLow risk4.8KRepository stars

Agentic video studio for Claude Code, Cursor and Codex: script, voiceover, images, stock footage, music, subtitles and render

WorkflowMedium risk62KRepository stars

README in short

The README opens with a ready setup prompt for Claude Code, Codex, Hermes or OpenClaw, and the agent installs the repo, ffmpeg and the ElevenLabs key on its own. Manual install covers cloning, the symlink, uv sync and brew install ffmpeg. The authors then describe the two layers the model uses to understand a video: a word-timestamped transcript and an on-demand filmstrip with a waveform. The pipeline runs from transcript to an edit decision list, render and a self-check repeated up to three times. Detailed editing rules live in SKILL.md and agent install steps in install.md.

README badge

Are you the author? Show in your README that the project is reviewed in the SkillFoxx catalog. The badge updates itself.

SkillFoxx badge for this project
[![Проверено SkillFoxx](https://skillfoxx.ru/badges/skills/video-use.svg)](https://skillfoxx.ru/skills/video-use?utm_source=github&utm_medium=badge&utm_campaign=readme)

SKILL.md

---
name: video-use
description: Edit any video by conversation. Transcribe, cut, color grade, generate overlay animations, burn subtitles — for talking heads, montages, tutorials, travel, interviews. No presets, no menus. Ask questions, confirm the plan, execute, iterate, persist. Production-correctness rules are hard; everything else is artistic freedom.
---

# Video Use

## Principle

1. **LLM reasons from raw transcript + on-demand visuals.** The only derived artifact that earns its keep is a packed phrase-level transcript (`takes_packed.md`). Everything else — filler tagging, retake detection, shot classification, emphasis scoring — you derive at decision time.
2. **Audio is primary, visuals follow.** Cut candidates come from speech boundaries and silence gaps. Drill into visuals only at decision points.
3. **Ask → confirm → execute → iterate → persist.** Never touch the cut until the user has confirmed the strategy in plain English.
4. **Generalize.** Do not assume what kind of video this is. Look at the material, ask the user, then edit.
5. **Artistic freedom is the default.** Every specific value, preset, font, color, duration, pitch structure, and technique in this document is a *worked example* from one proven video — not a mandate. Read them to understand what's possible and why each worked. Then make your own taste calls based on what the material actually is and what the user actually wants. **The only things you MUST do are in the Hard Rules section below.** Everything else is yours.
6. **Invent freely.** If the material calls for a technique not described here — split-screen, picture-in-picture, lower-third identity cards, reaction cuts, speed ramps, freeze frames, crossfades, match cuts, L-cuts, J-cuts, speed ramps over breath, whatever — build it. The helpers are ffmpeg and PIL. They can do anything the format supports. Do not wait for permission.
7. **Verify your own output before showing it to the user.** If you wouldn't ship it, don't present it.

## Hard Rules (production correctness — non-negotiable)

These are the things where deviation produces silent failures or broken output. They are not taste, they are correctness. Memorize them.

1. **Subtitles are applied LAST in the filter chain**, after every overlay. Otherwise overlays hide captions. Silent failure.
2. **Per-segment extract → lossless `-c copy` concat**, not single-pass filtergraph. Otherwise you double-encode every segment when overlays are added.
3. **30ms audio fades at every segment boundary** (`afade=t=in:st=0:d=0.03,afade=t=out:st={dur-0.03}:d=0.03`). Otherwise audible pops at every cut.
4. **Overlays use `setpts=PTS-STARTPTS+T/TB`** to shift the overlay's frame 0 to its window start. Otherwise you see the middle of the animation during the overlay window.
5. **Master SRT uses output-timeline offsets**: `output_time = word.start - segment_start + segment_offset`. Otherwise captions misalign after segment concat.
6. **Never cut inside a word.** Snap every cut edge to a word boundary from the Scribe transcript.
7. **Pad every cut edge.** Working window: 30–200ms. Scribe timestamps drift 50–100ms — padding absorbs the drift. Tighter for fast-paced, looser for cinematic.
8. **Word-level verbatim ASR only.** Never SRT/phrase mode (loses sub-second gap data). Never normalized fillers (loses editorial signal).
9. **Cache transcripts per source.** Never re-transcribe unless the source file itself changed.
10. **Parallel sub-agents for multiple animations.** Never sequential. Spawn N at once via the `Agent` tool; total wall time ≈ slowest one.
11. **Strategy confirmation before execution.** Never touch the cut until the user has approved the plain-English plan.
12. **All session outputs in `<videos_dir>/edit/`.** Never write inside the `video-use/` project directory.

Everything else in this document is a worked example. Deviate whenever the material calls for it.

## Directory layout

The skill lives in `video-use/`. User footage lives wherever they put it. All session outputs go into `<videos_dir>/edit/`.
Foxx AIvideo-use

I am Foxx AI and I have already vetted this tool. Ask about install, setup or anything else, and I will keep it simple.