YouTube Transcript MCP Server
MCP server for YouTube transcripts: an agent gets a video transcript by URL, with timestamps and metadata
Medium risk
We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.
Why this level
- Makes network calls to YouTube for captions and metadata
- May use third-party proxy servers with a username and password
Install
Manual install
uvx --from git+https://github.com/jkawamoto/mcp-youtube-transcript mcp-youtube-transcriptRequires uv installed. No keys are needed; set a proxy if YouTube access is restricted.
This is third-party code. Review the repository files before installing.
What it does
The server returns a YouTube video transcript to an agent by URL. It offers a plain transcript, a timestamped variant, video metadata and a list of available languages. You can set the language, and long transcripts are split into parts served by cursor to fit the model's context limit. The split threshold is set with the response-limit flag. Proxies are supported for regions with restricted YouTube access.
Who it is for. For content creators, marketers and technical writers who work with the content of YouTube videos.
Good fit when
- You need a video transcript for notes or an article
- You need a timestamped transcript for clipping or subtitles
- You want a video's metadata and available languages
Not a fit when
- The video has no available captions or transcript
- You need image or audio analysis rather than the video's text
Example request
Get this video's transcript with timestamps and turn it into a short summary by sectionLimitations
The server works only with videos that have captions or an automatic transcript. YouTube access from Russia is restricted and throttled, so without a proxy results may fail; the README provides proxy variables for this. It installs via uvx straight from the git repository, with a Docker Hub image and a Claude Desktop bundle. Very long transcripts are returned in parts by cursor.
How to disable. Remove the youtube-transcript block from your MCP client config and restart it. If you installed the bundle, remove it in Claude Desktop settings.
MCP
- Transport
- stdio
- Authentication
- not required
| Environment variables | |
|---|---|
| WEBSHARE_PROXY_USERNAME secret | Webshare residential proxy username when YouTube access is restricted |
| WEBSHARE_PROXY_PASSWORD secret | Webshare residential proxy password |
| HTTPS_PROXY | Proxy server URL for reaching YouTube |
Security check
- Makes network calls to YouTube for captions and metadata
- May use third-party proxy servers with a username and password
README in short
The README lists the tools get_transcript, get_timed_transcript, get_video_info and get_available_languages with URL, language and cursor parameters. Installation is shown via uvx from the git repository, a Claude Desktop bundle, options for goose and LM Studio, and a Docker Hub image. Pagination of long transcripts via next_cursor and the response-limit flag is explained separately. There is a proxy section for regions with restricted YouTube access using Webshare or the HTTP_PROXY and HTTPS_PROXY variables. MIT licensed.
FAQ
What about a very long video?
Transcripts longer than the set threshold are split. The response includes next_cursor; pass it in the next request to get the continuation. The threshold is changed with the response-limit flag.
What if YouTube access is restricted?
The README describes proxies: a Webshare username and password, or the HTTP_PROXY and HTTPS_PROXY variables for other proxy servers.
Related
Skills and CLI from HeyGen: the agent writes HTML compositions and renders them into MP4 videos, captions, motion graphics and decks
Blender MCP
MCP for Blender
An MCP server and Blender addon: the agent creates and edits objects, materials and scenes, pulls assets and runs Python in Blender
Official Remotion skills for making videos in React: compositions, animation, captions, maps and rendering
Agentic video studio for Claude Code, Cursor and Codex: script, voiceover, images, stock footage, music, subtitles and render