VoiceMode
An MCP server for voice conversations with Claude Code and other MCP agents: speech recognition and voice synthesis, local or cloud
Medium risk
We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.
Why this level
- Accesses the microphone and plays audio on the device
- In cloud mode it sends audio to an external recognition and synthesis service
- Installs and runs local speech services
Install
In your terminal, with SkillFoxx CLI
npx skillfoxx add mcp/voicemodeDetects the agents on your machine, checks the risk and pins the version.
Other ways to install
Run in a terminal
claude mcp add --transport stdio voicemode -- uvx voice-mode==8.12.0Or add to the file .mcp.json, in the project
{
"mcpServers": {
"voicemode": {
"command": "uvx",
"args": [
"voice-mode==8.12.0"
]
}
}
}If the file already exists, add the server inside the mcpServers key.
Keys and settings
OPENAI_API_KEYsecret, optional- OpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGoptional- Enable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSoptional- Skip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALoptional- Prefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMAToptional- Audio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELoptional- Whisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONoptional- Disable silence detection for continuous recording (true/false)
Replace the values in angle brackets with your own. Keys never go into install links and are not stored by us.
The button opens the agent and offers to add the server. If nothing happens, copy the config below.
Add to the file ~/.cursor/mcp.json, for all projects
{
"mcpServers": {
"voicemode": {
"command": "uvx",
"args": [
"voice-mode==8.12.0"
]
}
}
}If the file already exists, add the server inside the mcpServers key. For a single project, put the same block into .cursor/mcp.json.
Keys and settings
OPENAI_API_KEYsecret, optional- OpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGoptional- Enable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSoptional- Skip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALoptional- Prefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMAToptional- Audio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELoptional- Whisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONoptional- Disable silence detection for continuous recording (true/false)
Replace the values in angle brackets with your own. Keys never go into install links and are not stored by us.
The button opens the agent and offers to add the server. If nothing happens, copy the config below.
Run in a terminal
code --add-mcp '{"name":"voicemode","type":"stdio","command":"uvx","args":["voice-mode==8.12.0"]}'Or add to the file .vscode/mcp.json, in the project
{
"servers": {
"voicemode": {
"type": "stdio",
"command": "uvx",
"args": [
"voice-mode==8.12.0"
]
}
}
}If the file already exists, add the server inside the servers key.
Keys and settings
OPENAI_API_KEYsecret, optional- OpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGoptional- Enable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSoptional- Skip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALoptional- Prefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMAToptional- Audio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELoptional- Whisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONoptional- Disable silence detection for continuous recording (true/false)
Replace the values in angle brackets with your own. Keys never go into install links and are not stored by us.
Run in a terminal
codex mcp add voicemode -- uvx voice-mode==8.12.0Or add to the file ~/.codex/config.toml, for all projects
[mcp_servers.voicemode]
command = "uvx"
args = ["voice-mode==8.12.0"]If the file already exists, append the block to the end.
Keys and settings
OPENAI_API_KEYsecret, optional- OpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGoptional- Enable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSoptional- Skip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALoptional- Prefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMAToptional- Audio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELoptional- Whisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONoptional- Disable silence detection for continuous recording (true/false)
Replace the values in angle brackets with your own. Keys never go into install links and are not stored by us.
Run in a terminal
gemini mcp add -s user voicemode uvx voice-mode==8.12.0Or add to the file ~/.gemini/settings.json, for all projects
{
"mcpServers": {
"voicemode": {
"command": "uvx",
"args": [
"voice-mode==8.12.0"
]
}
}
}If the file already exists, add the server inside the mcpServers key.
Keys and settings
OPENAI_API_KEYsecret, optional- OpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGoptional- Enable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSoptional- Skip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALoptional- Prefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMAToptional- Audio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELoptional- Whisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONoptional- Disable silence detection for continuous recording (true/false)
Replace the values in angle brackets with your own. Keys never go into install links and are not stored by us.
Add to the file ~/.config/devin/mcp_config.json, for all projects
{
"mcpServers": {
"voicemode": {
"command": "uvx",
"args": [
"voice-mode==8.12.0"
]
}
}
}If the file already exists, add the server inside the mcpServers key. Legacy Cascade keeps the MCP config in ~/.codeium/windsurf/mcp_config.json.
Keys and settings
OPENAI_API_KEYsecret, optional- OpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGoptional- Enable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSoptional- Skip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALoptional- Prefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMAToptional- Audio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELoptional- Whisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONoptional- Disable silence detection for continuous recording (true/false)
Replace the values in angle brackets with your own. Keys never go into install links and are not stored by us.
Formerly Windsurf.
Add to the file cline_mcp_settings.json, for all projects
{
"mcpServers": {
"voicemode": {
"command": "uvx",
"args": [
"voice-mode==8.12.0"
]
}
}
}If the file already exists, add the server inside the mcpServers key. Open the settings file in Cline: MCP Servers tab, Configure MCP Servers.
Keys and settings
OPENAI_API_KEYsecret, optional- OpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGoptional- Enable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSoptional- Skip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALoptional- Prefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMAToptional- Audio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELoptional- Whisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONoptional- Disable silence detection for continuous recording (true/false)
Replace the values in angle brackets with your own. Keys never go into install links and are not stored by us.
Add to the file .roo/mcp.json, in the project
{
"mcpServers": {
"voicemode": {
"command": "uvx",
"args": [
"voice-mode==8.12.0"
]
}
}
}If the file already exists, add the server inside the mcpServers key.
Keys and settings
OPENAI_API_KEYsecret, optional- OpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGoptional- Enable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSoptional- Skip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALoptional- Prefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMAToptional- Audio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELoptional- Whisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONoptional- Disable silence detection for continuous recording (true/false)
Replace the values in angle brackets with your own. Keys never go into install links and are not stored by us.
A fork of Roo Code, same .roo folders.
Add to the file opencode.json, in the project
{
"mcp": {
"voicemode": {
"type": "local",
"command": [
"uvx",
"voice-mode==8.12.0"
]
}
}
}If the file already exists, add the server inside the mcp key.
Keys and settings
OPENAI_API_KEYsecret, optional- OpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGoptional- Enable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSoptional- Skip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALoptional- Prefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMAToptional- Audio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELoptional- Whisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONoptional- Disable silence detection for continuous recording (true/false)
Replace the values in angle brackets with your own. Keys never go into install links and are not stored by us.
Add to the file ~/.config/zed/settings.json, for all projects
{
"context_servers": {
"voicemode": {
"command": "uvx",
"args": [
"voice-mode==8.12.0"
]
}
}
}If the file already exists, add the server inside the context_servers key.
Keys and settings
OPENAI_API_KEYsecret, optional- OpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGoptional- Enable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSoptional- Skip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALoptional- Prefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMAToptional- Audio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELoptional- Whisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONoptional- Disable silence detection for continuous recording (true/false)
Replace the values in angle brackets with your own. Keys never go into install links and are not stored by us.
Add to the file .codeassistant/mcp.json, in the project
{
"mcpServers": {
"voicemode": {
"command": "uvx",
"args": [
"voice-mode==8.12.0"
]
}
}
}If the file already exists, add the server inside the mcpServers key.
Keys and settings
OPENAI_API_KEYsecret, optional- OpenAI API key for cloud-based STT/TTS (optional - local services can be installed)
VOICEMODE_DEBUGoptional- Enable debug mode with detailed logging (true/false)
VOICEMODE_SKIP_TTSoptional- Skip TTS and show text only for faster response (true/false)
VOICEMODE_PREFER_LOCALoptional- Prefer local services over cloud when available (true/false, default: true)
VOICEMODE_AUDIO_FORMAToptional- Audio format: pcm, mp3, wav, flac, aac, opus (default: pcm)
VOICEMODE_WHISPER_MODELoptional- Whisper model: tiny, base, small, medium, large (default: base)
VOICEMODE_DISABLE_SILENCE_DETECTIONoptional- Disable silence detection for continuous recording (true/false)
Replace the values in angle brackets with your own. Keys never go into install links and are not stored by us.
For Claude Code: claude plugin marketplace add mbailey/voicemode, then claude plugin install voicemode@voicemode, then /voicemode:install and /voicemode:converse. As an MCP server: uvx voice-mode-install, then claude mcp add --scope user voicemode -- uvx --refresh --from voice-mode voicemode-mcp-launcher.
Other ways from the author
claude plugin marketplace add mbailey/voicemode
claude plugin install voicemode@voicemodeAfter install run /voicemode:install for dependencies and local services, then /voicemode:converse.
This is third-party code. Review the repository files before installing.
What it does
VoiceMode gives the agent a converse tool: you speak into the microphone and the agent answers by voice. Speech recognition and synthesis can run fully local via Whisper and Kokoro, or use cloud services behind an OpenAI-compatible API, and it switches between them transparently. It detects silence to stop recording when you pause and focuses on low latency so the exchange feels like a real conversation. The server installs as a Claude Code plugin or as an MCP server through uvx, and runs on Linux, macOS and Windows including WSL. The package is published on PyPI as voice-mode.
Who it is for. For developers who prefer to talk to the agent when their hands or eyes are busy.
Good fit when
- You want to talk to Claude Code by voice instead of typing
- You need to work with the agent away from the screen: on the move, while cooking, when your eyes are tired
- You need local voice input and output without sending audio to the cloud
Not a fit when
- There is no microphone or speakers, or the environment has no audio access
- You only need text input and do not need voice
Example request
Let's switch to voice mode and talk through how to rewrite this moduleLimitations
You need a working microphone and speakers. Local speech services need ffmpeg, portaudio and similar dependencies, plus disk space and compute for the models. The cloud path through OpenAI needs a key and is not reachable from Russia without a VPN, while the local mode works without one. On WSL2 microphone access needs the pulseaudio packages.
How to disable. Remove the plugin via /plugin or detach the MCP server with claude mcp remove voicemode. The local speech services can be stopped and removed separately.
MCP
- Transport
- stdio
- Authentication
- not required
| Environment variables | |
|---|---|
| OPENAI_API_KEY secret | OpenAI key for cloud speech recognition and synthesis, optional; the local mode works without it. |
Security check
- Accesses the microphone and plays audio on the device
- In cloud mode it sends audio to an external recognition and synthesis service
- Installs and runs local speech services
README in short
The README presents VoiceMode as a way to hold voice conversations with Claude Code and other MCP agents. The quick start offers two paths: the Claude Code plugin via its marketplace, or installing the Python package and wiring the MCP server through uvx. Features include natural dialogue, offline use with local speech services, low latency, silence detection and a choice between local and cloud modes. Separate sections cover system dependencies per platform, permission setup, installing Whisper and Kokoro, and common audio problems. It is MIT licensed and the PyPI package is named voice-mode.
FAQ
Can it run offline without an OpenAI key?
Yes. With local Whisper and Kokoro the recognition and synthesis run on your machine. The OpenAI key is only for the cloud path and acts as a fallback.
Which agents does it work with?
The main path is the Claude Code plugin, but the server speaks the MCP protocol, so it also connects to other MCP-compatible agents.
Related
A self-hosted knowledge base with block-level references and a built-in MCP server for connecting AI agents to your notes
A CLI for every Google Workspace API with JSON output and agent skills: Drive, Gmail, Calendar, Sheets and more
Local search over Markdown notes, docs and meeting transcripts: keywords, semantic search and reranking, with an MCP server
A task manager for AI-driven development: breaks a PRD into dependent tasks and guides the agent through them via MCP or CLI