Claude Vision Skill
A skill that gives image recognition to models without native vision: it sends the image to a vision model and returns a text description
Medium risk
We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.
Why this level
- Runs the vision.js script, reads files and the clipboard, and sends images to an external vision service
- Requires an API key stored in a file next to the script
Install
In your terminal, with SkillFoxx CLI
npx skillfoxx add skills/claude-vision-skillDetects the agents on your machine, checks the risk and pins the version.
Other ways to install
Run in a terminal in the project folder
npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a claude-code -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a cursor -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a github-copilot -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a codex -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a gemini-cli -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a cline -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Run in a terminal in the project folder
npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a roo -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
A fork of Roo Code, same .roo folders.
Run in a terminal in the project folder
npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a opencode -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Clone the repository: git clone https://github.com/asuojun/claude-vision-skill.git. Copy vision.js into the project, set the API key, model name and, if needed, the API base URL, and add the CLAUDE.md content to the project. In SKILL.md replace the path to vision.js with your own.
Other ways from the author
git clone https://github.com/asuojun/claude-vision-skill.gitAfter cloning, ask the agent to read the README and set the skill up; it will ask for the service and key.
This is third-party code. Review the repository files before installing.
What it does
The skill adds image recognition to models that have no native vision, such as DeepSeek. When the user gives a file path, a URL or pastes an image, the agent runs the bundled vision.js script instead of trying to view the image directly. The script reads the image, encodes it as base64 and sends it to a vision model through an OpenAI-compatible format, then returns the answer as text. By default it recommends Qwen on the Aliyun DashScope platform, but any vision model with an OpenAI-compatible API fits, including gpt-4o-mini. It can read an image from the clipboard, using a bundled Swift helper on macOS and a PowerShell script on Windows, and falls back to the clipboard automatically when a file is not found.
Who it is for. For people working with a model that has no vision but who still want the agent to read screenshots and images.
Good fit when
- The main model cannot read images but you need to parse a screenshot or picture
- You want to paste an image into chat and get a description without manual commands
- You need a cheap or free vision model through an OpenAI-compatible API
Not a fit when
- Your model already reads images directly
- You cannot send images to an external vision service for privacy reasons
Example request
Look at this screenshot and describe it in textLimitations
The skill does not recognize images itself; it only forwards them to an external vision model, so you need an API key and pay the service's rates, and images leave for the provider. It needs Node, plus a Swift helper on macOS and a PowerShell script on Windows for the clipboard. By default descriptions come in Chinese unless you ask otherwise. The recommended Qwen on Aliyun is reachable from Russia without a VPN, while the OpenAI option does not work without one. SKILL.md hardcodes the author's absolute path, which you must replace with your own on install. No license is stated in the repository.
How to disable. Remove the skill folder from the skills directory, delete vision.js from the project and drop the skill's mentions from CLAUDE.md.
Security check
- Runs the vision.js script, reads files and the clipboard, and sends images to an external vision service
- Requires an API key stored in a file next to the script
README in short
The Chinese README explains that the skill gives vision to models that lack it by sending the image to a vision model and returning text. The core is the vision.js script in an OpenAI-compatible format that is not tied to one provider. It describes three scenarios for Claude Code: a plain project, cyberboss integration and a short explanation of how it works. It recommends Qwen on Aliyun for its free tier, but OpenAI or any compatible service works after changing the base URL and model. Installation is automatic via the agent following the README, or manual by copying vision.js and CLAUDE.md into the project.
SKILL.md
--- name: claude-vision-skill description: Use when the user shares, pastes, or references an image (local path or URL) and you need to describe, analyze, or recognize its content, especially when the current model cannot read images directly. Run the bundled vision.js helper to convert the image into text. --- # Vision Helper The current model may not support native image input. When the user provides an image path or URL, do not rely on viewing the image directly. Instead run: node /path/to/claude-vision-skill/vision.js "<absolute image path>" "<prompt>" For an image URL use --url, for a pasted image use --clipboard. The --clipboard mode reads the current image from the system clipboard (macOS uses a bundled Swift helper, Windows a bundled PowerShell script). If a local path does not exist, vision.js falls back to the clipboard automatically; pass --no-fallback to disable this. Rules: always use the absolute path to vision.js; use an absolute image path for local files or --url for remote images; configuration lives in .env next to vision.js (DASHSCOPE_API_KEY, VISION_MODEL, DASHSCOPE_BASE_URL); never print or commit the API key.
FAQ
Which vision model should I use?
The author recommends Qwen on Aliyun DashScope for its free tier for new users. Any other model with an OpenAI-compatible API also fits, including gpt-4o-mini; you then change the base URL and model name.
Do my images leave my machine?
Yes, the script sends the image to an external vision model. The API key is read from a .env next to vision.js and must not end up in the repository.
Related
Open-source personal AI assistant on your own machine: answers in Telegram, Slack, Discord and WhatsApp, extended with skills and plugins
A self-improving agent from Nous Research with a TUI, messaging gateway, cron jobs and skills it writes itself
An open source coding agent for the terminal and desktop with build and plan modes
Open prompt library with a Claude Code plugin, MCP server and CLI: search, fetch and improve prompts and skills from an agent