Claude Vision Skill

A skill that gives image recognition to models without native vision: it sends the image to a vision model and returns a text description

Skill

Medium risk

We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.

Why this level

  • Runs the vision.js script, reads files and the clipboard, and sends images to an external vision service
  • Requires an API key stored in a file next to the script
All reasons and checks

asuojun/claude-vision-skill

Install

In your terminal, with SkillFoxx CLI

npx skillfoxx add skills/claude-vision-skill

Detects the agents on your machine, checks the risk and pins the version.

Other ways to install

Run in a terminal in the project folder

npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a claude-code -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a cursor -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a github-copilot -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a codex -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a gemini-cli -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a cline -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Run in a terminal in the project folder

npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a roo -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

A fork of Roo Code, same .roo folders.

Run in a terminal in the project folder

npx skills add asuojun/claude-vision-skill --skill claude-vision-skill -a opencode -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

You will need: Node.js

Checked against the repository on Sep 25, 2026, commit fa5ca17.

Text for your agent

Clone the repository: git clone https://github.com/asuojun/claude-vision-skill.git. Copy vision.js into the project, set the API key, model name and, if needed, the API base URL, and add the CLAUDE.md content to the project. In SKILL.md replace the path to vision.js with your own.

Other ways from the author
git clone https://github.com/asuojun/claude-vision-skill.git

After cloning, ask the agent to read the README and set the skill up; it will ask for the service and key.

This is third-party code. Review the repository files before installing.

What it does

The skill adds image recognition to models that have no native vision, such as DeepSeek. When the user gives a file path, a URL or pastes an image, the agent runs the bundled vision.js script instead of trying to view the image directly. The script reads the image, encodes it as base64 and sends it to a vision model through an OpenAI-compatible format, then returns the answer as text. By default it recommends Qwen on the Aliyun DashScope platform, but any vision model with an OpenAI-compatible API fits, including gpt-4o-mini. It can read an image from the clipboard, using a bundled Swift helper on macOS and a PowerShell script on Windows, and falls back to the clipboard automatically when a file is not found.

Who it is for. For people working with a model that has no vision but who still want the agent to read screenshots and images.

Good fit when

  • The main model cannot read images but you need to parse a screenshot or picture
  • You want to paste an image into chat and get a description without manual commands
  • You need a cheap or free vision model through an OpenAI-compatible API

Not a fit when

  • Your model already reads images directly
  • You cannot send images to an external vision service for privacy reasons

Example request

Look at this screenshot and describe it in text

Limitations

The skill does not recognize images itself; it only forwards them to an external vision model, so you need an API key and pay the service's rates, and images leave for the provider. It needs Node, plus a Swift helper on macOS and a PowerShell script on Windows for the clipboard. By default descriptions come in Chinese unless you ask otherwise. The recommended Qwen on Aliyun is reachable from Russia without a VPN, while the OpenAI option does not work without one. SKILL.md hardcodes the author's absolute path, which you must replace with your own on install. No license is stated in the repository.

How to disable. Remove the skill folder from the skills directory, delete vision.js from the project and drop the skill's mentions from CLAUDE.md.

Security check

  • Runs the vision.js script, reads files and the clipboard, and sends images to an external vision service
  • Requires an API key stored in a file next to the script

README in short

The Chinese README explains that the skill gives vision to models that lack it by sending the image to a vision model and returning text. The core is the vision.js script in an OpenAI-compatible format that is not tied to one provider. It describes three scenarios for Claude Code: a plain project, cyberboss integration and a short explanation of how it works. It recommends Qwen on Aliyun for its free tier, but OpenAI or any compatible service works after changing the base URL and model. Installation is automatic via the agent following the README, or manual by copying vision.js and CLAUDE.md into the project.

SKILL.md

---
name: claude-vision-skill
description: Use when the user shares, pastes, or references an image (local path or URL) and you need to describe, analyze, or recognize its content, especially when the current model cannot read images directly. Run the bundled vision.js helper to convert the image into text.
---

# Vision Helper

The current model may not support native image input. When the user provides an image path or URL, do not rely on viewing the image directly. Instead run:

node /path/to/claude-vision-skill/vision.js "<absolute image path>" "<prompt>"

For an image URL use --url, for a pasted image use --clipboard. The --clipboard mode reads the current image from the system clipboard (macOS uses a bundled Swift helper, Windows a bundled PowerShell script). If a local path does not exist, vision.js falls back to the clipboard automatically; pass --no-fallback to disable this.

Rules: always use the absolute path to vision.js; use an absolute image path for local files or --url for remote images; configuration lives in .env next to vision.js (DASHSCOPE_API_KEY, VISION_MODEL, DASHSCOPE_BASE_URL); never print or commit the API key.

FAQ

Which vision model should I use?

The author recommends Qwen on Aliyun DashScope for its free tier for new users. Any other model with an OpenAI-compatible API also fits, including gpt-4o-mini; you then change the base URL and model name.

Do my images leave my machine?

Yes, the script sends the image to an external vision model. The API key is read from a .env next to vision.js and must not end up in the repository.

Editors’ pick

Open-source personal AI assistant on your own machine: answers in Telegram, Slack, Discord and WhatsApp, extended with skills and plugins

CLIHigh risk390.7KRepository stars
Editors’ pick

A self-improving agent from Nous Research with a TUI, messaging gateway, cron jobs and skills it writes itself

CLIHigh risk249.8KRepository stars
Editors’ pick

An open source coding agent for the terminal and desktop with build and plan modes

CLIHigh risk210.6KRepository stars
Editors’ pick

Open prompt library with a Claude Code plugin, MCP server and CLI: search, fetch and improve prompts and skills from an agent

PluginMedium risk171.5KRepository stars
Foxx AIClaude Vision Skill

I am Foxx AI and I have already vetted this tool. Ask about install, setup or anything else, and I will keep it simple.