agent-vision-toolkit

A set of CLI tools and a skill that give a text-only agent sight: image analysis, long-screenshot OCR and UI restoration

Skill

Medium risk

We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.

Why this level

  • Sends images to an external multimodal API
  • Runs CLIs and can automate a GUI
All reasons and checks

anionex/agent-vision-toolkit

Install

In your terminal, with SkillFoxx CLI

npx skillfoxx add skills/agent-vision-toolkit

Detects the agents on your machine, checks the risk and pins the version.

Other ways to install

Run in a terminal in the project folder

npx skills add anionex/agent-vision-toolkit --skill vision-skills -a claude-code -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit bf68366
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/anionex/agent-vision-toolkit.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/vision-skills/
git -C "$tmp" checkout bf68366d2a2250691ef42f3ca1b464d2eba1ab36
mkdir -p .claude/skills
cp -R "$tmp/skills/vision-skills" .claude/skills/vision-skills

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

npx skills add anionex/agent-vision-toolkit --skill vision-skills -a cursor -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit bf68366
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/anionex/agent-vision-toolkit.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/vision-skills/
git -C "$tmp" checkout bf68366d2a2250691ef42f3ca1b464d2eba1ab36
mkdir -p .agents/skills
cp -R "$tmp/skills/vision-skills" .agents/skills/vision-skills

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

npx skills add anionex/agent-vision-toolkit --skill vision-skills -a github-copilot -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit bf68366
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/anionex/agent-vision-toolkit.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/vision-skills/
git -C "$tmp" checkout bf68366d2a2250691ef42f3ca1b464d2eba1ab36
mkdir -p .github/skills
cp -R "$tmp/skills/vision-skills" .github/skills/vision-skills

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

npx skills add anionex/agent-vision-toolkit --skill vision-skills -a codex -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit bf68366
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/anionex/agent-vision-toolkit.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/vision-skills/
git -C "$tmp" checkout bf68366d2a2250691ef42f3ca1b464d2eba1ab36
mkdir -p .agents/skills
cp -R "$tmp/skills/vision-skills" .agents/skills/vision-skills

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

npx skills add anionex/agent-vision-toolkit --skill vision-skills -a gemini-cli -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit bf68366
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/anionex/agent-vision-toolkit.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/vision-skills/
git -C "$tmp" checkout bf68366d2a2250691ef42f3ca1b464d2eba1ab36
mkdir -p .agents/skills
cp -R "$tmp/skills/vision-skills" .agents/skills/vision-skills

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/anionex/agent-vision-toolkit.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/vision-skills/
git -C "$tmp" checkout bf68366d2a2250691ef42f3ca1b464d2eba1ab36
mkdir -p .devin/skills
cp -R "$tmp/skills/vision-skills" .devin/skills/vision-skills

Commands for macOS and Linux, on Windows run them in Git Bash.

Formerly Windsurf.

Run in a terminal in the project folder

npx skills add anionex/agent-vision-toolkit --skill vision-skills -a cline -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit bf68366
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/anionex/agent-vision-toolkit.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/vision-skills/
git -C "$tmp" checkout bf68366d2a2250691ef42f3ca1b464d2eba1ab36
mkdir -p .cline/skills
cp -R "$tmp/skills/vision-skills" .cline/skills/vision-skills

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

npx skills add anionex/agent-vision-toolkit --skill vision-skills -a roo -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit bf68366
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/anionex/agent-vision-toolkit.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/vision-skills/
git -C "$tmp" checkout bf68366d2a2250691ef42f3ca1b464d2eba1ab36
mkdir -p .roo/skills
cp -R "$tmp/skills/vision-skills" .roo/skills/vision-skills

Commands for macOS and Linux, on Windows run them in Git Bash.

A fork of Roo Code, same .roo folders.

Run in a terminal in the project folder

npx skills add anionex/agent-vision-toolkit --skill vision-skills -a opencode -y

The skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.

Without third-party tools, from commit bf68366
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/anionex/agent-vision-toolkit.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/vision-skills/
git -C "$tmp" checkout bf68366d2a2250691ef42f3ca1b464d2eba1ab36
mkdir -p .agents/skills
cp -R "$tmp/skills/vision-skills" .agents/skills/vision-skills

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/anionex/agent-vision-toolkit.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/vision-skills/
git -C "$tmp" checkout bf68366d2a2250691ef42f3ca1b464d2eba1ab36
mkdir -p .agents/skills
cp -R "$tmp/skills/vision-skills" .agents/skills/vision-skills

Commands for macOS and Linux, on Windows run them in Git Bash.

Run in a terminal in the project folder

tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/anionex/agent-vision-toolkit.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/vision-skills/
git -C "$tmp" checkout bf68366d2a2250691ef42f3ca1b464d2eba1ab36
mkdir -p .agents/skills
cp -R "$tmp/skills/vision-skills" .agents/skills/vision-skills

Commands for macOS and Linux, on Windows run them in Git Bash.

You will need: Node.js

Checked against the repository on Sep 25, 2026, commit bf68366.

Text for your agent

Install the skill: npx skills add Anionex/agent-vision-toolkit --skill vision-skills. Then set VISION_API_KEY, VISION_BASE_URL and VISION_MODEL in the config.

Other ways from the author
npx skills add Anionex/agent-vision-toolkit --skill vision-skills -g

After install, set VISION_API_KEY, VISION_BASE_URL and VISION_MODEL.

This is third-party code. Review the repository files before installing.

What it does

agent-vision-toolkit adds sight to agents on text-only models such as DeepSeek that lack multimodality. It consists of several CLIs and a skill that teaches the agent when to use each: image Q&A, long-screenshot OCR, frontend UI restoration and GUI automation. Any agent that can call shell commands can use them. There is also a transparent local proxy and single-file plugins so pasted images and built-in image tools work without extra setup.

Who it is for. For people running an agent on a text-only model who want to give it image handling.

Good fit when

  • The agent's model cannot see images but you need to work with them
  • You need OCR of a long screenshot or image analysis
  • You want to restore markup from a UI screenshot

Not a fit when

  • The agent's model is already multimodal
  • You have no access to an external multimodal API

Example request

Read the text on this long screenshot and describe what it shows

Limitations

It needs a multimodal API compatible with OpenAI Chat Completions, OpenAI Responses or Anthropic Messages: its base URL, key and model. Availability depends on the chosen provider.

How to disable. Remove the bin directory from PATH and delete the installed skill or plugins from the agent directory.

Security check

  • Sends images to an external multimodal API
  • Runs CLIs and can automate a GUI

README in short

The README explains that an agent's vision can move out of the model and into the harness: a set of CLIs and a skill teach a text-only agent to analyze images, OCR long screenshots, restore UI and automate GUIs. A local proxy and single-file plugins install optionally for seamless image handling. An external multimodal API is required. MIT licensed.

SKILL.md

---
name: vision-skills
description: >-
  Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a
  target, pixel box), detect (element inventory), trace (image to SVG
  geometry), crop (cut a pixel box to a file), and scripts/html_shot.py (HTML
  file to image). Use for any task involving an image — questions, text,
  splitting and transcribing long screenshots or chat histories, locating elements,
  comparing, rebuilding as HTML/SVG, digitizing a sketch or diagram, reading
  values off a chart, operating a GUI from screenshots — and to re-check an
  image yourself when a description you were given lacks a detail.
---

# vision-skills

Five local CLIs that give a text-only agent eyes. They read one shared
vision config (`VISION_API_KEY` / `VISION_BASE_URL` / `VISION_MODEL` /
`LANG`), plus the optional Python-client settings `VISION_API_PROTOCOL`,
`VISION_REASONING_EFFORT`, and `VISION_USER_AGENT` — no extra credentials.

Pick the tool by the question you are answering:

| Question | Tool |
|---|---|
| "What does this image show / say?" | `glance` |
| "Where is X?" — a thing you can name | `ground` |
| "Where are all the Xs?" — every instance of a kind | `detect` |
| "What is its exact shape, size, offset?" | `trace` |
| "Cut this box out as its own image file" | `crop` |
| "OCR this long screenshot / scrolling page / chat history" | `scripts/long_screenshot_ocr.py` |
| "Extract the icon/logo foreground as transparent PNG — manual region or auto (cropped+scaled screenshots)" | `scripts/extract_fg.py` |
| "Turn this HTML file into a viewport or full-page screenshot" | `scripts/html_shot.py` |
| "Which colours dominate a region, and which palette value fits it?" | `scripts/dominant_colors.py` |
| A relation none of them return — a gap, a distance between two located things | code over the pixels (Pillow) |

`glance` answers what something is; `ground` and `detect` answer where.
You give `ground` a description of a particular thing; you give `detect` a
kind and it enumerates the instances.

Both give real coordinates, but they are not pixel-exact: the box arrives
on a 0-1000 grid and is scaled to your image, so the last pixel or few are
not reliable. That is accurate enough to crop with, to click, to compare
positions against. When a number has to be exact, `trace` derives it from
the actual pixels — offsets, sizes, shapes.

## Use the provided tools before hand-rolled pixels

Everything this toolkit ships a tool for, call the tool — do not rewrite

FAQ

What do I need to prepare?

A multimodal API compatible with OpenAI or Anthropic: a base URL, key and model name.

Which agents does it support?

Codex, Claude Code, Pi, Oh My Pi and OpenCode.

Editors’ pick

Open-source personal AI assistant on your own machine: answers in Telegram, Slack, Discord and WhatsApp, extended with skills and plugins

CLIHigh risk390.7KRepository stars
Editors’ pick

A self-improving agent from Nous Research with a TUI, messaging gateway, cron jobs and skills it writes itself

CLIHigh risk249.8KRepository stars
Editors’ pick

An open source coding agent for the terminal and desktop with build and plan modes

CLIHigh risk210.6KRepository stars
Editors’ pick

Open prompt library with a Claude Code plugin, MCP server and CLI: search, fetch and improve prompts and skills from an agent

PluginMedium risk171.5KRepository stars
Foxx AIagent-vision-toolkit

I am Foxx AI and I have already vetted this tool. Ask about install, setup or anything else, and I will keep it simple.