ModLens
A skill that gives text-only models sight: an image by path or from a paste becomes structured JSON with transcribed text, layout and semantics
Medium risk
We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.
Why this level
- Runs a CLI and reaches a chosen vision provider over the network
- Can reuse signed-in CLIs and spend their quota
Install
In your terminal, with SkillFoxx CLI
npx skillfoxx add skills/modlensDetects the agents on your machine, checks the risk and pins the version.
Other ways to install
Run in a terminal in the project folder
npx skills add liustack/modlens --skill modlens -a claude-code -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Without third-party tools, from commit ffeee3e
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/liustack/modlens.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/modlens/
git -C "$tmp" checkout ffeee3e32de1c05315e334be97c9f087242fe2f3
mkdir -p .claude/skills
cp -R "$tmp/skills/modlens" .claude/skills/modlensCommands for macOS and Linux, on Windows run them in Git Bash.
Run in a terminal in the project folder
npx skills add liustack/modlens --skill modlens -a cursor -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Without third-party tools, from commit ffeee3e
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/liustack/modlens.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/modlens/
git -C "$tmp" checkout ffeee3e32de1c05315e334be97c9f087242fe2f3
mkdir -p .agents/skills
cp -R "$tmp/skills/modlens" .agents/skills/modlensCommands for macOS and Linux, on Windows run them in Git Bash.
Run in a terminal in the project folder
npx skills add liustack/modlens --skill modlens -a github-copilot -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Without third-party tools, from commit ffeee3e
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/liustack/modlens.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/modlens/
git -C "$tmp" checkout ffeee3e32de1c05315e334be97c9f087242fe2f3
mkdir -p .github/skills
cp -R "$tmp/skills/modlens" .github/skills/modlensCommands for macOS and Linux, on Windows run them in Git Bash.
Run in a terminal in the project folder
npx skills add liustack/modlens --skill modlens -a codex -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Without third-party tools, from commit ffeee3e
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/liustack/modlens.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/modlens/
git -C "$tmp" checkout ffeee3e32de1c05315e334be97c9f087242fe2f3
mkdir -p .agents/skills
cp -R "$tmp/skills/modlens" .agents/skills/modlensCommands for macOS and Linux, on Windows run them in Git Bash.
Run in a terminal in the project folder
npx skills add liustack/modlens --skill modlens -a gemini-cli -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Without third-party tools, from commit ffeee3e
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/liustack/modlens.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/modlens/
git -C "$tmp" checkout ffeee3e32de1c05315e334be97c9f087242fe2f3
mkdir -p .agents/skills
cp -R "$tmp/skills/modlens" .agents/skills/modlensCommands for macOS and Linux, on Windows run them in Git Bash.
Run in a terminal in the project folder
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/liustack/modlens.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/modlens/
git -C "$tmp" checkout ffeee3e32de1c05315e334be97c9f087242fe2f3
mkdir -p .devin/skills
cp -R "$tmp/skills/modlens" .devin/skills/modlensCommands for macOS and Linux, on Windows run them in Git Bash.
Formerly Windsurf.
Run in a terminal in the project folder
npx skills add liustack/modlens --skill modlens -a cline -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Without third-party tools, from commit ffeee3e
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/liustack/modlens.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/modlens/
git -C "$tmp" checkout ffeee3e32de1c05315e334be97c9f087242fe2f3
mkdir -p .cline/skills
cp -R "$tmp/skills/modlens" .cline/skills/modlensCommands for macOS and Linux, on Windows run them in Git Bash.
Run in a terminal in the project folder
npx skills add liustack/modlens --skill modlens -a roo -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Without third-party tools, from commit ffeee3e
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/liustack/modlens.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/modlens/
git -C "$tmp" checkout ffeee3e32de1c05315e334be97c9f087242fe2f3
mkdir -p .roo/skills
cp -R "$tmp/skills/modlens" .roo/skills/modlensCommands for macOS and Linux, on Windows run them in Git Bash.
A fork of Roo Code, same .roo folders.
Run in a terminal in the project folder
npx skills add liustack/modlens --skill modlens -a opencode -yThe skills tool installs the current version from the repository. Add the -g flag to use the skill in every project.
Without third-party tools, from commit ffeee3e
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/liustack/modlens.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/modlens/
git -C "$tmp" checkout ffeee3e32de1c05315e334be97c9f087242fe2f3
mkdir -p .agents/skills
cp -R "$tmp/skills/modlens" .agents/skills/modlensCommands for macOS and Linux, on Windows run them in Git Bash.
Run in a terminal in the project folder
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/liustack/modlens.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/modlens/
git -C "$tmp" checkout ffeee3e32de1c05315e334be97c9f087242fe2f3
mkdir -p .agents/skills
cp -R "$tmp/skills/modlens" .agents/skills/modlensCommands for macOS and Linux, on Windows run them in Git Bash.
Run in a terminal in the project folder
tmp=$(mktemp -d)
git clone --filter=blob:none --no-checkout https://github.com/liustack/modlens.git "$tmp"
git -C "$tmp" sparse-checkout set --no-cone /skills/modlens/
git -C "$tmp" checkout ffeee3e32de1c05315e334be97c9f087242fe2f3
mkdir -p .agents/skills
cp -R "$tmp/skills/modlens" .agents/skills/modlensCommands for macOS and Linux, on Windows run them in Git Bash.
Install the skill with npx -y skills add liustack/modlens --skill modlens --global, restart the agent, then configure a vision provider and run the health check with modlens doctor.
Other ways from the author
npx -y skills add liustack/modlens --skill modlens --globalUser-level skill install via skills.sh. After restarting the agent, ask it to configure modlens and run the health check.
This is third-party code. Review the repository files before installing.
What it does
ModLens plugs vision into agents whose main chat runs on a text-only model without multimodality, for example DeepSeek or GLM. When a path or link to an image, a pasted-image placeholder, or a request to read a screenshot appears in the conversation, the skill calls the modlens CLI on its own and returns the result as JSON: full transcription, reading-order layout regions, entities and relations. It draws vision from several sources: a Gemini key, any endpoint speaking the OpenAI chat completions protocol, an Anthropic key, the free Antigravity CLI, and reuse of an already signed-in Claude Code, Codex, OpenCode, Pi or Kimi. Without a pinned provider, every configured one forms a single failover chain, and each attempt is recorded in the response metadata.
Who it is for. For people working in text-only agents who want to read screenshots, diagrams and documents without switching to a multimodal model.
Good fit when
- The model in your agent cannot see images, but you need to read a screenshot or diagram
- You want to paste an image straight into the chat and get transcribed text and layout
- You need a strict, fact-based description of an image rather than the model guessing
Not a fit when
- The model is already multimodal and sees the image itself
- You need web search or page fetching rather than image reading
- No vision provider is available and there is nowhere to set one up
Example request
Read this interface screenshot and list what it says and where each part sitsLimitations
The skill needs at least one working vision source. Some providers are unavailable from Russia without a VPN: Gemini, Anthropic, OpenAI. At the same time any endpoint speaking the OpenAI chat completions protocol works, including reachable platforms such as DashScope with qwen-vl, the GLM platform or SiliconFlow, so availability comes down to provider choice. API keys are your own cost and quota, and reusing a signed-in CLI is flagged in the response warnings so it is clear whose quota was spent. It needs Node 22.19+ or bun and network access.
How to disable. Remove the modlens skill folder from the agent's skills directory. On DeepSeek Harness remove the @liustack/modlens plugin via the dsh plugin command. After that the agent returns to its default behavior.
Security check
- Runs a CLI and reaches a chosen vision provider over the network
- Can reuse signed-in CLIs and spend their quota
README in short
The README describes ModLens as a vision bridge for text-only models and coding agents. It reads an image by path or straight from a chat paste and returns structured JSON with transcription, layout and entities. It installs with one command as a DeepSeek Harness plugin or as a skill via skills.sh in Claude Code, Codex, OpenCode and Pi, and uninstalling is deleting a single folder. Vision comes from six built-in providers and four reusable CLIs that combine into one failover chain. MIT licensed.
SKILL.md
--- name: modlens description: "Plug-in vision for text-only models. Hard rule: when a file path or URL with an image extension (.png, .jpg, .jpeg, .webp, .gif, .heic, .heif) appears anywhere in the conversation (typed by the user, injected as a `[Image: source: <path>]` line, or inside a tag) and you cannot see that image's content, run this skill on it before any other approach: no self-built OCR, no PIL, no tesseract. Also triggers on pasted-image placeholders such as `[Image #1]` and `[Unsupported Image]`. If you can actually see the image, do not use this skill. When unsure, run `modlens guard` before the first read of a session: a deny verdict means the active model has native vision and must read the image itself. Runs the modlens CLI to convert the image into structured JSON evidence: every word transcribed, layout regions, semantics, visual clues. Also use when the user asks how to install, configure, or switch modlens providers (Gemini API key, OpenAI-compatible endpoints, Claude API or Claude Code CLI)." compatibility: Requires network access and one of node 22.19+/npx, bun/bunx, or a preinstalled modlens binary on PATH. allowed-tools: Bash --- # ModLens — Vision Bridge Skill Use this skill when an image is in play and you cannot see its content: a path or URL with an image extension (the path alone is the trigger, hand it to modlens, never Read the bytes or build your own OCR), a placeholder like `[Image #1]`, `[Unsupported Image]`, or a `[Image: source: <path>]` line, or the user asking to configure modlens. Do not use it for web search or fetch (that is `modsearch`), or for images you can already see natively.
FAQ
Which provider to pick in Russia?
The simplest is an endpoint speaking the OpenAI chat completions protocol on a reachable platform, for example DashScope with qwen-vl or the GLM platform. Its key and baseUrl are set through modlens config set.
Do I have to pay for a key?
No. You can take a free Gemini key, install the free Antigravity CLI with no key, or reuse an already signed-in Claude Code or Kimi.
Related
Open-source personal AI assistant on your own machine: answers in Telegram, Slack, Discord and WhatsApp, extended with skills and plugins
A self-improving agent from Nous Research with a TUI, messaging gateway, cron jobs and skills it writes itself
An open source coding agent for the terminal and desktop with build and plan modes
Open prompt library with a Claude Code plugin, MCP server and CLI: search, fetch and improve prompts and skills from an agent