DSH Vision Toolkit
Vision tools and a skill for DeepSeek Harness: image Q&A, long-screenshot OCR and UI restoration
Medium risk
We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.
Why this level
- Runs code and calls external vision model services
- Reads supplied screenshots and project files
Install
Manual install
dsh plugin --profile web add @anionex/dsh-vision-toolkitCommand for the DeepSeek Harness web profile, after install set a vision model provider in settings.
This is third-party code. Review the repository files before installing.
What it does
The plugin gives text-only models in DeepSeek Harness the ability to work with images. A pasted image is analysed against the task at hand: where the error is, where a button sits, what a long screenshot contains. The toolkit covers OCR, locating and cropping elements, pixel-level image diff and rebuilding an interface from a mockup into code. A bundled skill tells the agent what to look at for each task and how to verify the result. It installs with one command into a DeepSeek Harness profile.
Who it is for. For developers and designers who work in DeepSeek Harness with a text-only model but need to read screenshots and mockups.
Good fit when
- You need to pull text from a long screenshot
- You need to understand what is wrong on a UI screenshot
- You need to rebuild a layout from a mockup or image
Not a fit when
- You work in a different agent, not DeepSeek Harness
- You have no vision model access or no way to pay a provider
Example request
Look at this screenshot and tell me exactly where the error shows upLimitations
The toolkit targets DeepSeek Harness and its profiles and does not work outside that environment. It needs a vision model provider and a key, and some features rely on third-party services. Part of the documentation is in Chinese.
How to disable. Remove the plugin with dsh plugin remove for the relevant profile.
Security check
- Runs code and calls external vision model services
- Reads supplied screenshots and project files
README in short
The README presents the toolkit as the first large vision plugin in the DeepSeek Harness ecosystem. It integrates tools from the agent-vision-toolkit project directly into profiles, sessions, settings and the web UI. Pasting an image automatically switches the text-only model to its vision variant. Install is one command, then a vision provider is configured. The docs include OCR, UI analysis and layout restoration examples.
FAQ
Which agents does it work with?
It is built for DeepSeek Harness and its web and headless profiles. It does not run outside that environment.
Do I need a separate key?
Yes, you need a vision model provider. The README lists several options, some with free models.
Related
A skills library that gives coding agents a development process: brainstorming, planning, TDD, subagents and code review
Skills for real engineers by Matt Pocock
Skills For Real Engineers
Small composable skills for engineering with agents: plan grilling, TDD, bug diagnosis, code review and architecture
GitHub toolkit for spec-driven development: the specify CLI adds agent commands and skills to a project, from principles to implementation
Reference MCP servers
Model Context Protocol servers
Official reference MCP servers: Filesystem, Fetch, Git, Memory, Sequential Thinking, Time and Everything