# agent-vision-toolkit: giving text-only coding agents a way to see images

> A Python toolkit, agent skill and optional local proxy that route pasted images and built-in image tools to a multimodal model, so DeepSeek-class text-only agents can answer image questions, OCR long screenshots and rebuild frontend UIs.

**Anionex/agent-vision-toolkit** — 为纯文本模型"看图“设计更好的视觉工具箱和技能，支持多图理解，图片问答，前端UI还原、GUI 自动化等，并可选无缝接入多个主流agent，直接识别粘贴图片｜ A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

- Repository: https://github.com/Anionex/agent-vision-toolkit
- Website: https://agent-vision.anionex.me
- Stars: 1,217 · Forks: 48
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/anionex-agent-vision-toolkit

## The gap agent-vision-toolkit fills for text-only models

Coding agents built on text-only models such as DeepSeek can read a repository but not a screenshot. The README describes the failure mode directly: the agent is "unable to see images, with every attempt to use an image tool blocked by the system". Pasting a design mockup or an error dialog into the chat produces nothing useful.

The project's answer is to move vision out of the model and into the harness. Its stated position is that "an agent's vision capability doesn't have to live in the model, it can live in the harness". Concretely, the toolkit ships vision tool CLIs, an agent skill named vision-skills that tells the agent when to call each tool, and an optional local proxy. The intended audience is narrow and clear: people running Codex, Claude Code, Pi, Oh My Pi or OpenCode against a text-only model who still want image handling inside the same session.

## Two layers: shell-invokable CLIs and a transparent proxy

The repository is split into two components. The first is the vision tool CLIs plus the skill. Any agent that can invoke a shell can use them, which keeps the integration surface small. The README states the skill teaches the agent "what to inspect, which tool to choose, what sequence to follow, and how to verify the final result", so the skill is the decision layer rather than a wrapper around a single call.

The second layer is the optional upgrade: a transparent local proxy and single-file native plugins. The README's claim is that with it, "images we paste and the agent's built-in image tools both work seamlessly". The proxy is the piece that matters for pasted images, because it sits between the agent and its model provider. The .env.example confirms the split: only the vision API is configured there, while "DeepSeek auth is still sent by Codex and passed through by the proxy". So the proxy forwards the text-model credentials it receives and swaps in a separate vision endpoint for image content.

A related package, dsh-vision-toolkit, is tracked as a Git submodule and maintained separately. It brings the toolkit into DSH Web and Headless profiles as a native Profile Bundle, and the release note lists ten structured visual tools: intent-aware image Q&A, grounding, detection, tracing, cropping, pixel diff, long-screenshot OCR, foreground extraction, dominant-color analysis and HTML screenshots. It also adds DSH Credentials, a managed isolated runtime, previewable Artifacts, Web Settings and agent-scoped progressive tool exposure.

## Installing agent-vision-toolkit and running a first image question

The README advertises a one-sentence install: ask your agent to install it, and it follows AGENT_INSTALL.md. The repository also keeps a bin/ directory of shell entry points, so manual use is possible without an agent in the loop.

Because dsh-vision-toolkit is a submodule, a plain clone leaves that directory empty. The README gives two ways to fix it:

```bash
git clone --recurse-submodules https://github.com/Anionex/agent-vision-toolkit.git
```

Or, in a checkout you already have:

```bash
git submodule update --init --recursive
```

Next, create your .env from the example and fill in the vision side. Only three values are required to start:

```bash
VISION_API_KEY=
VISION_BASE_URL=https://openrouter.ai/api/v1
VISION_MODEL=google/gemini-3.6-flash
```

The example file recommends the Gemini Flash series and notes that a capable local model also works. It also lists Aliyun DashScope as an alternative base URL, https://dashscope.aliyuncs.com/compatible-mode/v1. LANG controls the output language, zh or en, and defaults to Chinese when unset, so an English-speaking user should set LANG=en explicitly. With the proxy running, the intended result is that a pasted image is answered in the same conversation where the text-only model is already working.

## Protocol choices and where configuration goes wrong

The client and proxy speak one of three protocols: chat_completions, which is the default, responses, or anthropic. The .env.example is unusually explicit about the failure mode here. For anthropic, the base URL must end in /v1, not /messages. That is the kind of mistake that produces a confusing error rather than a clear one.

Two optional settings carry similar warnings. VISION_REASONING_EFFORT applies to the responses protocol. VISION_ANTHROPIC_THINKING defaults to omit, which sends no thinking field and is described as having the broadest compatibility; disabled or adaptive should only be used when the selected model documents that mode, and the file says to restore omit first if the provider returns HTTP 400. There is also VISION_USER_AGENT, defaulting to a browser-compatible string to avoid gateways that block Python-urllib clients. That default is a quiet admission that some providers reject the standard Python client out of hand.

This is the part of the toolkit with the most moving parts, and it is also the part the README documents least in prose. The .env.example is effectively the reference. Treat it as the source of truth rather than the README's summary.

## What the toolkit does not do

The toolkit does not make a text-only model multimodal. It inserts a second model, reached through VISION_API_KEY and VISION_BASE_URL, to handle visual content. If you have no vision endpoint available, or your organisation will not approve a second provider, the whole approach collapses. That is a real constraint, not a footnote.

The proxy is a local process sitting between the agent and its provider. Anything that depends on the agent talking directly to the provider, such as a corporate gateway that pins the User-Agent or terminates TLS, may not survive that insertion. The README does not document rollback if the proxy misbehaves, and it does not describe what happens to a session when the vision endpoint rate-limits or returns an error mid-turn. Those are the questions to answer before putting this in front of a team.

Finally, the project's own framing is a bet on the harness. The README states the goal is for "a tool-equipped text-model agent" to outperform "a native multimodal agent that does not use this toolkit and its methods". That is a comparison the repository asserts, not one it demonstrates with published measurements.

## How it compares to just using a multimodal model

The obvious alternative is to run the agent on a multimodal model in the first place and skip the toolkit entirely. That removes the proxy, the second API key and the protocol negotiation. It also removes the need for a skill that decides which vision tool to call, because the model sees the image directly.

The difference in approach is where the intelligence sits. A native multimodal agent does perception inside the model. agent-vision-toolkit does it outside, in shell-invokable tools plus a skill that sequences them. The project's argument is that the outside approach can be better targeted, because the toolkit passes "the user's or model's latest intent" along with the image and returns "the details needed for the current turn instead of a broad, unfocused description".

That argument holds up best for structured visual work: long-screenshot OCR, grounding, cropping, pixel diff, frontend UI restoration. For a one-off "what is in this picture" question, the extra machinery is overhead. The toolkit earns its place when the task has a shape the skill can recognise, not when the image is incidental.

## Licence, maintenance and upgrade cost

The project is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. The repository also carries FUNDING.md and a sponsor section, and the README asks for stars and forks. None of that changes the licence terms. As always, the licence text in LICENSE governs, not the README's summary, and this is not legal advice.

The last push was on 2026-08-27, and the repository is not archived. Two releases exist: v0.1.0 on 2026-08-06 and v0.2.0 on 2026-08-14. The changelog shows a rename in that window, with the skill moving from vision-tools to vision-skills on 2026-08-18, and the addition of native DeepSeek Harness support through the dsh-vision-toolkit submodule on 2026-08-13. Anyone with a pre-rename checkout should expect the old skill name to be stale.

Upgrade cost is concentrated in the submodule and the .env file. Because dsh-vision-toolkit is maintained independently and tracked as a submodule, pulling the parent repository does not update it; you need git submodule update --init --recursive. Configuration keys have also moved between releases, so diff your .env against .env.example after each upgrade rather than assuming your existing keys are still read.

## Conclusion

Adopt it if your agent already runs on a text-only model such as DeepSeek and you need image Q&A, long-screenshot OCR or frontend UI restoration without switching models. Skip it if you already run a multimodal model natively, or if you cannot supply a separate vision endpoint, since the toolkit adds a second provider and a local proxy process. Before wiring it into a team workflow, clone with --recurse-submodules so dsh-vision-toolkit is present, then set VISION_API_KEY, VISION_BASE_URL and VISION_MODEL in .env and confirm the chosen base URL matches the protocol you select.

## FAQ

### Does agent-vision-toolkit work with Codex?

Yes. The README states the code has been verified in real Codex plus DeepSeek sessions, and the seamless integration layer includes a transparent local proxy with native plugins for Codex. The proxy passes through the DeepSeek auth that Codex sends and uses a separate vision API for image content.

### Which model do I need for agent-vision-toolkit?

Two models are involved. The agent keeps running on its text-only model, while VISION_MODEL points at a multimodal model; the .env.example recommends the Gemini Flash series and notes a capable local model also works. The example default is google/gemini-3.6-flash.

### Why is the dsh-vision-toolkit directory empty after cloning agent-vision-toolkit?

Because dsh-vision-toolkit is tracked as a Git submodule and is maintained independently. The README says to clone with --recurse-submodules, or to run git submodule update --init --recursive in an existing checkout.

### What is agent-vision-toolkit in DSH?

The README describes dsh-vision-toolkit as a linked package that brings this toolkit into DSH Web and Headless profiles as a native Profile Bundle, providing ten structured visual tools plus DSH Credentials, a managed isolated runtime, previewable Artifacts, Web Settings and agent-scoped progressive tool exposure.

## Sources

- [Anionex/agent-vision-toolkit on GitHub](https://github.com/Anionex/agent-vision-toolkit)
- [License: MIT](https://github.com/Anionex/agent-vision-toolkit/blob/main/LICENSE)
- [Project website](https://agent-vision.anionex.me)
- [README](https://github.com/Anionex/agent-vision-toolkit/blob/main/README.md)
- [Releases](https://github.com/Anionex/agent-vision-toolkit/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/anionex-agent-vision-toolkit
