# ModLens: a vision plugin that gives text-only coding agents sight

> ModLens is a TypeScript plugin for DeepSeek Harness and a skill for Claude Code, Codex, Pi and OpenCode that turns a pasted image into structured JSON evidence for models that cannot see. The install is one command, but the engine behind it decides how much setup you actually do.

**liustack/modlens** — The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件，为 DeepSeek、GLM 等纯文本模型外挂视觉能力，粘贴图片即得结构化 JSON 证据（OCR、版面、语义）。

- Repository: https://github.com/liustack/modlens
- Website: https://liustack.dev/tools/modlens
- Stars: 4,100 · Forks: 126
- Language: TypeScript
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/liustack-modlens

## The gap ModLens fills: text-only models in an image-heavy workflow

DeepSeek's flagship chat models and GLM-5.3 are text-only, as the README states plainly. Paste a screenshot of a stack trace or a scanned invoice into one of them and nothing useful comes back. The usual workaround is to save the image to disk, run a separate OCR tool, and paste the text into the chat. ModLens removes the middle step: the image goes into the composer, and the plugin turns it into structured evidence before the model answers.

The audience is narrow and specific. You are already using DeepSeek Harness, or one of the skill harnesses the README names (Claude Code, Codex, Pi, OpenCode), and you want image input without switching to a multimodal model. The README is explicit that the plugin targets text-only routes and excludes native vision models in the same families, including GLM-5.3-Flash. If your model already sees images, ModLens has nothing to add.

## How the vision bridge works: paste, route, structured evidence

There are two paste paths, and the difference matters. In the first, a pasted image lands as a private temp file and its path enters the composer, which the README compares to the interaction OpenCode and Pi ship. The `modlens_read_image` tool then reads that path. In the second, you pick a `(modlens vision)` entry in the model selector, the choice is remembered, and the image is converted to structured evidence at request time while the thumbnail stays visible in your message.

The plugin auto-discovers provider routes that carry eligible text-only DeepSeek, GLM or MiMo Pro models and adds a wrapped entry per route. A stock install gets `DeepSeek-V4-Flash (modlens vision)` and `DeepSeek-V4-Pro (modlens vision)`; extra routes such as opencode-go or zai get their own entries. The takeover rule is conservative: only a model whose metadata positively confirms it is text-only is wrapped, and anything unconfirmed is left alone. That is why GLM-5.3-Flash keeps its native paste.

The output is the part that separates ModLens from a caption generator. The README describes full transcription, reading-order layout regions, and entity and relation lists, so the model can quote specifics instead of paraphrasing a scene. That is the right shape for invoices, error screenshots and diagrams, and the wrong shape for questions like "does this photo look good."

## Installing ModLens in DeepSeek Harness and in skill harnesses

On DeepSeek Harness the README gives a single plugin command, pinned to the current release. Run it and restart the harness.

```bash
npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.26.2
```

For the skill harnesses, the README offers skills.sh at user level. This places the `modlens` skill folder where the harness can find it.

```bash
npx -y skills add liustack/modlens --skill modlens --global
```

After either install, the README says to restart the harness, ask your AI to configure modlens, and run its health check. The health check is the real first use: it reports whether an engine is reachable and which one. The README notes that an existing login in Claude Code, Codex, OpenCode or Pi can be enough, and that modlens asks before reusing any of them.

Only if the health check comes back empty should you add an engine. The recommended route is a free Gemini API key from Google AI Studio, which the README says takes about three minutes and makes every read 5 to 10 seconds. To avoid signing up anywhere, the alternative is Antigravity CLI.

```bash
curl -fsSL https://antigravity.google/cli/install.sh | bash
agy
```

The first line installs the CLI; the second signs you in and you exit. On DSH there is also a GUI path: Settings, then Plugins, then Plugin config, which carries a ModLens card for switching the engine.

## Where ModLens stops: engine dependency and the models it refuses to touch

The plugin is a bridge, not a model. It carries no vision weights of its own, so a skill folder with no reachable engine reads nothing. The README is honest about this ordering: install, run the health check, and only set up a free engine if the check comes back empty. Anyone expecting a self-contained OCR binary will be disappointed.

Second, the takeover rule cuts both ways. Because only positively confirmed text-only models are wrapped, an unconfirmed route is left with its native paste behavior. If you were hoping ModLens would intercept a model it does not recognize, it will not, and the README points to `docs/harness-setup.md` for the per-model details rather than promising blanket coverage.

Third, the free-engine path has a quota dimension. Reused logins join the engine pool as equals, and the README says every reused read is labeled with whose quota it spent. That labeling is useful, but it does not change the fact that heavy image work draws down someone's account. Teams that paste images all day should decide in advance whether that someone is a shared login or a dedicated key.

## ModLens compared with pointing the agent at an OCR CLI

The obvious alternative is to skip the plugin and shell out to a standalone OCR tool, then paste the text. The difference is in what reaches the model. A raw OCR dump loses reading order and layout, so a two-column invoice or a table becomes a wall of lines with no structure. ModLens returns layout regions and entity and relation lists, which gives the model something to reason over rather than something to re-parse.

The second difference is integration cost. A manual OCR step has no install footprint but adds a step to every image task, and it lives outside the harness. ModLens installs as one plugin on dsh or one skill folder elsewhere, and the README states that uninstalling is deleting a folder with no harness config changed. That claim is testable by inspection: the README says no hooks, no wrappers, no local proxy daemon. For a team that wants the capability without a permanent piece of infrastructure, that is the trade being offered.

## Release cadence, licence, and what upgrading costs you

The repository is not archived, and the last push was on 2026-09-18, five days before this writing. Three releases landed in September 2026: v3.26.0 on 2026-09-06, v3.26.1 on 2026-09-07, and v3.26.2 on 2026-09-18. That is a fast patch rhythm around a single minor line, which suggests active iteration but also means the version pin in your install command goes stale quickly.

The dsh install command pins `@liustack/modlens@3.26.2` explicitly, so upgrading means changing that version string rather than relying on a floating tag. The README points to `docs/harness-setup.md` for update details. The package requires Node.js 22.19 or newer, which is worth checking before you start, since an older runtime will fail before the health check ever runs.

The licence is MIT. That permits commercial use and modification, but it also means no warranty and no support obligation from the author. Nothing here is legal advice; if you redistribute ModLens inside a product, read the LICENSE file in the repository rather than this summary.

## Conclusion

ModLens fits teams already running a text-only DeepSeek, GLM or MiMo Pro model inside DSH, Claude Code, Codex, Pi or OpenCode who need OCR and layout evidence rather than a caption. Skip it if your model already accepts images natively, since the plugin deliberately leaves confirmed vision models alone, or if you cannot point it at any vision engine, because the skill folder alone reads nothing. Before adopting, run the health check the README describes and confirm which engine it selected, then check that the reused login it picked is one you are willing to spend quota on.

## FAQ

### What is ModLens and which models does it work with?

ModLens is a vision plugin that gives text-only models sight, returning structured JSON evidence such as OCR, layout and semantics from a pasted image. The README says it auto-discovers routes carrying eligible text-only DeepSeek, GLM or MiMo Pro models, and excludes native vision models in those families.

### How do I install ModLens in DeepSeek Harness?

The README gives one command: npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.26.2. After that, restart the harness, ask your AI to configure modlens, and run the health check.

### Does ModLens need an API key or a separate vision engine?

It needs a reachable engine, and the README says an existing login in Claude Code, Codex, OpenCode or Pi can be enough, with modlens asking before reusing any of them. Only if the health check comes back empty should you add a free Gemini key or install Antigravity CLI.

### What happens to images pasted into a model that already supports vision?

Nothing. The README states that only a model whose metadata positively confirms it is text-only is taken over, and anything unconfirmed is left alone, so native vision models such as GLM-5.3-Flash keep their native paste behavior.

## Sources

- [License: MIT](https://github.com/liustack/modlens/blob/main/LICENSE)
- [liustack/modlens on GitHub](https://github.com/liustack/modlens)
- [Project website](https://liustack.dev/tools/modlens)
- [README](https://github.com/liustack/modlens/blob/main/README.md)
- [Releases](https://github.com/liustack/modlens/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/liustack-modlens
