Model or dataset
itmisx/deepx-code avatar
itmisx/deepx-code

deepx-code: a Go coding agent that routes between two DeepSeek tiers

deepseek标配coding agent、原生支持模型路由、CodeGraph代码图谱、OCR截图识别、自动上下文压缩、最佳工作模式选择,workflow等功能,从根本上节省Token

393 stars41 forksGoMIT

At a glance

What is it?
One static binary, no Node or Python runtime, with a built-in code graph, offline OCR through PaddleOCR, and a flash to pro model handoff. The README's own comparison table is careful to say it does not compare model quality.
Who is it for?
This fits a developer who already pays for a DeepSeek style OpenAI compatible endpoint, wants a single binary with no runtime to install, and wants the agent to read a screenshot or jump to a symbol definition without a cloud round trip. It is a poor fit if you need the model itself to be Claude, since the project's own table is explicit that the tradeoff is cost, open source licensing, the single binary, the code graph, and offline OCR rather than model quality.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 11 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One Go binary, and the install is a single curl line

The distribution decision is stated first among the features, because it determines everything else. There is no Node and no Python runtime. The binary covers macOS, Linux, and Windows, and installation on macOS and Linux is one line:

bash
curl -fsSL https://raw.githubusercontent.com/itmisx/deepx-code/main/scripts/install.sh | bash && exec $SHELL

The Windows path uses PowerShell's alias form of the same idea:

powershell
irm https://raw.githubusercontent.com/itmisx/deepx-code/main/scripts/install.ps1 | iex

There is a second install route for users behind a slow connection to GitHub, which pulls the source and the binaries from a Gitee mirror and sets an environment variable in the same command. Once the mirror is used at install time, `deepx upgrade` follows it automatically, which is the detail that makes the mirror worth having rather than a one-off workaround.

The binary lands at `~/.local/bin/deepx` and upgrades with `deepx upgrade`. The `&& exec $SHELL` at the end of the install line is not decoration, since it reloads the shell so the new path is on your PATH without a manual step.

The build configuration backs the claim. A .goreleaser.yaml at the repository root handles cross-platform release builds, and the dependency list in go.mod is a pure Go set with no cgo requirement in the visible portion: Bubble Tea, Lipgloss, and Glamour for the terminal interface, a quickjs binding for running JavaScript workflows, an ONNX runtime binding for the OCR model, a tree-sitter binding for parsing, and a tokenizer binding for counting.

Two model roles, and the routing is what the cost argument rests on

The agent carries a flash role and a pro role rather than one model. A task starts on flash, and complex work is promoted to pro automatically. You can also pin the choice yourself with `/model flash` or `/model pro`, and switch between working modes with `/auto`, `/plan`, and `/review`.

First launch runs a wizard. You pick a provider with the arrow keys, DeepSeek or Xiaomi MiMo among the presets, fill in the API key, and the result is written to `~/.deepx/model.yaml`. The presets ship default flash and pro models and a 1M context window, with DeepSeek on `deepseek-v4-flash` and its pro variant, and MiMo on `mimo-v2.5` and its pro variant. You can edit that file directly and override `base_url`, `model`, `api_key`, `max_tokens`, and `context_window` per role, which means flash and pro do not have to point at the same provider.

That last point is the most flexible thing in the configuration. A cheap small model for routine edits and a larger one for architecture work can sit on two different endpoints, and nothing in the file format prevents it.

Multiple providers are archived by name into `~/.deepx/provider.yaml` on every `/config` run, covering deepseek, mimo, kimi, qwen, and a custom entry. After that, `/provider` switches between the ones already configured by writing the matching flash and pro pair back into `model.yaml`, so you are not re-entering a key each time.

The cache number is one real session, and the README says so

The headline cost claim is a cache hit rate, and the way it is presented is unusual enough to repeat precisely. The tip block reports a measured hit rate of about 99 percent on a long session, giving the actual figures: 41,591 tokens in the session, of which 41,472 were cache hits. DeepSeek bills cached input at a small fraction of the uncached rate, so the argument is that a long running session stops paying full price for the repeated prefix.

Two things are worth separating here. The 99 percent is a property of one session the author ran, not a benchmark over a range of workloads, and the numbers are given alongside it so a reader can judge the scale. A 41K token session is long enough for the prefix to dominate, and a short session would not show the same figure.

The design consequence underneath is that the prompt layout has to be cache-friendly, and the persistence layer is built to match. Sessions are stored as gob binary that keeps `tool_calls`, tool results, and `reasoning_content` intact, so a restart resumes without loss. When the window fills, the session is compressed in layers rather than truncated.

The on-disk layout is explicit about what each file is for. Under `~/.deepx/sessions/<sha1(workspace)[:16]>/` there is `meta.json` for workspace metadata, a `current` pointer where an empty value or `default` means the default conversation for that directory, `state.json` for compression state and a usage snapshot, dated `YYYY-MM-DD.jsonl` files as a text log used for memory search across conversations, `history.gob` for the full default conversation history, and a `conversations/` directory where each `/new` conversation keeps its own history, summary, and state.

The code graph is symbol level, and exact only for Go

The built-in code graph replaces repository-wide grep with structural queries: jump to a definition, find callers, find implementations of an interface, and compute the blast radius of a change. The precision claim is scoped to Go, where the parsing goes through `go/types`, and the README is explicit that this is what the graph is accurate for.

The rest of the repository follows that structure. There is a top-level `codegraph/` directory, a `tui/` directory for the terminal interface, `agent/` for the agent loop, `session/` for persistence, `mcp/` for the Model Context Protocol integration, `ocr/` for the image path, `skill/` for the skill system, `tools/` for tool definitions, `workflow/` for the JavaScript workflow engine, `config/` for configuration, and a `web/` directory alongside an `index.html` at the root.

The dependency list backs up the two headline features. A tree-sitter binding is what makes multi-language parsing possible, and the tokenizer binding is what a context window manager needs. Both are pure Go bindings, which is consistent with the no cgo claim.

Where the graph does not apply, the project names a different tool rather than implying parity. A file or directory can be pulled into context with an `@` prefix: typing `@` in the input opens a fuzzy local path picker, and choosing an entry inserts the `@path` into the message. The model then calls Read for a file or List for a directory, so the content is pulled on demand instead of the whole tree being pasted into the prompt.

Image reading is local, and the OCR runs through an ONNX model in the binary

Dropping a screenshot into the conversation does not send it to a multimodal API. OCR runs locally through PaddleOCR, and the comparison table contrasts this with sending images to a cloud multimodal model. For a workflow where a screenshot of an error or a whiteboard is part of the prompt, that removes both the per-image cost and the data leaving the machine.

The implementation detail is visible in the dependencies. A pure Go ONNX runtime binding appears in the main require block, and a wazero version appears among the indirect dependencies, which is the WebAssembly runtime that binding runs the model on. The whole thing stays inside the single binary, so the feature does not reintroduce a Python environment or an external service.

The file handling around it is broader than OCR. Capture covers photos, documents, videos, forwarded posts, and whole albums, and the stated promise is that nothing sent is ever silently dropped, meaning the agent reads the attachments and works from the takeaways rather than asking you to summarise first.

There is an `ocr/` directory at the top of the tree, and an `assets/` directory beside it, which is where the bundled model weights would live given the single binary constraint.

Three working modes are mutually exclusive on purpose

Methodology is locked with one command rather than blended. Three modes are offered: `karpathy`, described as pragmatic and craft focused, `openspec`, spec driven, and `superpowers`, rigorous end to end. Choosing one disables the other two, and the stated reason is to prevent methodologies from being mixed in the same run.

The switching mechanics are the part that shows care about context. A mode change is saved into the session, and the active mode is injected on every turn rather than appended to the history. That distinction matters for a long session: an injected instruction does not accumulate, while an appended one would keep re-stating itself and would also pollute the cached prefix that the cost argument depends on.

Alongside the modes sit the other structural tools. Multi step work runs against a visible sequential todo list that gets checked off item by item, while independent subtasks that could run in parallel can be split into a DAG and dispatched to concurrent sub-agents.

Reusable workflows go further and are written as JavaScript. Three functions are named for the shapes they express: `agent()` for a single call, `parallel()` for concurrent calls, and `pipeline()` for staged ones. A workflow can be generated and saved with `/ultracode <description>` and run by name with `/workflow <name>`. The stated properties are real concurrency, interruptible resume, structured output enforced through tool constraints, and a preflight that lists all stages before running with elapsed time shown live. The script convention is aligned with Claude Code's, so scripts are interchangeable between the two projects rather than being a private dialect.

The sandbox has three settings and the default uses the operating system

Isolation is not tied to containers. The default setting is `native`, which uses operating system facilities: Seatbelt on macOS and bubblewrap on Linux, with writes limited to the workspace and process separation applied on top. On a platform with no such mechanism the tool falls back to a softer policy rather than failing.

The other two settings are `docker` for container isolation, and `off` to disable it. The stated point is that you can give an agent a boundary without requiring a container runtime, which matters on a laptop where Docker is either not installed or not something you want running for an editor session.

A separate control sits on the actions rather than the environment. Review mode is the default for file writes and shell execution, which require human confirmation, and the feature is described as the thing that keeps the setup safe rather than as an opt-in hardening step.

Extensions come from the same ecosystem in two places. Skills go in `<workspace>/.deepx/skills/`, or an existing directory such as `~/.claude/skills/` can be reused as is, which is what makes the skill ecosystem compatible rather than proprietary. MCP servers are managed from inside the TUI, with `/mcp-add` to add one and `/mcp-list` to see what is configured.

The non-interactive path is the piece that turns this from a terminal toy into something scriptable. `deepx exec` runs a task to completion, prints only the result to stdout, and exits, with no TUI and no intermediate output shown. Piping data in is supported, so an error log can be fed to it directly. Output can be redirected, and the documentation points at scripts, CI, and cron as the intended homes for it.

Editorial conclusion

This fits a developer who already pays for a DeepSeek style OpenAI compatible endpoint, wants a single binary with no runtime to install, and wants the agent to read a screenshot or jump to a symbol definition without a cloud round trip. It is a poor fit if you need the model itself to be Claude, since the project's own table is explicit that the tradeoff is cost, open source licensing, the single binary, the code graph, and offline OCR rather than model quality. Before you install, check three things: which provider and key you will configure on first launch, since the wizard writes them to `~/.deepx/model.yaml`; which working mode you want, because `karpathy`, `openspec`, and `superpowers` are mutually exclusive and switching is saved per session; and whether the default `native` sandbox is what you want, given it uses Seatbelt on macOS and bubblewrap on Linux. The last push was 2026-09-21, the same day as the v0.2.111 through v0.2.113 tags.

Frequently asked questions

What is deepx-code used for?

It is a terminal coding agent for DeepSeek and other OpenAI compatible providers, packaged as a single Go binary with no Node or Python runtime. It adds a built-in code graph, local OCR, a flash to pro model handoff, and an OS level sandbox around the agent's file and shell actions.

Is deepx-code free to use?

The code is MIT licensed and the comparison table lists open source as the advantage over the closed source alternative it measures itself against. Running it still needs an API key for a provider, since the first launch wizard writes your key to ~/.deepx/model.yaml and the project describes cost in terms of cached input pricing on that provider.

How does deepx-code choose between its two models?

A task starts on the flash role and complex work is promoted to pro automatically, which is called dual model auto routing. You can pin the choice with /model flash or /model pro, and the two roles can be pointed at different providers and overridden per role for base_url, model, api_key, max_tokens, and context_window.

Does deepx-code send screenshots to a multimodal API?

No. Image text is read locally with PaddleOCR, and the ONNX runtime binding is part of the Go binary, so a dropped screenshot is read offline. The project's comparison table lists local offline OCR as the difference from sending images to a cloud model.

Official sources

  1. itmisx/deepx-code on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/itmisx-deepx-code.svg)](https://hysenlabs.com/projects/itmisx-deepx-code)