Open-source project
Larens94/codedna avatar
Larens94/codedna

CodeDNA: An In-Source Annotation Standard for Faster AI Codebase Navigation

A lightweight annotation standard that helps AI agents navigate codebases faster, with fewer file reads and tool calls

151 stars24 forksPythonMIT

At a glance

What is it?
CodeDNA embeds architectural context directly into source files, letting AI agents navigate with fewer reads and tool calls. This review covers its mechanism, setup, limitations, and alternatives.
Who is it for?
Adopt CodeDNA if you work with AI agents that frequently read and edit a shared codebase, especially in teams with mixed tools like Claude Code, Codex, and Aider. Avoid it if your team does not use AI agents regularly, or if you cannot enforce the pre-commit hook and annotation discipline; the protocol only helps if files stay annotated.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Problem: AI Agents Waste Reads on Unfamiliar Code

When an AI agent starts a session on a codebase it has never seen, it cannot know which files matter. It reads broadly, opens files that turn out to be irrelevant, and makes tool calls to search for symbols that are not where it expects. The README frames this as a cost problem: too many file reads and tool calls before the agent can act. CodeDNA targets exactly that. The intended user is a developer or team that runs AI coding agents regularly, across different models or tools, and wants those agents to spend less time exploring and more time editing. The project's claim is that embedding context in the files themselves, rather than in an external index, gives the next agent what it needs on the first read.

The Mechanism: Headers as a Communication Protocol

CodeDNA works by adding structured headers to the top of each source file. These headers carry fields like `exports:` and `used_by:`, which are derived from the code's actual structure via AST or tree-sitter. The `rules:` field is optional and can be generated by an LLM when you run `codedna init --model`. The key idea is that the header is the channel: an agent reads the first 10 to 15 lines of a file and gets the architectural context without reading the whole file. The `codedna manifest` command builds a project-level map called Level 0, which lists packages and dependencies. This is a two-tier design: per-file headers for immediate context, and a project map for global orientation. The README calls this "an in-source communication protocol" and stresses that no infrastructure or retrieval pipeline is needed. The code itself carries its own context.

Installation and First Commands

The CLI install is straightforward if you have Python 3.11 or newer. The README shows `pipx install git+https://github.com/Larens94/codedna.git`, then `codedna install --path . --tools codex` to set up the integration for a specific agent. The `--tools` flag accepts values like `claude`, `codex`, `opencode`, `aider`, and more. Multiple tools can be passed at once, for example `--tools claude codex opencode`. The install step creates a `.codedna` directory, adds a Git pre-commit hook, and writes the agent-specific instruction file such as `CLAUDE.md` or `AGENTS.md`. It claims to preserve existing instruction files and hooks. After install, you run `codedna init . --no-llm` for a free structural pass, or `codedna init . --model deepseek/deepseek-chat` to add LLM-generated rules, which the README estimates at about $0.40 per 200 files. There is also a local option with `--model ollama/llama3`. The `codedna doctor` command checks that everything is set up correctly.

The Daily Workflow: Commands That Keep Headers Honest

The protocol only works if the headers stay accurate. CodeDNA provides a set of commands to keep them in sync. `codedna refresh` recalculates `exports:` and `used_by:` via AST or tree-sitter, and it costs no LLM tokens. It preserves `rules:` and `agent:` fields. `codedna verify` detects stale references and is read-only, with a `--json` flag for CI. `codedna check` produces a coverage report and exits with code 1 if any files are incomplete, which makes it a natural CI gate. `codedna impact <file>` shows the transitive callers of a file before you edit it, which is useful when you are about to change a public contract. The README suggests running `codedna impact` before changing a public contract and `codedna verify` after structural edits. This workflow is concrete and testable, and it is the part of the project that feels most like a real engineering practice rather than a marketing promise.

Multi-Language and Agent Integration: The Practical Table

CodeDNA does not lock you into one editor or agent. The README includes a table of supported tools with the exact `--tools` value and what gets installed. For Claude Code it installs `CLAUDE.md` plus active hooks. For Codex it installs a cross-vendor `AGENTS.md`. Aider users must start with `aider --read AGENTS.md` or configure `read: AGENTS.md`, which is a manual step. Cursor gets `.cursorrules`, and Copilot gets its own instruction file. Windsurf only gets instructions, no hooks. This variety means the same annotation standard works across a mixed team, which is a real advantage if your team already uses different agents. The language support is broad: Python, PHP, TypeScript/JavaScript, Go, Java, Kotlin, Ruby, Rust, C#, Swift, and supported templates. The format adapts to the language, so PHP uses `//` comments, Python uses docstrings, and Blade uses `{{-- --}}`. That is a thoughtful touch, but it also means the parser must handle each language correctly, and any gap will show up as missing annotations.

Limitations and Failure Modes

The biggest limitation is that the protocol depends on discipline. If a developer edits a file and does not run `codedna refresh` or `codedna verify`, the headers go stale. The pre-commit hook helps, but it only runs on commit, and not every edit goes through a commit. The README does not describe an automatic mechanism to detect staleness outside of `verify`, so the burden is on the team to run it. Another limitation is the LLM cost. The `--model` path is not free, and the README gives a rough figure of $0.40 per 200 files, which scales linearly with file count. For a large monorepo with tens of thousands of files, that cost could become significant. The no-LLM path is free, but it only adds structural fields like `exports:` and `used_by:`, not the semantic `rules:` that might be more useful for an agent. Finally, the project is very new. The only release is v1.0.0 from March 2026, and the README claims performance gains like +17pp F1 on SWE-bench, but those numbers come from the project's own evidence section, not from independent verification. Treat them as claims, not facts.

Alternatives: In-Source vs. External Context

The main alternative to CodeDNA is an external context system, such as a retrieval-augmented generation (RAG) pipeline that indexes your codebase into a vector database. Tools like LlamaIndex or custom embeddings allow an agent to query for relevant files semantically without modifying the source. The difference is fundamental: CodeDNA puts the context inside the files, so the agent reads a header and gets the exact dependency information. A RAG system requires building and maintaining an index, which is infrastructure that can drift from the code. CodeDNA's approach has the advantage of zero infrastructure, but it only works if the annotations are present. Another alternative is the agent's native memory, such as Claude Code's CLAUDE.md or Cursor's rules, which give global instructions but do not provide per-file dependency data. CodeDNA is more granular. The trade-off is that CodeDNA requires a one-time annotation pass and ongoing maintenance, while a RAG pipeline requires ongoing indexing but no source changes. For a small codebase, the annotation overhead is trivial; for a huge one, the LLM cost and the risk of stale headers might make RAG more attractive.

Maintenance, License, and the Bottom Line

Maintenance cost is real. The CLI has a `self-update` command that reinstalls via pip, and the `refresh` command is designed to keep headers accurate without LLM cost. The project ships a CI workflow (the README links to `actions/workflows/ci.yml`), which suggests the maintainers run tests. The license is MIT, which is permissive and allows commercial use, modification, and redistribution. That is a low-license-risk choice for most teams. However, the project is young: one release, one maintainer, and a single push date. That means you should expect API changes and potential bugs. The README's evidence section includes specific metrics like "1.6x team velocity" and "98.2% protocol adoption" but these are self-reported and lack methodology details. The core idea is sound: putting architectural context in the source reduces the need for agents to explore. Whether it works for you depends on your ability to keep the annotations current. If you can enforce `codedna verify` in CI and run `codedna refresh` after major refactors, the protocol could genuinely cut down agent tool calls. If you cannot, the headers will rot and the agents will ignore them.

Editorial conclusion

Adopt CodeDNA if you work with AI agents that frequently read and edit a shared codebase, especially in teams with mixed tools like Claude Code, Codex, and Aider. Avoid it if your team does not use AI agents regularly, or if you cannot enforce the pre-commit hook and annotation discipline; the protocol only helps if files stay annotated. Before adopting, verify that the CLI supports your language (Python, PHP, TypeScript, Go, Java, Kotlin, Ruby, Rust, C#, Swift, and templates) and that your CI can run `codedna check` and `codedna verify` as gates. Also confirm the LLM cost model: the `--model` path charges about $0.40 per 200 files, but the `--no-llm` structural pass is free. Test the `codedna impact` command on a critical file to see if the dependency chain is accurate for your codebase.

Frequently asked questions

Does CodeDNA need a model API key to annotate a repository?

No. `codedna init . --no-llm` runs a structural pass over exports and used_by only and needs no key. Passing a model adds the rules block: `codedna init . --model deepseek/deepseek-chat` is priced at roughly $0.40 for 200 files, and `codedna init . --model ollama/llama3` runs it locally for free.

Which files does codedna install write into my repository?

It creates a `.codedna` directory, installs a Git pre-commit gate, and writes the instruction file for the agent named by `--tools`: `CLAUDE.md` for Claude Code, `AGENTS.md` for Codex, OpenCode, Aider, and Antigravity, `.cursorrules` for Cursor, `.windsurfrules` for Windsurf, `.roorules` for Roo Code. Existing instruction files and Git hooks are preserved rather than overwritten.

How do I run CodeDNA with Claude Code instead of the CLI?

Add the marketplace and install the plugin with `claude plugin marketplace add Larens94/codedna` followed by `claude plugin install codedna@codedna`, start a new session or run `/clear`, then run `/codedna:init`. The `/codedna:*` slash commands belong to the plugin, and every other agent uses the `codedna ...` CLI commands.

Which codedna commands cost money after the first annotation?

None of them. `codedna refresh <path>` recalculates exports and used_by through tree-sitter at zero LLM cost, and `codedna check <path>`, `codedna verify <path>` with `--json`, and `codedna impact <file-or-symbol>` are all read only. The model call happens in `codedna init` when `--model` is passed.

Which languages does CodeDNA auto-detect and how does it annotate them?

Python, PHP, TypeScript/JavaScript, Go, Java, Kotlin, Ruby, Rust, C#, VB.NET, Swift, and supported templates. The marker adapts per language, with `//` for PHP, docstrings for Python, and `{{-- --}}` for Blade, and the per-language details live in docs/languages.md.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/larens94-codedna.svg)](https://hysenlabs.com/projects/larens94-codedna)