codesight claims fourteen languages and parses two of them properly
Universal AI context generator. Saves thousands of tokens per conversation in Claude Code, Cursor, Copilot, Codex, and more.
At a glance
- What is it?
- A zero-dependency Node tool that maps a codebase so an AI assistant does not have to rediscover it each session, with a generated wiki and fourteen MCP tools. The AST path covers TypeScript and JavaScript; the other twelve languages are handled by pattern matching.
- Who is it for?
- codesight is worth running on a TypeScript or JavaScript repository, where the generated wiki is derived from real structure rather than guesses, and the token saving on a large codebase is substantial. Set your expectations by language before you measure anything.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 68 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Fourteen languages are supported and two of them are parsed
The headline lists TypeScript, JavaScript, Python, Go, Ruby, Elixir, Java, Kotlin, Rust, PHP, Dart, Swift, C# and BrightScript, which is the scripting language used on Roku devices. Then it says which of them get real analysis:
TypeScript projects get full AST precision. Everything else uses battle-tested regex detection across the same 30+ frameworks.That single sentence is the most important line in the project. On a TypeScript or JavaScript repository the tool builds a syntax tree and reads routes, components, models and relationships out of it. On a Python repository it looks for the patterns those thirty-odd framework integrations produce, which finds the routes and the models in the common cases and will miss the unusual ones.
Nothing about the output format changes between the two. A wiki article looks identical whether it was derived from a tree or from a pattern, which means a reader cannot tell from the article which one they got. That is the practical hazard: check the language of your repository against the first line of the README before trusting a generated article about behaviour rather than about file locations.
The remaining counts follow from the same design: thirty-odd framework detectors, fourteen ORM parsers, fourteen MCP tools, and a stated test basis of twenty-five or more open source projects.
One npx call, no keys, and an init that writes four config files
The install story is one command with nothing to configure:
npx codesightNo config file, no setup step, no API keys. That is a real advantage over tools that need a key before they will tell you anything, and it follows from the design: the analysis is local and static, so there is nothing to authenticate against.
The flags are where the decisions are. The initialisation flag is the one to look at closely, because it writes four separate agent configuration files at once:
npx codesight --init # Generate CLAUDE.md, .cursorrules, codex.md, AGENTS.mdRunning it means committing to the conventions of four different tools simultaneously, whether or not your team uses all four. There is a per-tool alternative, a profile flag that generates an optimised configuration for one named assistant, and a flag that opens an interactive report in a browser.
Two others are worth knowing before you need them: a blast radius flag that takes a file path and reports what depends on it, and a mode flag that starts the tool as an MCP server exposing fourteen tools, which is how it is meant to be used by an assistant rather than run by hand.
The wiki replaces a full reload with a two hundred token index
The knowledge base is the part that changes how a session starts. The generated directory holds an index, an overview, one article per domain, and a log:
index.md — catalog of all articles (~200 tokens) — read this at session startThe index is the entry point. An assistant reads that short catalog, then pulls exactly one article instead of loading everything. The comparison the project gives is stark:
| "How does auth work?" | ~12K tokens (reads 8+ files) | ~300 tokens (`auth.md`) |
| New session start | ~5K tokens (full reload) | ~200 tokens (`index.md`) |The mechanism is an assistant that is told to read the index first and then request a specific article, which is why three MCP tools exist for it: one to fetch the index, one to fetch a named article, and one that health-checks the wiki for orphan articles, missing cross-links and stale content.
Two operational details decide whether this stays true. The directory is committed to git, so every session on every machine sees the same knowledge, and it only stays correct if it is regenerated. A watch mode keeps it current while you code, and a hook mode regenerates on every commit.
That health check is the tool admitting the failure mode: a wiki nobody regenerates is a wiki that quietly lies.
Compiled from a syntax tree, not from a model, in two hundred milliseconds
The design decision that makes the tool predictable is that nothing here calls a model. The project positions itself against a known pattern for building a code knowledge base with an LLM and says the difference is that this one is compiled from a syntax tree instead. Zero API calls, and a stated runtime of two hundred milliseconds.
The practical consequences run in both directions. In its favour, the same input always produces the same output, the tool works offline and on an air-gapped machine, there is no per-scan cost, and nothing about your source code leaves the machine. For a tool whose output is fed straight into a model, that is the right set of properties.
Against it, the generated prose cannot exceed what the detectors found. The project is candid about this: the wiki is described as a narrative layer on top of data the codebase already contains, with the structure coming from analysis rather than from reading. So the articles are as good as the route table, the schema parse and the dependency graph behind them, and on a language handled by pattern matching, that is the weakest link.
Which is the argument for the AST and regex distinction in the first place, stated more sharply: you are trading narration quality for structure quality, and the structure is what the token saving depends on.
Knowledge mode classifies your notes with the same keyword heuristics
There is a second mode that does not look at code at all. Point it at a folder of markdown and it produces a context primer covering decisions, open questions and a note index:
npx codesight --mode knowledge # Scan current directory for .md filesThe targets are a personal knowledge vault, a project docs folder, or any directory of markdown. It understands Obsidian conventions including frontmatter, wiki-style backlinks and tags, Notion exports, and the output of ADR tooling.
Classification is by pattern, and the patterns are listed: a decision record is recognised by a decision heading or by phrases like decided to, going with, or chose one thing over another. Meeting notes by an attendees line, an action items line, or a filename containing standup, sync or one on one. Retrospectives by what went well, stop doing, or a retro filename. Specs by goals and requirements headings. Research and session logs by filename.
So the same property holds here as in the code path. Classification is cheap, deterministic and fast, and it misses anything phrased differently. A decision recorded without one of those phrases will not appear in the map, and the map gives no indication that it is missing. Treat the decision list as a starting index and confirm anything load-bearing.
The syntax tree comes from a checksummed prebuilt artifact
The AST path needs a parser that is not written in the same language as the tool, and the build scripts show how it is obtained. There is a reference build script that invokes a standalone parser compiler against a configuration in the reference directory, and a separate script that generates checksums for the result.
"build:reference": "asc --config reference/ast-plugin/asconfig.json""checksums:reference": "node reference/ast-plugin/gen-checksums.mjs"Generating checksums for a prebuilt parser artifact is a supply-chain precaution that most tools of this size skip. It means the binary the analysis depends on can be verified rather than trusted, which matters for a tool that runs on your source tree with no configuration and no prompts.
The rest of the manifest is conventional but slightly unusual in one respect. The prepare script runs the compiler, so installing from a checkout compiles TypeScript rather than shipping prebuilt JavaScript:
"prepare": "tsc"The package also exports several plugin surfaces as separate entry points, covering continuous integration, git hooks, skills and Terraform, which is how the tool extends beyond the one-shot scan. Tests run the build first and then the suite through a TypeScript-aware runner.
The repository carries its own generated output and a promotional README
Two things in the repository are worth noting. First, the tree includes a codesight output directory at its own root, alongside a documentation directory, an evaluation directory and a plugin directory. The tool is run on itself and the result is committed, which is either good dogfooding or an invitation to merge a stale map, depending on whether a hook regenerates it.
Second, the README reads as a landing page. It leads with a download count, then a dense line of capability claims, then the author's social accounts, two company sites, and links to two sibling projects in the same house, one of them a collection of skills for a coding agent and the other a plugin aimed at search visibility.
None of that makes the numbers wrong. Zero dependencies, a Node floor of 18, MIT licensing, one hundred and forty-nine tests, fourteen MCP tools and a stated test basis of twenty-five or more open source projects across fourteen languages are all checkable claims, and the licence is genuinely MIT with a citation file alongside it.
What the presentation does mean for an evaluator is that the adoption numbers and the capability counts come from the author rather than from an independent index, so weight them accordingly. The version in the manifest is 1.19.0, there are no published releases, and the last push is dated 27 July 2026.
Editorial conclusion
codesight is worth running on a TypeScript or JavaScript repository, where the generated wiki is derived from real structure rather than guesses, and the token saving on a large codebase is substantial. Set your expectations by language before you measure anything. On the other twelve languages the same output is produced by pattern matching, so treat the wiki as a map of where to look rather than as an account of what the code does. And run the wiki in watch mode from the start, because a stale map is worse than no map.
Frequently asked questions
How do I run codesight?
Run npx codesight in any project root. There is no config file, no setup step and no API key to supply. Node.js 18 or newer is required, and the package declares zero dependencies.
Does codesight use an AI model to analyse my code?
No. The generated wiki is compiled from a syntax tree rather than produced by a model, with zero API calls and a stated runtime of about 200 milliseconds. The TypeScript and JavaScript paths use AST analysis; the other supported languages use regex detection.
What does codesight --init generate?
It writes four agent configuration files at once: CLAUDE.md, .cursorrules, codex.md and AGENTS.md. If you only use one assistant, the profile flag generates an optimised configuration for a single named tool instead.
How does codesight use MCP?
Started with the mcp flag it runs as a server exposing fourteen tools, and three of those are for the generated wiki: fetching the index at session start, fetching a single named article, and a health check for orphan articles, missing cross-links and stale content.
Can codesight read my notes as well as my code?
Yes, in a separate mode. Pointing the knowledge mode at a folder of markdown produces a context primer covering decisions, open questions and a note index, recognising decision records, meeting notes, retrospectives, specs, research and session logs, and it understands Obsidian frontmatter and backlinks.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/houseofmvps-codesight)