# FastCode reads your repository so the model does not have to

> An HKUDS project that indexes code with tree-sitter, embeddings and BM25, exposes the result over MCP and a web UI, and treats token spend as a budget to be managed rather than a cost to be paid.

**HKUDS/FastCode** — "FastCode: Accelerating and Streamlining Your Code Understanding"

- Repository: https://github.com/HKUDS/FastCode
- Website: https://arxiv.org/abs/2603.01012
- Stars: 2,289 · Forks: 276
- Language: Python
- License: not declared
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/hkuds-fastcode

## A framework for spending fewer tokens on a repository

The README states the goal in the repository description and the title both: accelerating and streamlining code understanding. The specific claim is about context. FastCode is built on the observation that pasting files into a prompt is the expensive part of asking a model about a codebase, and that most of what gets pasted is never relevant to the question being asked.

The reported numbers frame it. The README puts FastCode at roughly three times faster than Cursor and four times faster than Claude Code, at 55 percent lower cost than Cursor and 44 percent lower than Claude Code, with up to ten times fewer tokens consumed. Those are the project's own benchmark figures, and they are the numbers to check against the published paper rather than to take on faith, since speed and cost comparisons in this area move quickly as the tools being compared move.

The repository is Python, with 2,289 stars, 276 forks and 16 open issues, and it carries a homepage pointing at an arXiv paper rather than a documentation site. The default branch is `main`, the last recorded push is 2026-07-06, and the topics list is empty, so the README and the paper are the two places where the design is explained.

## Semantic and structural indexing, held together

The representation layer is where most of the engineering went, and the README describes it in three parts.

The first is hierarchical code units. Instead of chunking files into fixed token windows, the index is built across levels: files, then classes, then functions, then documentation, using AST parsing. The dependency list makes that concrete, with `tree-sitter` and a grammar package per language, covering Python, JavaScript, TypeScript, Java, Go, C, C++, Rust and C#. `libcst` sits alongside the grammars, which is a hint that Python gets a real concrete syntax tree rather than an approximation.

The second is a hybrid index. Semantic embeddings and keyword search with BM25 are combined rather than used as alternatives, which is the usual answer to the weakness of embeddings alone: that they match concepts but miss exact identifiers, and that a query containing a function name the author chose unhelpfully will not retrieve that function. `sentence-transformers`, `faiss-cpu` and `chromadb` are the vector-side dependencies, and `rank-bm25` is the keyword side.

The third is graph modelling, and it is the part that makes structural navigation possible. Three relationship graphs are maintained: call, dependency and inheritance. That is what lets the system answer what a function touches by walking edges rather than by reading files, and `networkx` is the graph library in the requirements for exactly that reason.

## Navigation that tries not to read whole files

Given those indices, the README describes the retrieval behaviour in four pieces, and they read like a description of how a person looks through unfamiliar code.

Search runs in two stages. A first pass finds code that might be relevant, and a second pass ranks and organises those candidates for the specific question. Retrieval and reranking separated this way is common practice, but the reason it matters here is cost: the expensive model is only applied to a shortlist, while the cheap lexical pass handles the whole repository.

File browsing is deliberately limited to structure, so folder organisation and file patterns can be understood without file contents being pulled in. Connections are followed up to two steps, which is a bounded traversal rather than a graph-wide one. And code skimming reads the file's headlines instead of its body: function names, class definitions and type hints. That last one is the highest-leverage idea in the whole design, since a type hint tells a model most of what it needs about a signature at a small fraction of the cost of the body.

Taken together these are four ways of not paying for tokens. The system is not trying to understand more of the repository. It is trying to understand exactly enough of the part the question touches.

## Cost treated as a value to optimise

The retrieval layer has a companion in the decision layer, and the README describes it as budget-aware behaviour rather than as a heuristic.

Before processing anything, the system weighs five factors: confidence level, query complexity, codebase size, resource cost, and iteration count. That is a fairly explicit statement that retrieval is a control problem. Given the same query on a small repository with a confident first hit, the answer is to stop; given a broad question on a large repository with weak initial matches, the answer is to keep gathering.

Two follow-on behaviours are described alongside it. Learning is described as resource-optimised, adapting during the run to decide what else is worth fetching. And selection is described as value-first, prioritising high-impact and low-cost information and continuing until a stopping point is reached.

Whether those five factors are weighted by hand or learned is the kind of detail the paper presumably settles and the README does not, and it is worth reading the paper for if you plan to modify the loop. The practical upshot for a user is that query phrasing changes the bill, which is true of every tool in this space but more pronounced when the loop is explicitly tracking spend.

## Entry points: MCP, a web UI and an API

FastCode is meant to be used from somewhere else rather than as a destination. The README lists three ways in: an MCP server for editors and agents, a web interface for exploring a codebase, and an API for wiring it into a workflow.

The MCP path is the one that matters most for token spend, because it puts the index in front of an agent as a tool rather than in front of a human as a page. The release notes confirm this is recent rather than original: version 1.0.1, published 2026-02-25, adds MCP support as a new feature alongside two fixes, an OpenAI compatibility layer for `max_tokens` and `max_completion_tokens`, and an embedder change to use Apple Silicon MPS when it is available.

The entry points map onto the repository layout. `api.py` is the API, `mcp_server.py` is the MCP server, `web_app.py` with `web_interface.html` is the browser interface, `main.py` is the command line path, and `config/` holds configuration with `env.example` documenting environment variables. FastAPI and Uvicorn appear in the requirements for the API, Flask alongside them for the web layer, and `mcp[cli]` is the MCP dependency.

## The container build explains a lot

The Dockerfile is short enough to read in full, and it makes two decisions that are worth knowing about before you start the thing.

The base image is `python:3.12-slim-bookworm`, which matches the Python 3.12 badge in the README, and the system layer installs git and build-essential because tree-sitter grammars compile from source. Requirements are copied and installed before the application code so that the pip layer stays cached across code changes.

The interesting step comes next. The image pre-downloads the sentence-transformers model before the app code is copied, with a comment noting the cached layer is roughly 470MB and that copying the application code afterwards means code changes do not invalidate it. That is a small piece of build hygiene with a large practical consequence: without it, every rebuild of the app re-downloads a few hundred megabytes of model weights.

The embedding model named there is `paraphrase-multilingual-MiniLM-L12-v2`, and the effect of the split is visible in the port and start command:

```dockerfile
FROM python:3.12-slim-bookworm
WORKDIR /app
COPY requirements.txt ./
RUN pip install --no-cache-dir --retries 5 --timeout 60 -r requirements.txt
EXPOSE 8001
CMD ["python", "api.py", "--host", "0.0.0.0", "--port", "8001"]
```

The compose file adds a second service, `nanobot`, running as a gateway on a separate port, and the two share a repositories volume so the gateway can read what FastCode indexed. An environment variable points the gateway at the internal API address rather than the published port. The project also lists a small local model as a supported option, qwen3-coder-30b, which is the clearest signal that the design is meant to work without a frontier model behind it.

## What the benchmarks actually cover

Four benchmarks are named, and the README gives each a focus area, which is more informative than the names alone.

SWE-QA covers software engineering question answering, LongCodeQA covers extended code analysis, LOC-BENCH covers code localisation for bug reports and feature requests, and GitTaskBench covers real production repository workflows. Read together, they are four different shapes of the same job: answer a question about behaviour, summarise a long file without losing the important part, find the right place to change something, and carry out a multi-step task in a repository that is not a toy.

Code localisation is the benchmark most directly aligned with what the index is for, and it is also the one most sensitive to the retrieval design, since a two-stage search that returns the right file in the top three and the wrong one in the top one scores very differently depending on how the scoring window is set. Token efficiency and accuracy are reported together in the README, which is the right way to report them, since a system that returns less and also returns less is not saving anything.

The project also links to a demo video, and the README's stated accuracy result is that it outperforms the compared baselines across all four. Verify that claim against the paper before planning around it.

## Conclusion

FastCode is an argument about where the tokens should go, and the repository makes that argument concretely enough to disagree with. The claim is that a model does not need your codebase in its context window if it has an index it can query, and every design choice follows from that: structural parsing rather than chunking, a hybrid retrieval index rather than embeddings alone, sketches of files rather than whole files, and a budget-aware loop that decides when it has learned enough. The honest way to evaluate it is on the four benchmarks the project names, SWE-QA, LongCodeQA, LOC-BENCH and GitTaskBench, against the speed and cost numbers the README puts next to Cursor and Claude Code, on a repository you know well enough to tell when a search result is wrong. Start with the MCP server rather than the web UI, because that is where the token savings actually land, and pin it to a small local model such as qwen3-coder-30b if you want to see how much of the reported benefit is indexing and how much is the model behind it.

## FAQ

### What does FastCode do?

FastCode indexes a repository so a language model can query it instead of reading it wholesale. It builds an AST-based index across files, classes, functions and documentation, combines embeddings with BM25 keyword search, and maintains call, dependency and inheritance graphs, then serves results through an MCP server, a web UI and an API.

### How does FastCode reduce token usage?

Four mechanisms. Search runs in two stages so the expensive model only sees a shortlist, file browsing is limited to structure, connections are followed up to two steps, and code skimming reads only signatures, class definitions and type hints rather than file bodies. The decision layer also weighs confidence, query complexity, codebase size, cost and iteration count before fetching more.

### Which languages can FastCode index?

The README lists Python, JavaScript, TypeScript, Java, Go, C and C++, Rust and C#, and the dependency list carries a tree-sitter grammar package for each of those plus libcst for Python concrete syntax trees.

### How do I run FastCode?

Docker Compose builds an image from python:3.12-slim-bookworm and starts the API on port 8001, alongside a nanobot gateway service that shares the repositories volume. The image pre-downloads the sentence-transformers model before copying application code, so rebuilding after a code change does not re-fetch the weights. Without Docker, the API is started directly with api.py.

### Can FastCode use a local model instead of a hosted one?

Yes. The README lists small model support and names qwen3-coder-30b as a local option, and the v1.0.1 release notes show the embedder gaining Apple Silicon MPS support. That combination is worth testing if you want to know how much of the reported gain comes from indexing rather than from the model behind it.

## Sources

- [HKUDS/FastCode on GitHub](https://github.com/HKUDS/FastCode)
- [Issues](https://github.com/HKUDS/FastCode/issues)
- [Project website](https://arxiv.org/abs/2603.01012)
- [README](https://github.com/HKUDS/FastCode/blob/main/README.md)
- [Releases](https://github.com/HKUDS/FastCode/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hkuds-fastcode
