SmallCode: a terminal coding agent built for 8B to 35B local models
AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.
At a glance
- What is it?
- SmallCode is an npm-installable terminal agent that trades frontier-model assumptions for context budgeting, a forgiving tool-call parser and patch-based edits. The design is coherent, but the README leaves the RAG and escalation paths thin.
- Who is it for?
- Adopt SmallCode if you already run an 8B to 35B model on your own hardware and want a terminal agent that budgets context, decomposes work into a TODO file and applies search-and-replace patches rather than rewriting whole files. Do not adopt it if your model is 4B or smaller, if you need a documented rollback story, or if you expect the RAG harness to work without Python 3 and Git installed.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 50 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem SmallCode targets: agents that assume a frontier model
Most terminal coding agents are written for models with 128k or more of context and near-perfect JSON tool calling. SmallCode starts from the opposite premise. Its README states the target explicitly: local models between 8B and 35B parameters on consumer hardware. The comparison table in the README draws the line at OpenCode, which it describes as assuming frontier models such as Claude and GPT-5, while SmallCode compensates for small-model limitations through architecture rather than prompting.
The README is equally direct about who should not use it. Models of 4B parameters or fewer, it says, struggle with multi-step tool use and lose context across turns. Models above 35B do not need the adaptations and are better served by tools built for frontier models. That is a narrow band, and it is the honest one: an 8B model that can follow a two-step edit is a different customer from a 70B model that can plan a refactor.
The audience is therefore engineers who already run LM Studio, llama.cpp or Ollama locally and want an agent loop that fits the model they have, not the model they wish they had. The privacy column in the README's table is the other half of the pitch: fully local, no network needed.
How SmallCode compensates: context budget, TODO decomposition, patch edits
Four mechanisms show up in the README and in .env.example, and they are the real substance of the project.
Context is budget-managed rather than dumped. SMALLCODE_CONTEXT_BUDGET defaults to 70, and SMALLCODE_MAX_TOOL_RESULT_CHARS defaults to 8000, roughly 240 lines, which the comments say fits most source files in one read_file call. A read guard returns the first N lines plus an explicit instruction to use grep or read a smaller line range instead of silently truncating the middle of a file. The head size is tunable through SMALLCODE_READ_GUARD_HEAD_LINES, default 30.
Tool calling goes through what the README calls a forgiving multi-format parser, rather than assuming reliable JSON. Planning is decomposed into steps tracked in a TODO file instead of a single-shot plan. Editing is search-and-replace patch based, not a full file write.
The cache-split setting is the most interesting detail in .env.example. It keeps the system prompt stable across turns so llama.cpp can reuse its KV cache. Without it, the comment says, dynamic content such as memory and knowledge injected into the system prompt invalidates llama.cpp's context checkpoints every turn, producing the "erased invalidated context checkpoint" loop and full re-processing. That is a concrete llama.cpp behaviour, not a marketing claim, and it is the kind of thing only someone running small models locally would have hit.
Installing SmallCode and pointing it at a local model
The README gives two install routes. The npm route is the shortest. The package exposes the smallcode binary as well as an alias, smolv2.
npm install -g smallcode
cd my-project
smallcodeNode.js 18 or newer is required, with 20.x or 22.x recommended because prebuilt SQLite binaries exist for those LTS lines. For machines without Node, the README also documents prebuilt tarballs for Windows, macOS and Linux that bundle Node.js and native addons, installed through install.sh on Linux and macOS or install.ps1 on Windows. The script extracts to ~/.smallcode and adds it to PATH.
Configuration lives in a .env file in your project root. Two keys are required.
SMALLCODE_MODEL=your-model-name
SMALLCODE_BASE_URL=http://localhost:1234/v1The base URL points at any OpenAI-compatible endpoint, and 1234 is the port the README uses for LM Studio. If you cloned the repository instead of installing the package, the README's checkout quick start runs npm install and then node bin/smallcode.js directly.
If the fullscreen TUI misbehaves in your terminal, the README offers node bin/smallcode.js --classic for a readline fallback. The TUI enables raw mode, mouse tracking and bracketed paste; the README states that SmallCode restores the terminal on exit, on Ctrl+Z suspend, on termination and on crash, and suggests running reset if a kill -9 leaves the shell echoing escape sequences.
One optional dependency deserves a note before you install. better-sqlite3 powers the code graph and FTS5 memory search. Prebuilt binaries cover Node LTS on Linux, macOS and Windows; on non-LTS Node you need a compiler toolchain. The README is clear that if the build fails, SmallCode still works and falls back to JSON-based memory. To skip the native dependency entirely, the README gives npm install -g smallcode --omit=optional, at the cost of FTS5 memory search.
The RAG harness needs Python, Git and a config file you have to write
SmallCode ships a local GitHub RAG database, and the setup is heavier than the agent install. The README lists Python 3 and Git as requirements for the scraper and indexer. Indexing runs through an npm script, with a broader preset for a larger multi-language corpus.
npm run rag:index
npm run rag:index -- --preset broadAfter a global install the same job is available as smallcode-rag-index --preset broad. For custom repositories the README says to create .smallcode/rag/repos.json with preset, repos and chunking limits. It does not document the schema of that file beyond those three field names, and it points to docs/rag-harness.md for the full LM Studio and llama.cpp setup, the UI walkthrough, RAG config, indexing and web-fallback flow. If you intend to rely on retrieval rather than on the model's own context, that document is the one to read before you start, because the README alone does not tell you what a valid repos.json looks like.
Where SmallCode is the wrong tool
The model-size guidance is the first boundary, and the project states it against its own interest: 4B and below, and above 35B, are both out of scope. A 4B model will lose context across turns regardless of how the harness budgets it. A 70B model gains nothing from a forgiving tool-call parser and loses the frontier-oriented features that OpenCode and similar tools provide.
The optional dependency is the second boundary. On a non-LTS Node release, better-sqlite3 needs python3, make and a C++ compiler on Linux, Xcode Command Line Tools on macOS, or Visual Studio Build Tools on Windows. The fallback to JSON memory keeps the agent running, but memory search behaviour differs between a machine that compiled SQLite and one that did not, so two developers on the same team may not get the same agent.
The README also documents a warning that will confuse people: [email protected]: No longer maintained, traced through budget-aware-mcp to better-sqlite3. The README calls it a harmless upstream deprecation notice and notes there is no newer published version, so it cannot be silenced by a version bump. It is cosmetic, but it appears on every install and will generate support questions.
Finally, the README says nothing about rollback. The patch-based editing model is presented as a benefit over full file writes, which it is for token cost, but there is no documented undo, no checkpoint command and no stated recovery path if a patch lands wrong. For a tool that edits source files in place, that silence matters more than any feature on the comparison table.
SmallCode versus OpenCode: different assumptions, different failure modes
The README's own comparison puts OpenCode on the frontier-model side: cloud API calls, full file writes, single-shot planning and an assumption of reliable JSON tool calling. SmallCode inverts each of those. The difference is not quality, it is which failure mode you prefer.
With OpenCode and a frontier model, the failure mode is usually cost, latency and sending your code to a provider. With SmallCode and an 8B model, the failure mode is a turn that produces a malformed tool call, an edit that misses, or a plan that stalls. SmallCode's architecture is a set of mitigations for that second failure mode: parse leniently, budget the context, decompose the plan, patch instead of rewrite. If your model is strong enough that those mitigations are unnecessary, they are overhead.
The escalation option is the middle path. SmallCode can fall back to a cloud provider on hard failure, configured through ANTHROPIC_API_KEY, OPENAI_API_KEY, OPENROUTER_API_KEY or DEEPSEEK_API_KEY in .env. The README's truncated example also describes routing each model tier to a different endpoint, keeping fast work local and sending complex tasks to a larger OpenRouter model. That is a reasonable design, but it is documented as an optional block in .env.example rather than as a worked configuration, so expect to experiment with the routing keys yourself.
Licence, maintenance and upgrade cost
SmallCode is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is a permissive arrangement with no copyleft obligation on your own code. It says nothing about the licences of your local model weights or of the model server you point SMALLCODE_BASE_URL at, and those are separate questions you have to answer yourself. Nothing here is legal advice.
The repository is not archived. Its last push was on 2026-08-12, about five weeks before this writing, and the most recent tagged release in the repository is v1.6.0 from 2026-05-31, following v1.5.2 and v1.5.0 in the same two-day window. That pattern, two releases in two days and then a gap, suggests bursts of work rather than a steady cadence. The package version in package.json matches v1.6.0, so the published npm package and the repository tag line up.
Upgrade cost is dominated by the native dependency, not by the JavaScript. If you are on Node LTS, upgrading is an npm install. If you are on a non-LTS line, every Node upgrade can invalidate the prebuilt better-sqlite3 binary and force a rebuild, which is the kind of recurring cost that does not show up in a changelog. The configuration surface is also wide: .env.example documents context budget, cache split, context window, tool result size, read guard and read guard head lines, each with a default that may or may not suit your model. Treat the defaults as a starting point to measure, not as a tuned configuration.
Editorial conclusion
Adopt SmallCode if you already run an 8B to 35B model on your own hardware and want a terminal agent that budgets context, decomposes work into a TODO file and applies search-and-replace patches rather than rewriting whole files. Do not adopt it if your model is 4B or smaller, if you need a documented rollback story, or if you expect the RAG harness to work without Python 3 and Git installed. Before committing, verify three things yourself: that your local server answers at the SMALLCODE_BASE_URL you set, that your model name matches what the server actually loads, and whether better-sqlite3 compiled on your Node version, because the JSON memory fallback changes what the agent remembers between turns.
Frequently asked questions
What is SmallCode?
SmallCode is a terminal-native AI coding agent optimized for small LLMs in the 8B to 35B parameter range, installable from npm as the smallcode package. The README states it targets local models on consumer hardware and compensates for their limitations through context budgeting, a forgiving tool-call parser, TODO-file planning and search-and-replace edits.
Which AI platform is best for code?
The README does not rank platforms. SmallCode's README positions itself against OpenCode, describing OpenCode as built for frontier models such as Claude and GPT-5 over cloud APIs, while SmallCode targets 8B to 35B local models and states that models above 35B are better served by tools designed for frontier models.
Is coding going away due to AI?
SmallCode's documentation does not address this. What it does show is a tool that assumes a human runs it in a project directory, reviews patch-based edits and configures a local model server, which is a different picture from fully autonomous code generation.
What are some good examples of small coding projects for beginners?
The README does not offer beginner project examples. It describes SmallCode itself as a coding agent, and its own bench suite names smoke, polyglot-mini and tool-use suites, which are evaluation harnesses rather than tutorial projects.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/doorman11991-smallcode)