ChatLLM Web: 4K contexts, eight agent steps, and an agent that excludes MCP
Private local model studio, AI chat, and agent workspace powered by WebGPU and WebLLM.
At a glance
- What is it?
- ChatLLM Web is a static web application that runs language models in the browser through a WebGPU runtime, with a chat workspace, an agent workspace and a model studio in one page and no backend at all. Its limits are stated precisely enough to plan around: a four-thousand-token window on every curated model, an eight-step cap on agent runs, and a runtime boundary that deliberately excludes several things people expect.
- Who is it for?
- ChatLLM Web suits someone who wants a model studio that runs entirely on their own machine, keeps every conversation and artefact in the browser, and never sends a prompt to a server it did not choose. Before you point it at a large codebase, read the two ceilings: every curated model has a four-thousand-token window, and exceeding it drops whole earlier turns rather than summarising them.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 55 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Every curated model has a 4K window, and older turns disappear
One constraint shapes everything else in this product: every model in the curated list is offered with a four-thousand-token context window. That is small next to what a hosted assistant will give you, and it is fixed rather than negotiated per model. The mitigation is honest but lossy. When the input budget for a conversation is exceeded, the application omits complete older turns rather than truncating the most recent one mid-thought, and the activity trace reports that it happened. So a long conversation does not error; it quietly becomes a shorter conversation with a note in the log. For a tool meant to read files you attach, that is workable, and for a tool meant to reason over a large repository it is the constraint that decides whether the project fits. The per-conversation settings, including temperature, top-p, output limit and system prompt, are stored per conversation, which is the other half of how you keep two different jobs apart inside a 4K budget.
Sixty-five logical models, one hundred and sixty-three records
The catalog is the part of the product that most needs explaining, and the numbers are the explanation. The curated set is twenty models split across a stable tier and an experimental tier, with declared memory requirements from a few hundred megabytes to several gigabytes. Behind it, an advanced view exposes sixty-five logical models, and those are backed by one hundred and sixty-three records in the upstream runtime directory. The gap between sixty-five and one hundred and sixty-three is not a discrepancy; it is quantization and context-length variants grouped under one logical model so the picker does not show you four copies of the same thing. The selection logic is explicit: if the graphics stack reports support for half-precision shaders, it picks the sixteen-bit quantization, and where the directory offers a thirty-two-bit variant it falls back to that. The original model identifier is kept attached to every conversation and every generated message, which is what stops a swap mid-conversation from quietly rewriting which model produced a reply.
The agent reviews diffs and stops itself after eight steps
The agent workspace is where the design is most deliberate. Six tools are exposed and each carries an access scope and an approval mode. Reading the selected context files, searching them, running arithmetic in a sandboxed parser and reading back artefacts from the conversation are automatic. Creating and updating an artefact require a diff approval first, which means the model proposes text and you see the line diff before anything is written to the browser database. Artefacts have a hard size and length ceiling, stated as the reason the approval diff stays reviewable rather than as a quota. A run rail records planning, tool calls, approvals, completions, stops and errors, and each run is capped at eight steps. Stopping preserves the steps that finished, and retrying starts a fresh run from the original task rather than resuming a half-finished one, which is the safer default when the failure might have been caused by the accumulated state.
An agent workspace with MCP, RAG and shell deliberately left out
The runtime boundary is stated as a list of exclusions, and the list is long enough to be the most important paragraph in the documentation. The agent's tools never receive arbitrary disk access. Shell execution, direct filesystem writes, network search, protocol servers, retrieval augmentation and third-party data sources all stay outside the current version's boundary. That is a strong constraint for something billed as an agent workspace, and it is stated without hedging rather than as a roadmap item. The reasoning is visible in the design of what remains: every tool is scoped to files you selected for this run or artefacts created in this conversation, arithmetic runs in a parser rather than an interpreter, and the only writes are ones you approved. In exchange, the application needs no keys, no accounts and no server, which is also why it can be served as static files.
Model choice is driven by reported memory, not by a leaderboard
The recommendation logic is device-first and the readme says so. If the graphics stack is available and the browser reports at least eight gigabytes of device memory, the suggested model is a two-billion-parameter general assistant. On an unknown device, or one reporting less, it starts with a one-billion-parameter chat model instead. The model studio shows the evidence behind that decision as cards covering the graphics stack, reported memory, browser storage quota and the runtime's own status, which is more useful than a bare model name because it tells you why the answer is what it is. The curated tiers are labelled by capability as well as stability, with filters for coding, reasoning, vision, tools and experimental entries, and downloads require explicit approval. Nothing here is selected by a benchmark table, so if you need a specific capability you are picking from the capability labels rather than from a score.
No backend, no account, and model weights from two declared hosts
The deployment story is short because there is almost nothing to deploy. The application is static and served from a page-hosting service, model assets are fetched directly from the addresses the model entries declare on a model hub and on the runtime's own library host, and the project states plainly that there is no application backend, no account, no API key, no analytics and no telemetry. Conversations, selected files, tool results and artefacts stay on the device, and the app can be installed as a progressive web application so cached models survive the first load and work offline afterwards. The consequence for a reader is a specific privacy property rather than a general promise: nothing about your prompts has anywhere to go, because there is no server component that receives them. The consequence for an evaluator is that every feature that needs a network call, including fetching a model and the advanced manifests, is visible as a download you approve first.
Three TypeScript configs, a legacy folder, and no workflow directory
The repository layout is modest for a product with this many screens. There is a source directory, a public directory with the static shell, a tests directory, two configuration files for the bundler and the test runner, an entry HTML file, and three TypeScript configurations split between the application and the tooling. Two entries deserve comment. There is a legacy directory kept at the root, which usually means an earlier build of the same product is preserved for reference rather than deleted, so anyone reading the source should check which of the two is live. There is also a document in the root that reads like a design review checklist, which is the sort of file that tells you what the team considers finished. What is absent is equally informative: no workflow directory appears in the tree, so the lint, typecheck and test scripts the manifest declares are run by hand or by an external service rather than by a pipeline in the repository.
Editorial conclusion
ChatLLM Web suits someone who wants a model studio that runs entirely on their own machine, keeps every conversation and artefact in the browser, and never sends a prompt to a server it did not choose. Before you point it at a large codebase, read the two ceilings: every curated model has a four-thousand-token window, and exceeding it drops whole earlier turns rather than summarising them. Decide whether the agent's narrow toolset is enough, because shell access, direct file writes, network search, protocol servers and retrieval are all outside the current boundary by design. And check your graphics support and reported memory before choosing a model, since the recommendation is driven by device capability rather than by benchmarks.
Frequently asked questions
What is ChatLLM and is it related to other products with similar names?
This project is a browser-only model studio that runs models locally through a WebGPU runtime, with no subscription, account or backend. It is unrelated to similarly named products that people also search for, which is worth knowing before you look for pricing.
Why does ChatLLM Web drop earlier messages from a conversation?
Every curated model is offered with a four-thousand-token window. When that budget is exceeded, complete older turns are omitted rather than partially kept, and the activity trace reports that it happened.
What can the ChatLLM Web agent actually do?
Six tools: read and search the files you selected, run arithmetic in a sandboxed parser, list and read artefacts from the conversation, and create or update an artefact behind a diff approval. Shell execution, direct file writes, network search, protocol servers and retrieval are all outside the boundary.
How does ChatLLM Web choose which model to recommend?
From the device rather than a benchmark: if WebGPU is available and the browser reports at least eight gigabytes of memory it suggests a two-billion-parameter general model, and unknown or lower-memory devices start with a one-billion-parameter chat model.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ryan-yang125-chatllm-web)